A geodesic in Wasserstein space · hover the path · drag to rotate
Scientific data increasingly arrive as objects such as probability distributions, symmetric positive definite matrices, networks, that live on curved spaces, not in ℝⁿ. My research builds the statistical foundations to model them where they live.
I am an Assistant Professor in the Department of Industrial and Systems Engineering at the University at Buffalo. My group works on distribution-in-distribution-out analytics, regression, inference, and change detection when both inputs and outputs are random objects on manifolds, and on nonparametric estimation on Riemannian manifolds.
On the applied side, I collaborate closely with anesthesiologists and surgeons on AI for perioperative medicine, kidney transplantation, and with manufacturers on computation pipelines, interpretable neural networks, and AI incubation for smart manufacturing.
Education
2021
Ph.D., Industrial & Systems EngineeringVirginia Tech
2026Gold Reviewer, International Conference on Machine Learning (ICML)
2025Runner-up, INFORMS Data Mining Best Paper Competition (General Track)
2023Collaborative Science Award, American Heart Association
2019Doctoral Student of the Year, Dept. of ISE, Virginia Tech
02
Research
Statistics on Manifolds & Distribution-valued Data
Regression, inference, and online monitoring when data are probability distributions or points on Riemannian manifolds. Current work includes distribution-in-distribution-out models, harmonic map regression with topological recovery, Riemannian OLS, and intrinsic cross-covariance.
Wasserstein geometryFréchet meanschange point detection
AI for Perioperative Medicine and Transplantation
Machine learning with clinicians in the loop. We develop real-time forecasting of intraoperative hypotension, prediction of cardiac-surgery-associated acute kidney injury, transferable discriminant analysis, and language models for kidney-donation decision support.
hypotensionCSA-AKIorgan transplant
Industrial AI & Smart Manufacturing
Adaptive computation pipelines for cyber-manufacturing, interpretable neural networks for quality modeling, ensemble active learning for AI incubation, and fog and edge analytics, deployed with partners in additive manufacturing, fiber, and food supply chains.
AutoML pipelinesinterpretable NNfog computing
The Interactive Playground
Each figure below distills one of the papers into a small interactive demonstration that runs live in your browser.
S² · drag to rotate
FOUNDATIONS · WHY GEOMETRY MATTERS
The Fréchet Mean on the Sphere
Averaging data that live on a curved space is already non-trivial. The Euclidean average of points on a sphere falls inside the sphere and is therefore not a valid data point. The Fréchet mean instead minimizes the sum of squared geodesic distances, staying on the manifold.
Spread the cluster and watch the two means diverge. The flatter your view of the world, the more Euclidean shortcuts cost you.
samplesFréchet meanEuclidean mean (inside!)
What am I looking at?
The Fréchet mean is computed live by Riemannian gradient descent. Points are lifted to the tangent space at the current estimate via the logarithm map, averaged, and mapped back with the exponential map, and the update is iterated until convergence. This intrinsic-vs-extrinsic gap is the starting point for all of the statistics below.
response space · drag to rotate
PREPRINT · arXiv:2405.11626 · NSF CAREER TOPIC
Distribution-in-Distribution-out Regression
What if both the predictors and the response are probability measures? Our DIDO regression model ν = (β1⊙T1 ⊕ ⋯ ⊕ βp⊙Tp)♯(ν̄) ⊕ ε defines the addition ⊕ and scalar multiplication ⊙ of optimal transport maps via parallel transport, which makes the operations commutative and additive, so a linear structure exists in Wasserstein space. The coefficients are estimated by Fréchet least squares, and the Fréchet Gauss-Markov theorem shows the estimator is the best linear unbiased estimator.
The response ν is itself a probability measure and appears here as a single point on the response space. Prediction starts at the Fréchet barycenter ν̄ of the responses and is pushed forward along one geodesic per predictor. Each segment of the polyline is driven by the predictor measure μⱼ annotated beside it. With three predictors, the response is pushed forward three times before it arrives at the prediction ν̂.
The surface stands in for the space where response measures live, and every point of it is an entire distribution. In the DIDO model, Tⱼ is the optimal transport map associated with the j-th predictor measure, each contribution βⱼ⊙Tⱼ is combined through the parallel-transport addition ⊕, and the combined map is applied at the barycenter ν̄. Move a slider and watch its geodesic stretch, bend the polyline, and relocate ν̂. The Fréchet least squares estimator has a closed form based on the Wasserstein covariance, and in the Gaussian special case the model reduces to the simple closed-form Gauss-DIDO regression. This is the theme of my NSF CAREER project on distribution-in-distribution-out analytics.
streaming · one distribution per tick
ICML 2026 · ACCEPTED (26.6% RATE)
Beyond Euclidean Summaries: Online Change Point Detection for Distributions
Each time step delivers a whole distribution, such as a batch of sensor readings or a histogram of intraoperative blood pressures. Classical monitoring tracks a mean and misses changes in shape. Our detector compares windows of distributions through their Wasserstein-Fréchet summaries, catching location, scale, and shape shifts online.
Press inject change and watch the statistic climb past the threshold. Try a pure variance change, where the mean barely moves but the detector still fires.
The statistic shown is a windowed Wasserstein-2 distance between the Fréchet means of a reference window and the most recent window of distributions. In the ICML 2026 paper (with Y. Zeng and Y. Huang), the detector comes with false-alarm control and detection-delay guarantees for general distribution-valued streams.
γ̂ : S¹ → T² · drag to rotate
PREPRINT · arXiv:2604.09513
Harmonic Map Regression
Regression where the response lives on a manifold. The task is to estimate a smooth loop on the torus from noisy observations. Our estimator minimizes a penalized harmonic (Dirichlet) energy, achieving rate-optimal estimation and recovering the topology of the underlying map, which here means the loop’s winding numbers around the torus.
Slide the smoothing parameter λ. Small values chase every noisy point, while larger values relax the loop toward the harmonic representative of its homotopy class. The readout tracks the fitted winding numbers, the topological signature the estimator is designed to recover.
noisy observationsharmonic fit γ̂true curve γ
Chen, X. “Harmonic Map Regression: A Rate-optimal Nonparametric Estimation on Manifolds with Topological Recovery.” arXiv:2604.09513
What am I looking at?
The fit minimizes a discretized penalized objective that combines data fidelity with Dirichlet (harmonic) energy, using gradient descent in the flat metric of T² with angle wrap-around handled intrinsically. Because smoothing acts on the manifold rather than in an ambient embedding, the loop cannot tear. For suitable λ the estimator attains the optimal nonparametric rate and provably recovers the homotopy class, which here is the winding pair (2,3). An extrinsic smoother in ℝ³ would happily cut through the tube and destroy exactly this structure.
X ∈ S² , Y ∈ S² · paired samples
PREPRINT · arXiv:2606.10212
Footpoint-invariant Riemannian Cross-covariance
How do you measure dependence between two manifold-valued variables? Naïve tangent-space covariances depend on the base point where you compute them, known as the footpoint. Our intrinsic construction is invariant to the footpoint thanks to parallel transport.
Dial the coupling ρ and watch paired samples on the two spheres co-move; the glyph in the middle is the cross-covariance operator’s strength and principal direction.
Samples are generated by a shared latent deformation plus independent noise on each sphere; the glyph summarizes the empirical tangent-space cross-covariance after parallel transport to the Fréchet means. In the paper, this operator is shown to be independent of the chosen footpoints, which enables well-defined canonical correlations between manifold-valued variables.
03
Publications
* graduate student advisee · bold = X. Chen
04
Students
Current advisees
Ph.D. · exp. 2026
Zehua (Jerry) Dong
Long-short term dual agent reinforcement learning for transferable knowledge preservation.
Aug 2023 – present
Ph.D. · exp. 2028
Yujing (Zipan) Huang
Deep distribution-in-distribution-out modeling and inference.
Jan 2024 – present
Ph.D. · exp. 2029
Jungyoon Moon
Topic in development.
Aug 2025 – present
Alumni
M.S. 2026
Nithin Rogan Arulsamy
Knowledge-graph-based hybrid agentic recommender system for dietary microbiome intervention.
2024 – 2026
Ph.D. 2025
Joshua Nielsen
Living kidney donation, addressing the transplant supply gap with machine learning and NLP. Houchens Prize for outstanding dissertation. Co-advised with Dr. Monica Gentili (UofL).
2021 – 2025
M.S. 2025
Shuoling Li
Transferable survival analysis via expectation maximization. Winner, IISE 2026 Graduate Research Award.
2023 – 2025
M.S. 2022
Jodie Ritter
Forecasting hypotension by learning from multivariate mixed responses.
2021 – 2022
Student honors
Shuoling LiWinner, IISE Graduate Research Award, 2026
Joshua NielsenHouchens Prize for outstanding dissertation, University of Louisville, 2025
Yujing (Zipan) HuangRunner-up, INFORMS Data Mining Best Paper Competition (General Track), 2025