STA 701S meets here

Day
Mondays
Time
11:45 AM–1:00 PM
Location
Old Chem 116

STA 701S Duke Department of Statistical Science

Graduate Research
Seminar

Research in progress. Ideas in conversation.

An advanced seminar at the research frontiers of statistical science. Third-year students develop ideas that may grow into research; fourth- and fifth-year students present research in progress or completed work. Every speaker builds an accessible introduction to the problem, area, and central ideas, creating space for students and faculty to ask sharper questions and deepen understanding through discussion.

Course STA 701S
Offered Fall and Spring
A seminar conversation A conceptual diagram showing a statistical idea being explained, questioned, and understood more deeply through seminar discussion. 01 EXPLAIN 02 QUESTION 03 DEEPEN new connection
Explain the idea. Invite the room in. Leave with a deeper understanding.

01 / The seminar

A forum for statistical ideas, clearly explained.

A strong presentation makes an interesting idea accessible.

STA 701S is a working forum for learning across areas and practicing the clear communication of technical ideas. Presenters choose a topic worth understanding, build a thoughtful and engaging talk, and create a conversation that helps speaker and audience think more carefully.

Choose

Select an interesting statistical idea, establish why it matters, and make it accessible without sacrificing depth.

Explain

Build a focused, accurate, and engaging presentation that gives the audience a clear path into the idea.

Discuss

Use constructive questions to test understanding, surface assumptions, and connect ideas across statistical science.

02 / Participation

Every talk is a shared responsibility.

Speakers bring an idea worth understanding; the room helps make sense of it.

  1. Presenters prepare a focused, high-quality talk that gives the audience enough context to engage deeply.

  2. Third-year students should develop and present ideas that may grow into research; fourth- and fifth-year students should present their research in progress or completed work. In every case, the talk should serve as an accessible introduction to the problem, area, and central ideas—not as a dense technical treatise.

  3. Participants listen actively, ask constructive questions, and help the room reach a deeper understanding.

Each presentation is no more than 20 minutes, followed by a 10–15 minute discussion period.

03 / Calendar

Fall 2026 seminar calendar.

The current speaker calendar is shown below. Presentations are grouped by meeting date.

Student submission guide

All seminar meetings

Mondays 11:45 AM–1:00 PM Old Chem 116

    1. Alexander Volfovsky

      Year 5 · fall-2026-17

      Example of what output looks like: What Does a Model Expect? Posterior Predictive Checks for Statistical Practice

      Read abstract

      Statistical models are useful because they let us compare what we observed with what the model says could have happened. This talk introduces posterior predictive checks as a practical way to make that comparison visible, using simple examples to connect simulation, model criticism, and uncertainty. The goal is an accessible introduction to how a model can reveal its own failures, and to what these checks can and cannot tell us about a completed analysis.

    2. Alexander Volfovsky

      Year 5 · fall-2026-18

      Introduction to bad talks

      Read abstract

      This second example illustrates a talk title without an uploaded presentation, so students can see how a non-linked schedule entry appears on the website.

    1. Lorenzo Mauri

      Year 5 · fall-2026-01

      Overfitted high-dimensional matrix factorizations via adaptive spectral shrinkage

      Read abstract

      Factor models are popular approaches for analyzing high-dimensional data to extract low-rank signals and estimate covariances. They decompose the covariance matrix as the sum of low-rank and diagonal components. A key issue is how to choose the latent dimension $k$, which is particularly challenging when the factor model only holds approximately and in low signal-to-noise scenarios. Bayesian overfitted factor models specify an upper bound on $k$ and rely on structured shrinkage priors to effectively remove extra components. Such approaches are popular and effective, but computationally expensive. We propose a much faster EigenBayes approach that provides valid uncertainty quantification, based on spectral estimation of latent factors and adaptive empirical Bayes calibration of key hyperparameters. The resulting posterior distribution factorizes across outcomes and is analytically tractable, bypassing Markov chain Monte Carlo. We show that EigenBayes adapts to the signal-to-noise ratio of each outcome and latent dimension, while shrinking superfluous latent components to zero. We establish favorable asymptotic properties and demonstrate strong empirical performance in numerical experiments and a genomics application, where EigenBayes outperforms state-of-the-art alternatives.

    1. Leah Johnson

      Year 5 · fall-2026-03

      Exploratory Permutation Testing for Treatment-effect Heterogeneity with XGBoost

      Read abstract

      Treatment effects rarely apply uniformly across patients, and average effects can obscure clinically important heterogeneity. We consider global testing of treatment-effect heterogeneity in randomized trials, where flexible models like XGBoost capture nonlinear covariate interactions but complicate inference, since the null distribution of any resulting test statistic is intractable analytically. We address this with permutation testing, which calibrates the full model-fitting pipeline by refitting under each permutation. We develop two approaches: an general A-learning framework that uses SHAP-based feature importance for clinical interpretability, and a pseudo-outcome approach that directly targets a chosen treatment-effect contrast at the cost of outcome-specific construction. We compare both methods across continuous, binary, and time-to-event outcomes in simulation, and discuss their trade-offs for exploratory precision-medicine analyses in pharmaceutical development.

    2. Piotr Suder

      Year 5 · fall-2026-04

      Alleviating Sim-to-Real Gap in Bayesian Optimization via Online Digital Twin Calibration

      Read abstract

      Bayesian optimization (BO) is widely used for optimizing expensive black-box systems, but its performance often degrades when relying on simulators/digital twins that exhibit model mismatch due to unmodeled dynamics, noise, or physical effects. This sim-to-real gap can lead to poor real-world performance despite strong results in simulation. We propose an online digital twin calibration approach for BO that adaptively aligns simulations with real-world feedback. The true objective and sim-to-real gap / discrepancy are modeled independently using Gaussian process (GP) surrogates, yielding a principled multi-fidelity surrogate that combines cheap, biased simulations with intermittent, unbiased high-fidelity evaluations. We formulate a deep GP surrogate which allows us to flexibly model the lengthscales of the discrepancy function and search specifically for the simulator calibration most useful for the underlying optimization task. We introduce the expected improvement variance reduction (EIVAR) acquisition function which targets reducing posterior uncertainty of the true objective in promising regions by jointly selecting design variables and calibration parameters. We provide an efficient approximation algorithm for computing and optimizing EIVAR. Our method improves sample efficiency and outperforms standard multi-fidelity and offline calibration baselines.

    1. Yen-Chun Liu

      Year 5 · fall-2026-05

      QuIP: Experimental design for expensive simulators with many Qualitative factors via Integer Programming

      Read abstract

      The need to explore and/or optimize expensive simulators with many qualitative factors arises in broad scientific and engineering problems. Our motivating application lies in path planning – the exploration of feasible paths for navigation – which plays an important role in robotics and assembly planning. For complex settings, the parameter space for path exploration can be discrete and high-dimensional, and the evaluation of path feasibility requires expensive virtual simulations. A carefully selected experimental design is thus essential for timely decision-making. We propose here a novel framework called QuIP, for experimental design of a Gaussian process (GP) surrogate with Qualitative factors via Integer Programming. QuIP leverages a GP surrogate with an exchangeable covariance function. For initial design, we show that its maximin design can be formulated as an assignment problem from operations research, which can be efficiently and globally optimized via state-of-the-art integer programming solvers. For sequential design (specifically, for active learning or black-box optimization), we show that its design problem can similarly be formulated as an assignment problem, which facilitates efficient and reliable optimization with state-of-the-art solvers. We demonstrate the effectiveness of QuIP over existing design methods in a suite of path planning experiments and an application to rover trajectory optimization.

    2. Gwen Jacobson

      Year 4 · fall-2026-06

      Initial Assessments of Covariate-Adjusted Average Treatment Effect Estimation under Misspecified Multiple Imputation Models

      Read abstract

      In completely randomized experiments, regression adjustment for pre-treatment covariates can effect greater efficiency in average treatment effect (ATE) estimation compared to using the unadjusted difference-in-means estimator. However, analysts sometimes contend with covariate missingness, complicating the adjustment approach. To overcome this missingness without discarding data, analysts use multiple imputation (MI) to generate multiple completed data sets. However, another question arises: What happens if the imputation model is mis-specified? Use of a mis-specified imputation model may, or may not, still improve the inferences. In this research, we investigate this question with the goal of providing guidance to analysts on when they might benefit from using multiple imputation with missing covariates in randomized experiments. We present preliminary results from a simulation study where the mis-specified model corresponds to an unintentional omission of an important interaction, as might happen in default applications of popular routines like multiple imputation by chained equations. We examine the relationship between estimation efficiency and interaction strength. We also preview ongoing work on additional experimental settings, future methodological comparisons, and potential diagnostics to provide guidance on the use of MI in ATE estimation.

    1. Caitrin Murphy

      Year 5 · fall-2026-08

      Multi-Armed Bandits for Batch Recommendations in Online Retail

      Read abstract

      To remain competitive and maximize revenue, online retailers must respond to customer search queries with personalized lists of items that are both relevant to the query and appealing to the customer’s individual preferences. Discrete choice models are commonly used to model purchase behavior in this setting. However, such approaches typically rely on the assumptions that customers examine every displayed item before making a purchase decision, and that utilities are determined entirely by observed information. In practice, customers may lose interest before reaching the end of a list, making the composition and ordering of the display important decisions for retailers. Additionally, a search query may reveal only part of a customer’s preferences, creating uncertainty about the utility assigned to each item. Failure to account for limited attention or unobserved preferences can lead to suboptimal recommendations and reduced revenue. We propose two extensions to the multinomial logit model that address these limitations. First, each customer is associated with a latent preference vector that represents tastes not revealed by the query. Second, the number of items examined follows a discrete-time survival process, so position effects arise from browsing behavior rather than from exogenous weights. We formulate the construction of personalized displays as a contextual bandit problem and propose a sequential recommendation policy using the SquareCB algorithm. We show that, with high probability, the regret is sublinear. Simulation experiments quantify the value of modeling latent preferences and limited attention.

    2. Vivek Singh

      Year 4 · fall-2026-07

      A Stable Matching Approach to Distribution-free Independence Testing

      Read abstract

      Testing independence between random vectors requires detecting dependence without specifying their marginal distributions. Multivariate ranks based on optimal transport address this problem by matching observations to a fixed reference set, extending the distribution-free properties of univariate ranks. We introduce an alternative construction of multivariate ranks based on stable matching. Instead of minimizing a global transport cost, the assignment satisfies a local stability condition. Under standard regularity conditions, the stable assignment is unique and the resulting rank vector is distribution-free. At the population level, measure transportation, preservation of dependence, and convergence of empirical ranks provide the basis for consistent independence tests using test statistics such as distance covariance. Stable matching also permits different choices of utility, producing rank maps that can lead to different power against dependence alternatives. We compare stable matching with optimal transport in terms of computation, the resulting maps, and testing power, with particular interest in heavy-tailed settings.

    1. Kateryna “Kat” Husar

      Year 5 · fall-2026-09

      Title and abstract forthcoming.

    2. DongKyu “Derek” Cho

      Year 4 · fall-2026-10

      Title and abstract forthcoming.

    1. Suchismita Roy

      Year 5 · fall-2026-11

      Title and abstract forthcoming.

    2. Shuo Wang

      Year 4 · fall-2026-12

      Title and abstract forthcoming.

    1. Gyeonghun “Hun” Kang

      Year 4 · fall-2026-16

      Title and abstract forthcoming.

    2. Marie Neubrander

      Year 4 · fall-2026-14

      Title and abstract forthcoming.

    1. Li Fan

      Year 4 · fall-2026-13

      Title and abstract forthcoming.

    1. Patrick Woitschig

      Year 4 · fall-2026-02

      Title and abstract forthcoming.