Seminar

Making a Training Dataset from Multiple Data Distributions

Over time we might accumulate lots of data from several different populations: e.g., the spread of a virus across different countries. Yet what we wish to model is not any one of these populations. One might want a model for the spread of the virus that is robust to the different countries, or is predictive on a new location we have only limited data for. We overview and formalize the objectives these present for mixing different distributions to make a training dataset, which have historically been hard to optimize. We show that by assuming we train models near "optimal" for our training distribution these objectives simplify to convex objectives, and provide methods to optimize these reduced objectives. Experimental results show improvements across language modeling, bio-assays, and census data tasks.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Efficient Estimation and Closed-Form Uncertainty Quantification for Net Benefit of Algorithms that Predict Individualized Treatment Benefit

Treatment benefit predictors (TBPs) quantify the expected treatment benefit given individual characteristics. The net benefit function evaluates the expected gain in clinical utility from using a TBP to guide treatment decisions, relative to the default decisions of treating no one and treating everyone. The existing estimator for the net benefit of a given TBP implicitly assumes a 1:1 randomization design in the randomized controlled trial data. When this assumption is violated, the estimator can exhibit biased and unstable finite-sample behaviour. We identify the source of this implicit design restriction and propose a corrected congruent-based modification of the net benefit estimator that remains valid under arbitrary randomization schemes. We further introduce an alternative net benefit formulation based on the average treatment effect among individuals recommended for treatment by a TBP (ATT-based). The ATT-based estimator is asymptotically equivalent to the corrected congruent-based estimator and makes more efficient use of the full sample by estimating the proportion recommended for treatment using all individuals. Three uncertainty quantification procedures are developed for the ATT-based estimator: a large-sample variance approximation, an aggregated nonparametric bootstrap, and a Bayesian analysis. A Monte Carlo simulation study evaluates point estimation of net benefit curves using mean squared error, while uncertainty quantification is assessed through confidence interval length and coverage probabilities for 95% intervals constructed using asymptotic, bootstrap, and Bayesian methods. Simulation studies show that the corrected congruent-based estimator removes the bias under unequal randomization, while the ATT-based estimator demonstrated improved finite-sample efficiency. An empirical application to the GUSTO randomized controlled trial illustrates the ATT-based estimator and accompanying uncertainty intervals in a real-world setting. Overall, the proposed methodology enables valid estimation and uncertainty quantification of net benefit for treatment benefit predictors under arbitrary randomization schemes.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Sasha Sharma

Topics in Trend Filtering with Poisson Loss

Many types of data come as counts — disease cases per day, website visits per hour, or pixel intensities in images. A common goal is to recover the smooth trend underlying these noisy counts. Trend filtering is a nonparametric method that fits flexible, piecewise polynomial curves which adapt automatically to abrupt changes in the signal without prespecifying where they occur. However, existing methods assume Gaussian noise, whereas count data follow Poisson-type models whose variability grows with the signal magnitude, making the effective noise heteroscedastic.

This dissertation develops scalable algorithms for trend filtering under Poisson loss and methodology for solving real-life applications. We propose two proximal algorithms that extend the estimator from simple time series to general graph structures. We apply this framework to epidemic surveillance, producing an R package (rtestim) that estimates time-varying reproduction numbers with principled, cross-validated tuning. We further identify and resolve a numerical instability in the linear system solvers that arise as inner subproblems of these algorithms, by recasting the system as a linear Gaussian state-space model, yielding a solver that is both stable and efficient. Finally, an ongoing work of ours shows that observation-dependent penalty weights can recover minimax optimal rates under the heteroscedastic noise inherent in exponential-family models.

Event Photo
Jiaping (Olivia) Liu

Extreme Value Theory: A Projection Estimator for the Angular Dependence Function

Extreme value theory provides a principal framework for modeling rare and extreme events, with applications in fields such as environmental science, finance, and engineering. Classical multivariate extreme value theory describes the limiting behaviour of normalised random vectors, particularly when the components are asymptotically dependent. However, many practical applications require a more flexible description of extremal dependence that can accommodate both asymptotic dependence and asymptotic independence. The Angular Dependence Function (ADF), introduced by Wadsworth and Tawn (2013), provides a flexible and interpretable way to characterize extremal dependence by describing how the rate of joint tail decay varies with the relative contribution of each component. Accurate estimation of the ADF is critical for understanding and modeling joint extremes.

While several estimators have been proposed for the ADF, they face significant limitations in finite samples, including high variability, irregular behavior across the domain, and violations of key theoretical constraints. The violations of the theoretical constraints are problematic and current approaches to addressing these issues are typically ad hoc, involving post hoc adjustments.

To resolve the issue of the violations, this work introduces a projection-based estimator for the ADF. Inspired by projection methods for the Pickands dependence function introduced in Fils-Villetard et al. (2008), the proposed method projects an initial non-parametric estimate onto a closed, convex set of admissible functions in L²([0, 1]). By construction, this estimator strictly enforces the theoretical upper and lower bounds. A simulation study across various Gaussian dependence levels demonstrates its ability to preserve validity and eliminate violations. This proposed method focuses on convex ADFs under positive quadrant dependence, reflecting the dependence structures most commonly observed in practice and other literature.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Rolling Extrapolation of Censored Survival Data and Its Applications to Lifetime Outcome Estimation

Estimation of lifetime outcomes is a fundamental problem in biostatistics, epidemiology, and health economic evaluation. In many cohort studies, however, follow-up durations are limited and survival data are heavily censored, making direct estimation of lifetime survival, life expectancy, and cumulative disease burden impossible. Conventional approaches typically rely on parametric survival models, but long-term extrapolations are often highly sensitive to model misspecification and may produce substantial bias.

In this talk, I will present a novel statistical framework, termed the Rolling Extrapolation Algorithm (REA), for extrapolating censored survival data beyond the observed follow-up period. The method incorporates external population information through a matched reference cohort and models relative survival between the study and reference populations. A key observation is that the logit transformation of relative survival often exhibits approximate linearity under broad classes of excess hazard models. Rather than performing a single long-term extrapolation, REA fits a restricted cubic spline model and iteratively predicts one step ahead, updating the fitted model in a rolling fashion until a lifetime horizon is reached.

Simulation studies demonstrate that REA substantially improves extrapolation accuracy compared with conventional one-shot spline and parametric approaches under a variety of hazard patterns. The resulting lifetime survival estimates can be combined with longitudinal quality-of-life, disability, healthcare expenditure, and productivity data to estimate life expectancy, years of life lost, disability-adjusted life years, and lifetime economic burden. Applications will be illustrated using nationwide cohort studies in Taiwan, including analyses of long-term PM2.5 exposure and healthy lifestyle factors.

The talk will focus on the statistical principles underlying REA, its theoretical motivation, empirical performance, practical implementation, and remaining methodological challenges in survival extrapolation and lifetime outcome estimation.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Jing-Shiang Hwang

A novel class of mixed Poisson distributions and wastewater-based epidemiology

Mixed Poisson families are widely used to model count data with overdispersion, zero inflation, or heavy tails in a variety of applications including finance, biology, and the physical sciences. The mixing distribution assigned to the Poisson rate is typically restricted to have nonnegative support. Surprisingly, this assumption is unnecessary. For example, the Hermite distribution is analogous to mixing a Poisson with an untruncated Gaussian and can be derived using generating functions so long as constraints on the natural parameter are satisfied. I will give a general characterization of this unusual class as well as several concrete examples, including an apparently novel generalization of the discrete stable family. I will also briefly present some applied work in wastewater-based epidemiology examining spatiotemporal variation of the pepper mild mottle virus biomarker. 

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Will Townes

Two MSc student presentations (Zhili Jiang & Zachary Lau)

Presentation 1

Time: 11:00am - 11:30am

Speaker: Zhili Jiang, UBC Statistics MSc student

Title: A Joint Model for Longitudinal and Survival Data with Nonlinear Trajectories and Interval-Censored Dropout, with Application to HIV Vaccine Studies

Abstract: Joint modeling of longitudinal biomarkers and time-to-event outcomes provides an important framework for understanding vaccine-induced immune responses and their relationship with clinical outcomes. In vaccine trials, dropout is often assumed to be non-informative and exactly observed, which may lead to biased inference when these assumptions are violated. In this study, we extend existing joint modeling approaches by incorporating a biologically motivated nonlinear mixed-effects model for longitudinal antibody trajectories and modeling dropout under both right- and interval-censored settings. The proposed framework provides a more realistic characterization of the association between immune dynamics and dropout through shared random effects. The method is applied to data from the VAX004 HIV-1 vaccine trial. The results suggest that dropout is associated with the underlying longitudinal antibody processes through shared random effects, supporting the presence of informative dropout under the proposed joint modeling framework. This association is consistently observed across Cox right-censored, Weibull right-censored, and Weibull interval-censored specifications. For the longitudinal component, the exponential-decay model provides a substantially better representation of antibody dynamics than linear and power-law alternatives. Simulation studies demonstrate reliable parameter estimation, although Hessian-based standard errors may underestimate uncertainty for parameters associated with the nonlinear component of the model. Overall, the proposed framework provides a flexible and biologically interpretable approach for joint modeling in vaccine studies, offering improved handling of realistic dropout mechanisms and the potential for extension to more complex longitudinal and survival settings. 

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Zachary Lau, UBC Statistics MSc student

Title: Scalable Gaussian Processes and Active Learning for Emulator Design in Solar Wind Simulation

Abstract: In this work, we discuss emulator design for solar wind simulators. Our work focuses on two areas. Firstly, we focus on scaling Gaussian Process regression to work well on simulator grids with millions of points. We accomplish this by extending existing work on Kronecker product covariance based algorithms to work efficiently with a dataset larger than working memory. Secondly, we implement and experiment with existing acquisition functions for active learning in the large data regime found in simulators. We find encouraging, though not definitive, results in favour of the Expected Predictive Information Gain acquisition function, particularly when it targets a prior concentrated in a particular part of the search space. To the best of our knowledge, this work is the first time that Gaussian Process Regression has been applied at this scale in Solar Wind modelling, the first time that these acquisition functions have been implemented at this scale for Gaussian Process models, and the first time that active learning has been applied to the problem of emulator design for the solar wind.

To join these seminars virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Manifold Sampling with Automatic Tuning

Many statistical and applied problems involve sampling from distributions constrained to curved lower-dimensional spaces, or manifolds. Standard MCMC methods are inapplicable in these settings because they do not naturally respect the constraint geometry, while existing manifold samplers can be highly sensitive to step-size tuning.

Our main contribution is an automatically tuned manifold sampler with a local step-size selection procedure that adapts to the geometry of the manifold. Under regularity conditions, we show that our method is invariant using the involutive MCMC framework. We further implement a contour-based sampling method with automatic tuning that achieves strong performance in terms of effective sample size per second while maintaining stable acceptance rates on several challenging target distributions. Empirical results show that automatic tuning can make manifold sampling more reliable and less sensitive to step-size choice for constrained and contour-based inference problems.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Junsong Tang

Introduction to the Computer-Based Testing Facility (CBTF)

The Computer-Based Testing Facility (CBTF) aims to solve a key problem in teaching and learning: helping instructors run digital assessments at scale securely and equitably. The easiest way to describe the CBTF is to imagine it as a network-filtered computer lab dedicated to running digital assessments in 50-minute increments throughout the day, invigilated by trained proctors. Students are typically given a multi-day window to write their tests at a time and location convenient for them. With network filtering, centralized invigilation, a distributed exam model, and flexibility for students and instructors, the mission of the CBTF is to spur pedagogical innovation at the university in a broad range of classes, programs, and departments. Though not the initial motivation, recently the CBTF has also been used to maintain exam integrity in the face of modern AI tools, particularly for computer-based exams where any element of programming is needed under controlled environments. This session will be useful for a range of people including faculty members teaching courses, administrators, IT staff, grad students as well as anyone with an interest in pedagogy and innovative teaching methods. We’ll also hear about the experience of several Statistics faculty members in using the CBTF in their courses as pilots in previous terms. There will be plenty of time for Q&A and a larger conversation around migration to computer-based testing and different learning technologies. The CBTF currently supports a variety of assessment options including Canvas, PrairieLearn, MTA and others. We will also discuss the advantages and disadvantages of different learning technologies in the CBTF from a pedagogical, logistical, and financial perspective.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Tags

UBC Statistics Department Colloquium: Statistical methods for single-cell and spatial data science

The UBC Statistics Department Colloquium Series features talks that are broad, accessible, and engaging - and open to everyone!

The third talk of our series will take place on Monday, June 8th where we will welcome Stephanie Hicks, Associate Professor of Biomedical Engineering and Biostatistics at Johns Hopkins University.

Date: Monday, June 8, 2026
Time: 3 - 4 PM
Location: ESB 5104/5106

Title: Statistical methods for single-cell and spatial data science

Abstract: Genomics is going through a data revolution where we can now profile gene expression at a single-cell or 2D spatial resolution. However, these data present unique challenges that have required the development of specialized statistical and computational methods and software infrastructure to successfully derive biological insights. Compared to bulk RNA-seq, there is an increased scale of the number of observations (or cells) that are measured and there is increased sparsity of the data, or fraction of observed zeros. Furthermore, as single-cell technologies mature, the increasing complexity and volume of data require fundamental changes in data access, management, and infrastructure alongside specialized methods to facilitate scalable analyses. I will discuss some challenges in the analysis of data and present some solutions that we have made towards addressing these challenges.

This colloquium series is sponsored in part by the Constance van Eeden Endowment.

Tags