|
Identifiability in epidemic models with prior immunity and under-reporting
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.07825v2 Announce Type: replace Abstract: Identifiability is the property in mathematical modelling that determines if model parameters can be uniquely estimated from data. For infectious disease models, failure to ensure identifiability can lead to misleading parameter estimates and unreliable policy recommendations. We examine the identifiability of a modified SIR model that accounts for under-reporting and pre-existing immunit....
|
|
Consistent Infill Estimability of the Regression Slope Between Gaussian Random Fields Under Spatial Confounding
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.09267v3 Announce Type: replace Abstract: The problem of estimating the slope parameter in regression between two spatial processes under confounding by an unmeasured spatial process has received widespread attention in the recent statistical literature. Yet, a fundamental question remains unresolved: when is this slope consistently estimable under spatial confounding, with existing insights being largely empirical or estimator-s....
|
|
arXiv:2506.10677v3 Announce Type: replace Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline. Traditional A/B testing treats both systems as black boxes, ignoring potential similarities between them. In practice, however, new and baseline systems are rarely radically different and often share significant structure, which can be captured by their propensit....
|
|
The fundamental problem of risk prediction for individuals: health AI, uncertainty, and personalized medicine
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.17141v2 Announce Type: replace Abstract: Background and Objective: Clinical prediction models are commonly evaluated regarding performance for a population, although decisions are made for individuals. The classic view relates uncertainty in risk estimates for individuals to sample size (estimation uncertainty) while other sources are model uncertainty (variability in modeling choices) and applicability uncertainty (variability ....
|
|
Simultaneous estimation of the effective reproduction number and the time series of daily infections: Application to Covid-19
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.21027v3 Announce Type: replace Abstract: The time-varying effective reproduction number is an important parameter for communication and policy decisions during an epidemic. In this paper, we present new statistical methods for estimating the reproduction number based on the popular model of \citet{cori2013new} which defines the effective reproduction number based on self-exciting dynamics of new infections. Such a model is conce....
|
|
Hyperspherical Variational Autoencoders Using Efficient Spherical Cauchy Distribution
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.21278v3 Announce Type: replace Abstract: We propose spherical Cauchy (spCauchy) latent variables for variational autoencoders on hyperspherical latent spaces. The spCauchy family has heavy-tailed global behavior and admits an exact differentiable reparameterization by applying a M\"obius transformation to uniform samples on the sphere. We show that, in the high-concentration limit, spCauchy recovers the local tangent-space geome....
|
|
Covariance scanning for adaptively optimal change point detection in high-dimensional linear models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.02552v4 Announce Type: replace Abstract: This paper investigates the detection and estimation of a single change in high-dimensional linear models. We derive minimax lower bounds for the detection boundary and the estimation rate, which uncover a phase transition governed by the sparsity of the covariance-weighted differential parameter. This form of "inherent sparsity" captures a delicate interplay between the covariance struct....
|
|
Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.07339v2 Announce Type: replace Abstract: Decisions about managing patients on the heart transplant waitlist are currently made by committees of doctors who consider multiple factors, but the process remains largely ad-hoc. With the growing volume of longitudinal patient, donor, and organ data collected by the United Network for Organ Sharing (UNOS) since 2018, there is increasing interest in analytical approaches to support clin....
|
|
Exact conditional goodness-of-fit tests for the mixed membership stochastic block model
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.14464v2 Announce Type: replace Abstract: We propose exact conditional goodness-of-fit tests for directed mixed membership stochastic block models. Given dyad-level sender and receiver roles, the block-pair edge totals are sufficient for the block probability matrix; conditioning on these totals gives a nuisance-free uniform law on a finite fiber. This yields finite-sample randomization tests for residual sender and receiver hete..
|
|
Signal Detection under Composite Hypotheses with Identical Distributions for Signals and for Noises
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.21692v2 Announce Type: replace Abstract: In this paper, we consider the problem of detecting signals in multiple, sequentially observed data streams, where the distribution of each stream lies in one of two common composite spaces, depending on whether it is a signal or a noise. For this problem, we study a practical yet underexplored setting where it is a priori known that all signals have an identical distribution and so do al....
|
|
arXiv:2508.01973v4 Announce Type: replace Abstract: This article demonstrates how recent developments in the theory of empirical processes allow us to construct a new family of asymptotically distribution-free smooth tests. Their distribution-free property is preserved even when the parameters are estimated, model selection is performed, and the sample size is only moderately large. A computationally efficient alternative to the classical ..
|
|
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.03456v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimators with improved statistical properties, assuming that better estimators inherently yield superior policies. Although theoretically justified, this estimator-centric approach neglects a critical practical obstac....
|
|
Geometry-preserving and interpretable dimension reduction for compositional data
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.05563v2 Announce Type: replace Abstract: High-dimensional compositional data pose unique statistical challenges due to the simplex constraint and excess zeros. While dimension reduction is indispensable for analyzing such data, conventional approaches often rely on log-ratio transformations that compromise interpretability and distort the data through ad hoc zero replacements. To address these issues, we introduce a geometry-pre....
|
|
Adaptive clinical trial design with delayed treatment effects using elicited prior distributions
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.07602v2 Announce Type: replace Abstract: Clinical trials with time-to-event endpoints, such as overall survival (OS) or progression-free survival (PFS), are fundamental for evaluating new treatments, particularly in immuno-oncology. However, modern therapies, such as immunotherapies and targeted treatments, often exhibit delayed effects that challenge traditional trial designs. These delayed effects violate the proportional haza....
|
|
A Statistical Test for Comparing the Linkage and Admixture Model Based on Central Limit Theorems
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.12734v4 Announce Type: replace Abstract: In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in $K$ ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model (HMM) that extends the Admixture Model by incorporating....
|
|
arXiv:2509.23544v2 Announce Type: replace Abstract: Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-definite matrices. These outputs are naturally modeled as elements of general metric spaces, where classical regression techniques that rely on vector space structure no longer apply. We introduce E2M (End-to-End Metric regression), a deep learn....
|
|
arXiv:2510.05566v2 Announce Type: replace Abstract: Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real-world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading....
|
|
arXiv:2510.09288v2 Announce Type: replace Abstract: The vulnerability of machine learning models to adversarial attacks remains a critical societal security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-case loss. These deterministic approaches do not account for uncertainty in the adversary's attack. While stochastic defenses placing a probability distribution on the advers....
|
|
arXiv:2510.15762v3 Announce Type: replace Abstract: The estimand framework is increasingly established to pose research questions in confirmatory clinical trials. In evidence synthesis, the uptake of estimands has been modest, and the PICO (Population, Intervention, Comparator, Outcome) framework is more often applied. While PICOs and estimands have overlapping elements, the estimand framework explicitly considers different strategies for ....
|
|
Generalized Guarantees for Variational Inference in the Presence of Even and Elliptical Symmetry
-
arxiv.org
-
1 month ago
-
eng
arXiv:2511.01064v3 Announce Type: replace Abstract: Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions. The best variational approximation is found by minimizing a divergence between distributions, $D(p||q)$, and several divergences have been proposed as objective functions for VI, with different choices leading to different approximations. We show that even when these ....
|
|
arXiv:2511.04873v2 Announce Type: replace Abstract: Prototype selection methods compress a training set, but the existing taxonomy of condensation, edition, hybrid, competence-based, optimization-based, and clustering-based families does not include methods that operate on the multi-scale topological structure of the data. This paper introduces two different persistence-based prototype selector variants, Topological Prototype Selector (TPS....
|
|
Sequential Bootstrap for Out-of-Bag Error Estimation: A 100-Seed Replication Study and Variance-Structure Analysis
-
arxiv.org
-
1 month ago
-
eng
arXiv:2511.18065v2 Announce Type: replace Abstract: Out-of-Bag (OOB) estimation is the standard internal diagnostic for bootstrap-aggregated tree ensembles. Under the classical multinomial bootstrap, the number of distinct training observations in each replicate, $U_b$, is itself random, but its contribution to OOB-based variability has rarely been isolated empirically. We use Sequential Bootstrap (SB) -- a resampling scheme that holds $U_....
|
|
arXiv:2601.11229v4 Announce Type: replace Abstract: Qualitative Comparative Analysis (QCA) requires researchers to choose calibration and dichotomization thresholds, and these choices can substantially affect truth tables, minimization, and resulting solution formulas. Despite this dependency, threshold sensitivity is often examined only in an ad hoc manner because repeated analyses are time-intensive and error-prone. We present ThSQCA, an....
|
|
Estimating conditional Mann-Whitney effects using pseudo-observation-based regression
-
arxiv.org
-
1 month ago
-
eng
arXiv:2601.15880v3 Announce Type: replace Abstract: The Mann-Whitney effect is an effect measure for the order of two sample-specific outcome variables. It has the interpretation of a probability and also a connection to the area under the ROC curve. In the literature it has been considered for both ordinal and right-censored time-to-event outcomes. For both cases, the present paper introduces a distribution-free regression model that rela....
|
|
arXiv:2601.21696v2 Announce Type: replace Abstract: Advances in data collection are producing growing volumes of temporal count observations, making adapted modeling increasingly necessary. In this work, we introduce a generative framework for independent component analysis of temporal count data, combining regime-adaptive dynamics with Poisson log-normal emissions. The model identifies disentangled components with regime-dependent contrib..
|
|
arXiv:2601.21959v2 Announce Type: replace Abstract: We develop a near-optimal testing procedure under the framework of Gaussian differential privacy for simple as well as one- and two-sided tests under monotone likelihood ratio conditions. Our mechanism is based on a private mean estimator with data-driven clamping bounds, whose population risk matches the private minimax rate up to logarithmic factors. Using this estimator, we construct p..
|
|
arXiv:2601.22784v2 Announce Type: replace Abstract: We introduce a rank-statistic approximation of $f$-divergences that avoids explicit density-ratio estimation by working directly with the distribution of ranks. For a resolution parameter $K$, we map the mismatch between two univariate distributions $\mu$ and $\nu$ to a rank histogram on $\{ 0, \ldots, K\}$ and measure its deviation from uniformity via a discrete $f$-divergence, yielding ....
|
|
arXiv:2601.22945v2 Announce Type: replace Abstract: We propose a novel framework for measuring privacy from a Bayesian game-theoretic perspective. This framework enables the creation of new, purpose-driven privacy definitions that are rigorously justified, while also allowing for the assessment of existing privacy guarantees through game theory. We show that pure and probabilistic differential privacy are special cases of our framework, an..
|
|
arXiv:2602.00878v2 Announce Type: replace Abstract: Slice sampling is a standard Monte Carlo technique for Dirichlet process (DP)-based models, widely used in posterior simulation. However, formal assessments of the scalability of posterior slice samplers have remained largely unexplored, primarily because the computational cost of a slice-sampling iteration is random and potentially unbounded. In this work, we obtain high-probability boun....
|
|
Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.03970v3 Announce Type: replace Abstract: We study the statistical behavior of reasoning probes in a stylized model of iterative computation inspired by neural algorithmic reasoning. The underlying computation is given by a looped Boolean circuit whose graph is a perfect $\nu$-ary tree ($\nu\ge 2$), with outputs recursively fed back as inputs across computation rounds. A probe observes a sampled subset of internal nodes and seeks....
|
|
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.03972v3 Announce Type: replace Abstract: The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixed-confidence setting (FC). For $K$-armed bandits with a unique best arm, the optimal sample complexities for both settings have been settled down, and they match up to logarithmic factors. This prompts an intere....
|
|
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.05395v2 Announce Type: replace Abstract: A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this paper we leverage Bayesian prior information to save on sampling costs, stopping once sufficient consistency is reached. Although the exact posterior is computationally intractable, we further introduce an efficie..
|
|
Deep networks learn to parse uniform-depth context-free languages from local statistics
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.06065v3 Announce Type: replace Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make the....
|
|
arXiv:2602.09651v2 Announce Type: replace Abstract: Diffusion models do not recover semantic structure uniformly over time. Instead, samples transition from semantic ambiguity to class commitment within a narrow regime. Recent theoretical work attributes this transition to dynamical instabilities along class-separating directions, but practical methods to detect and exploit these windows in trained models are still limited. We show that tr....
|
|
How Accurately Can a Gaussian Approximate Stochastic Approximation Iterates?
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.13906v2 Announce Type: replace Abstract: Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise. The focus of this paper is studying the distribution of SA iterates in finite time. In general, it is not possible to characterize the exact distribution, and therefore our goal is to find an approximation which can yield useful tail bounds. Inspired by the rich literature on the asymptotic n....
|
|
arXiv:2602.16794v2 Announce Type: replace Abstract: Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, w....
|
|
arXiv:2602.22768v2 Announce Type: replace Abstract: Multi-armed bandit (MAB) processes constitute a foundational subclass of reinforcement learning problems and represent a central topic in statistical decision theory. Yet, conducting valid sequential testing under adaptive allocation remains challenging due to the lack of asymptotic theory under non-i.i.d. reward sequences and sublinear sample sizes for some arms. To address this open cha....
|
|
Asymptotic theory for multiple samples with flexible random membership
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.24219v2 Announce Type: replace Abstract: A statistic can be a function of multiple samples. There is little existing work on asymptotic theory for such statistics when group membership is random. We propose a flexible framework that can handle both deterministic and random membership. We prove some asymptotic properties and apply the framework to the stratified sampling context.
|
|
arXiv:2603.07563v2 Announce Type: replace Abstract: In this paper, we address a fundamental limitation of the classical Wasserstein barycenter -- its sensitivity to outliers. To overcome these issues, we propose the robust Wasserstein barycenter (RWB) based on a recent concept of the robust optimal transport. Theoretical guarantees, including existence and consistency, are established for the proposed RWB. Through extensive numerical exper..
|
|
A Bayesian adaptive enrichment design using aggregate historical data to inform individualized treatment recommendations
-
arxiv.org
-
1 month ago
-
eng
arXiv:2603.09919v2 Announce Type: replace Abstract: Adaptive enrichment trials aim to identify and recruit participants most likely to benefit from treatment based on evolving biomarker evidence, with the goal of informing individualized treatment recommendations. Bayesian methods are well suited to these designs because they allow external information to be incorporated in a principled manner. In practice, prior studies often provide only....
|
|
Preconditioned One-Step Generative Modeling for Bayesian Inverse Problems in Function Spaces
-
arxiv.org
-
1 month ago
-
eng
arXiv:2603.14798v2 Announce Type: replace Abstract: We propose a machine-learning algorithm for Bayesian inverse problems in the function-space regime. Based on one-step generative transport, the method learns an amortized neural operator whose pushforward of a Gaussian source approximates the posterior distribution conditioned on each new observation. We show that white-noise sources are incompatible with the function-space limit, and the....
|
|
arXiv:2603.22215v2 Announce Type: replace Abstract: Joint modeling of multiview graphs with a common set of nodes between views and auxiliary predictors is an essential, yet less explored, area in statistical methodology. Traditional approaches often treat graphs in different views as independent or fail to adequately incorporate predictors, potentially missing complex dependencies within and across graph views and leading to reduced infer....
|
|
Tackling the 6/49 Lottery and Debunking Common Myths with Probabilistic Methods and Combinatorial Designs
-
arxiv.org
-
1 month ago
-
eng
arXiv:2603.24170v3 Announce Type: replace Abstract: At the end, the house always wins! This simple truth holds for all public games of chance. Nevertheless, since lotteries have existed, people have tried everything to give luck a helping hand. This article compares objective scientific approaches to tackle the 6/49 lottery: probabilistic methods and combinatorial designs. The mathematical models developed herein can be modified and applie..
|
|
arXiv:2605.00696v2 Announce Type: replace Abstract: We study adaptive querying for learning user-dependent quantities of interest, such as responses to held-out items and psychometric indicators, within tight query budgets. Classical Bayesian design and computerized adaptive testing typically rely on restrictive parametric assumptions or expensive posterior approximations, limiting their use in heterogeneous, high-dimensional, and cold-sta....
|