|
arXiv:2411.03383v3 Announce Type: replace Abstract: How hard is it to estimate a discrete-time signal $(x_{1}, ..., x_{n}) \in \mathbb{C}^n$ satisfying an unknown linear recurrence relation of order $s$ and observed in i.i.d. complex Gaussian noise? The class of all such signals is parametric but extremely rich: it contains all exponential polynomials over $\mathbb{C}$ with total degree $s$, including harmonic oscillations with $s$ arbitra....
|
|
B-MASTER: Scalable Bayesian Multivariate Regression for Master Predictor Discovery in Colorectal Cancer Microbiome-Metabolite Profiles
-
arxiv.org
-
1 month ago
-
eng
arXiv:2412.05998v4 Announce Type: replace Abstract: Motivation: The gut microbiome shapes cancer therapy response through its influence on host metabolism. While prior studies examine pairwise associations between individual genera and metabolites, there is limited methodology for identifying microbial genera that systematically regulate the overall metabolome. Scalable statistical tools are needed to uncover such system-level 'master pred....
|
|
Highest Posterior Density Intervals of Unimodal Distributions As Analogues to Profile Likelihood Ratio Confidence Intervals
-
arxiv.org
-
1 month ago
-
eng
arXiv:2412.06528v5 Announce Type: replace Abstract: In Bayesian statistics, the highest posterior density (HPD) interval is often used to describe properties of a posterior distribution. As a method for estimating confidence intervals (CIs), the HPD has two main desirable properties. Firstly, it is the shortest interval to have a specified coverage probability. Secondly, every point inside the HPD interval has a density greater than every ..
|
|
Targeted Data Fusion for Region-Specific Survival Effects in the AMP HIV Prevention Trials
-
arxiv.org
-
1 month ago
-
eng
arXiv:2501.18798v3 Announce Type: replace Abstract: The Antibody Mediated Prevention (AMP) trials opened a new scientific frontier by showing that passively administered monoclonal broadly neutralizing antibodies (bnAbs) could prevent HIV-1 acquisition. Conducted across multiple geographic regions, including the United States, Brazil, Peru, Switzerland, and sub-Saharan Africa, the AMP trials revealed substantial regional heterogeneity in t....
|
|
A Unified Framework for Multiple-Try Metropolis: Construction and Empirical Benchmarks
-
arxiv.org
-
1 month ago
-
eng
arXiv:2503.11583v2 Announce Type: replace Abstract: The multiple-try Metropolis (MTM) algorithm uses a compound proposal with multiple candidate draws to improve local sampling efficiency. While several methodological works have continued to develop MTM and the multi-candidate mechanism that characterizes it, the literature lacks a unified comparison of these components. This paper presents a structured formulation of MTM within the involu....
|
|
arXiv:2504.06108v3 Announce Type: replace Abstract: Causal inference in connected populations is complicated by contagion and other real-world processes inducing dependence among outcomes. We address a gap in the literature on causal inference under contagion: while there is a growing body of work on estimating causal effects under contagion, little is known about how contagion impacts causal effects and inference. We provide insight into ....
|
|
Assessing Racial Disparities in Healthcare Expenditures via Mediator Distribution Shifts
-
arxiv.org
-
1 month ago
-
eng
arXiv:2504.21688v4 Announce Type: replace Abstract: Racial disparities in healthcare expenditures are well-documented, yet the underlying drivers remain complex. This study develops a framework to decompose such disparities through shifts in the distributions of mediating variables, rather than treating race itself as a manipulable exposure. We define disparities as differences in covariate-adjusted outcome distributions across racial grou....
|
|
arXiv:2505.19925v2 Announce Type: replace Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers. These can be casewise outliers, such as cases belonging to a different population, or cellwise outliers, which are deviating cells (entries) of the data matrix. Recently some robust covariance estimators have been developed that can handle both types of outliers, but their com....
|
|
A longitudinal Bayesian framework for estimating causal dose-response relationships
-
arxiv.org
-
1 month ago
-
eng
arXiv:2505.20893v4 Announce Type: replace Abstract: Existing causal methods for time-varying exposure and time-varying confounding focus on estimating the average causal effect of a time-varying binary treatment on an end-of-study outcome, offering limited tools for characterizing marginal causal dose-response relationships under continuous exposures. We propose a scalable, nonparametric Bayesian framework for estimating marginal longitudi....
|
|
Position: Stop Chasing the C-index when Evaluating Survival Analysis Models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.02075v3 Announce Type: replace Abstract: The current state of evaluation in survival analysis is plagued by the persistent use of evaluation metrics in ways that are misaligned with the stated modeling objective. In addition, many such evaluations are based on censoring assumptions that are left implicit or unjustified. This means that the reported performance can be misleading and may fail to answer the scientific or modeling q....
|
|
Identifiability in epidemic models with prior immunity and under-reporting
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.07825v2 Announce Type: replace Abstract: Identifiability is the property in mathematical modelling that determines if model parameters can be uniquely estimated from data. For infectious disease models, failure to ensure identifiability can lead to misleading parameter estimates and unreliable policy recommendations. We examine the identifiability of a modified SIR model that accounts for under-reporting and pre-existing immunit....
|
|
Consistent Infill Estimability of the Regression Slope Between Gaussian Random Fields Under Spatial Confounding
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.09267v3 Announce Type: replace Abstract: The problem of estimating the slope parameter in regression between two spatial processes under confounding by an unmeasured spatial process has received widespread attention in the recent statistical literature. Yet, a fundamental question remains unresolved: when is this slope consistently estimable under spatial confounding, with existing insights being largely empirical or estimator-s....
|
|
arXiv:2506.10677v3 Announce Type: replace Abstract: We study A/B testing, the standard protocol for measuring the performance gain of a new decision system relative to a baseline. Traditional A/B testing treats both systems as black boxes, ignoring potential similarities between them. In practice, however, new and baseline systems are rarely radically different and often share significant structure, which can be captured by their propensit....
|
|
The fundamental problem of risk prediction for individuals: health AI, uncertainty, and personalized medicine
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.17141v2 Announce Type: replace Abstract: Background and Objective: Clinical prediction models are commonly evaluated regarding performance for a population, although decisions are made for individuals. The classic view relates uncertainty in risk estimates for individuals to sample size (estimation uncertainty) while other sources are model uncertainty (variability in modeling choices) and applicability uncertainty (variability ....
|
|
Simultaneous estimation of the effective reproduction number and the time series of daily infections: Application to Covid-19
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.21027v3 Announce Type: replace Abstract: The time-varying effective reproduction number is an important parameter for communication and policy decisions during an epidemic. In this paper, we present new statistical methods for estimating the reproduction number based on the popular model of \citet{cori2013new} which defines the effective reproduction number based on self-exciting dynamics of new infections. Such a model is conce....
|
|
Hyperspherical Variational Autoencoders Using Efficient Spherical Cauchy Distribution
-
arxiv.org
-
1 month ago
-
eng
arXiv:2506.21278v3 Announce Type: replace Abstract: We propose spherical Cauchy (spCauchy) latent variables for variational autoencoders on hyperspherical latent spaces. The spCauchy family has heavy-tailed global behavior and admits an exact differentiable reparameterization by applying a M\"obius transformation to uniform samples on the sphere. We show that, in the high-concentration limit, spCauchy recovers the local tangent-space geome....
|
|
Covariance scanning for adaptively optimal change point detection in high-dimensional linear models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.02552v4 Announce Type: replace Abstract: This paper investigates the detection and estimation of a single change in high-dimensional linear models. We derive minimax lower bounds for the detection boundary and the estimation rate, which uncover a phase transition governed by the sparsity of the covariance-weighted differential parameter. This form of "inherent sparsity" captures a delicate interplay between the covariance struct....
|
|
Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.07339v2 Announce Type: replace Abstract: Decisions about managing patients on the heart transplant waitlist are currently made by committees of doctors who consider multiple factors, but the process remains largely ad-hoc. With the growing volume of longitudinal patient, donor, and organ data collected by the United Network for Organ Sharing (UNOS) since 2018, there is increasing interest in analytical approaches to support clin....
|
|
Exact conditional goodness-of-fit tests for the mixed membership stochastic block model
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.14464v2 Announce Type: replace Abstract: We propose exact conditional goodness-of-fit tests for directed mixed membership stochastic block models. Given dyad-level sender and receiver roles, the block-pair edge totals are sufficient for the block probability matrix; conditioning on these totals gives a nuisance-free uniform law on a finite fiber. This yields finite-sample randomization tests for residual sender and receiver hete..
|
|
Signal Detection under Composite Hypotheses with Identical Distributions for Signals and for Noises
-
arxiv.org
-
1 month ago
-
eng
arXiv:2507.21692v2 Announce Type: replace Abstract: In this paper, we consider the problem of detecting signals in multiple, sequentially observed data streams, where the distribution of each stream lies in one of two common composite spaces, depending on whether it is a signal or a noise. For this problem, we study a practical yet underexplored setting where it is a priori known that all signals have an identical distribution and so do al....
|
|
arXiv:2508.01973v4 Announce Type: replace Abstract: This article demonstrates how recent developments in the theory of empirical processes allow us to construct a new family of asymptotically distribution-free smooth tests. Their distribution-free property is preserved even when the parameters are estimated, model selection is performed, and the sample size is only moderately large. A computationally efficient alternative to the classical ..
|
|
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.03456v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimators with improved statistical properties, assuming that better estimators inherently yield superior policies. Although theoretically justified, this estimator-centric approach neglects a critical practical obstac....
|
|
Geometry-preserving and interpretable dimension reduction for compositional data
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.05563v2 Announce Type: replace Abstract: High-dimensional compositional data pose unique statistical challenges due to the simplex constraint and excess zeros. While dimension reduction is indispensable for analyzing such data, conventional approaches often rely on log-ratio transformations that compromise interpretability and distort the data through ad hoc zero replacements. To address these issues, we introduce a geometry-pre....
|
|
Adaptive clinical trial design with delayed treatment effects using elicited prior distributions
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.07602v2 Announce Type: replace Abstract: Clinical trials with time-to-event endpoints, such as overall survival (OS) or progression-free survival (PFS), are fundamental for evaluating new treatments, particularly in immuno-oncology. However, modern therapies, such as immunotherapies and targeted treatments, often exhibit delayed effects that challenge traditional trial designs. These delayed effects violate the proportional haza....
|
|
A Statistical Test for Comparing the Linkage and Admixture Model Based on Central Limit Theorems
-
arxiv.org
-
1 month ago
-
eng
arXiv:2509.12734v4 Announce Type: replace Abstract: In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in $K$ ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model (HMM) that extends the Admixture Model by incorporating....
|
|
arXiv:2509.23544v2 Announce Type: replace Abstract: Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-definite matrices. These outputs are naturally modeled as elements of general metric spaces, where classical regression techniques that rely on vector space structure no longer apply. We introduce E2M (End-to-End Metric regression), a deep learn....
|
|
arXiv:2510.05566v2 Announce Type: replace Abstract: Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real-world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading....
|
|
arXiv:2510.09288v2 Announce Type: replace Abstract: The vulnerability of machine learning models to adversarial attacks remains a critical societal security challenge. Traditional defenses, such as adversarial training, typically robustify models by minimizing a worst-case loss. These deterministic approaches do not account for uncertainty in the adversary's attack. While stochastic defenses placing a probability distribution on the advers....
|
|
arXiv:2510.15762v3 Announce Type: replace Abstract: The estimand framework is increasingly established to pose research questions in confirmatory clinical trials. In evidence synthesis, the uptake of estimands has been modest, and the PICO (Population, Intervention, Comparator, Outcome) framework is more often applied. While PICOs and estimands have overlapping elements, the estimand framework explicitly considers different strategies for ....
|
|
Generalized Guarantees for Variational Inference in the Presence of Even and Elliptical Symmetry
-
arxiv.org
-
1 month ago
-
eng
arXiv:2511.01064v3 Announce Type: replace Abstract: Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions. The best variational approximation is found by minimizing a divergence between distributions, $D(p||q)$, and several divergences have been proposed as objective functions for VI, with different choices leading to different approximations. We show that even when these ....
|
|
arXiv:2511.04873v2 Announce Type: replace Abstract: Prototype selection methods compress a training set, but the existing taxonomy of condensation, edition, hybrid, competence-based, optimization-based, and clustering-based families does not include methods that operate on the multi-scale topological structure of the data. This paper introduces two different persistence-based prototype selector variants, Topological Prototype Selector (TPS....
|
|
Sequential Bootstrap for Out-of-Bag Error Estimation: A 100-Seed Replication Study and Variance-Structure Analysis
-
arxiv.org
-
1 month ago
-
eng
arXiv:2511.18065v2 Announce Type: replace Abstract: Out-of-Bag (OOB) estimation is the standard internal diagnostic for bootstrap-aggregated tree ensembles. Under the classical multinomial bootstrap, the number of distinct training observations in each replicate, $U_b$, is itself random, but its contribution to OOB-based variability has rarely been isolated empirically. We use Sequential Bootstrap (SB) -- a resampling scheme that holds $U_....
|
|
arXiv:2601.11229v4 Announce Type: replace Abstract: Qualitative Comparative Analysis (QCA) requires researchers to choose calibration and dichotomization thresholds, and these choices can substantially affect truth tables, minimization, and resulting solution formulas. Despite this dependency, threshold sensitivity is often examined only in an ad hoc manner because repeated analyses are time-intensive and error-prone. We present ThSQCA, an....
|
|
Estimating conditional Mann-Whitney effects using pseudo-observation-based regression
-
arxiv.org
-
1 month ago
-
eng
arXiv:2601.15880v3 Announce Type: replace Abstract: The Mann-Whitney effect is an effect measure for the order of two sample-specific outcome variables. It has the interpretation of a probability and also a connection to the area under the ROC curve. In the literature it has been considered for both ordinal and right-censored time-to-event outcomes. For both cases, the present paper introduces a distribution-free regression model that rela....
|
|
arXiv:2601.21696v2 Announce Type: replace Abstract: Advances in data collection are producing growing volumes of temporal count observations, making adapted modeling increasingly necessary. In this work, we introduce a generative framework for independent component analysis of temporal count data, combining regime-adaptive dynamics with Poisson log-normal emissions. The model identifies disentangled components with regime-dependent contrib..
|
|
arXiv:2601.21959v2 Announce Type: replace Abstract: We develop a near-optimal testing procedure under the framework of Gaussian differential privacy for simple as well as one- and two-sided tests under monotone likelihood ratio conditions. Our mechanism is based on a private mean estimator with data-driven clamping bounds, whose population risk matches the private minimax rate up to logarithmic factors. Using this estimator, we construct p..
|
|
arXiv:2601.22784v2 Announce Type: replace Abstract: We introduce a rank-statistic approximation of $f$-divergences that avoids explicit density-ratio estimation by working directly with the distribution of ranks. For a resolution parameter $K$, we map the mismatch between two univariate distributions $\mu$ and $\nu$ to a rank histogram on $\{ 0, \ldots, K\}$ and measure its deviation from uniformity via a discrete $f$-divergence, yielding ....
|
|
arXiv:2601.22945v2 Announce Type: replace Abstract: We propose a novel framework for measuring privacy from a Bayesian game-theoretic perspective. This framework enables the creation of new, purpose-driven privacy definitions that are rigorously justified, while also allowing for the assessment of existing privacy guarantees through game theory. We show that pure and probabilistic differential privacy are special cases of our framework, an..
|
|
arXiv:2602.00878v2 Announce Type: replace Abstract: Slice sampling is a standard Monte Carlo technique for Dirichlet process (DP)-based models, widely used in posterior simulation. However, formal assessments of the scalability of posterior slice samplers have remained largely unexplored, primarily because the computational cost of a slice-sampling iteration is random and potentially unbounded. In this work, we obtain high-probability boun....
|
|
Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.03970v3 Announce Type: replace Abstract: We study the statistical behavior of reasoning probes in a stylized model of iterative computation inspired by neural algorithmic reasoning. The underlying computation is given by a looped Boolean circuit whose graph is a perfect $\nu$-ary tree ($\nu\ge 2$), with outputs recursively fed back as inputs across computation rounds. A probe observes a sampled subset of internal nodes and seeks....
|
|
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.03972v3 Announce Type: replace Abstract: The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixed-confidence setting (FC). For $K$-armed bandits with a unique best arm, the optimal sample complexities for both settings have been settled down, and they match up to logarithmic factors. This prompts an intere....
|
|
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.05395v2 Announce Type: replace Abstract: A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this paper we leverage Bayesian prior information to save on sampling costs, stopping once sufficient consistency is reached. Although the exact posterior is computationally intractable, we further introduce an efficie..
|
|
Deep networks learn to parse uniform-depth context-free languages from local statistics
-
arxiv.org
-
1 month ago
-
eng
arXiv:2602.06065v3 Announce Type: replace Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning. Studies of the internal representations of Large Language Models (LLMs) support their ability to parse text when predicting the next word, while representing semantic notions independently of surface form. Yet, which data statistics make the....
|
|
arXiv:2602.09651v2 Announce Type: replace Abstract: Diffusion models do not recover semantic structure uniformly over time. Instead, samples transition from semantic ambiguity to class commitment within a narrow regime. Recent theoretical work attributes this transition to dynamical instabilities along class-separating directions, but practical methods to detect and exploit these windows in trained models are still limited. We show that tr....
|