|
arXiv:2606.00913v1 Announce Type: new Abstract: Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge. After deploying bandits, a natural question is whether one can construct a confidence interval for its mean reward and assess whether it reliably outperforms a baseline policy. The total reward a....
|
|
Efficient Synthetic Network Generation via Latent Embedding Reconstruction
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.00934v1 Announce Type: new Abstract: Network data are ubiquitous across the social sciences, biology, and information systems. Generating realistic synthetic network data has broad applications from network simulation to scientific discovery. However, many existing black-box approaches for network generation tend to overfit observed data while overlooking characteristic network structure, and incur substantial computational over....
|
|
Design-based edge-level causal inference with machine learning assisted covariate adjustment
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.00965v1 Announce Type: new Abstract: We study design-based causal inference for edge-level outcomes in directed networks under dyadic interference. In this setting, outcomes are defined on directed edges and depend on the joint treatment assignments of pairs of units, inducing a complex dependence structure that invalidates standard estimation and inference procedures developed for node-level data. We construct Horvitz--Thompson....
|
|
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.00984v1 Announce Type: new Abstract: We study linear contextual bandits under rare parameter updates: the learner may incorporate reward feedback into its parameter estimate only at a small number of update times, while still observing contexts online and selecting actions sequentially. This viewpoint clarifies a practical distinction that is often blurred in the literature: many "strictly batched" methods additionally restrict ....
|
|
arXiv:2606.01002v1 Announce Type: new Abstract: Engression is a recently proposed and effective framework for conditional distribution learning. Its multi-step Reverse Markov extension further improves generative flexibility by decomposing complex conditional sampling into sequential reverse transitions. Despite their strong empirical performance, rigorous finite-sample statistical guarantees for these methods remain unavailable. In this p..
|
|
Semiparametric Efficiency of Residual Correlation Testing under Gaussian Additive Noise Models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01011v1 Announce Type: new Abstract: This paper studies conditional independence testing under the Gaussian additive noise model (GANM), where two variables are modeled as nonlinear functions of covariates with independent bivariate Gaussian regression errors. Under this framework, conditional independence can be characterized by the correlation coefficient of the regression errors, which motivates a test based on the Pearson co....
|
|
arXiv:2606.01090v1 Announce Type: new Abstract: Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds. On a controlled C_n-symmetric task, we report three findings. First, a wrong-group control with identical orbit size and matched compute is worse than no constraint (j....
|
|
arXiv:2606.01184v1 Announce Type: new Abstract: Many interventions alter the structure of an outcome distribution rather than its mean: they can split a population into disconnected regimes, create loops or holes, generate branches, or reorganize an outcome cloud while leaving the average response nearly unchanged. In such settings, mean-based causal estimands such as the average treatment effect may miss important structural effects. We ....
|
|
Markovianity-Based Conditioning Depth Diagnostics for Hidden Confounding in Observational Datasets
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01214v1 Announce Type: new Abstract: Reliable causal discovery in time series depends on whether the conditioning set adequately represents the system state. If relevant history or unobserved processes are omitted, residual dependence can appear as direct causal links. We study this failure mode on promnient constraint-based causal discovery methods through a simple premise: how much does the inferred graph change as conditionin....
|
|
Functional Clustering of Survival Data via Smoothed Log-Hazard Trajectories: A Risk-Dynamics Perspective
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01239v1 Announce Type: new Abstract: This paper investigates clustering in survival data by shifting the analytical focus from cumulative survival probabilities to instantaneous risk, as characterized by the hazard function. We model smoothed log-hazard trajectories as functional objects that capture the temporal evolution of risk and propose a clustering framework based on Functional Principal Component Analysis applied to B-sp....
|
|
Efficient Approximation for Encoder--Decoder Neural Operators via Variation Spaces
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01244v1 Announce Type: new Abstract: We study operator learning using encoder--decoder neural networks. Inspired by the function-space theory of neural networks, we introduce a variation space as an infinite-dimensional structural class for nonlinear operators. This space is defined through vector-valued measures directly on the input and output spaces. For operators in this space, we establish approximation bounds for encoder--....
|
|
Distribution-free changepoint localization after sequential change detection
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01256v1 Announce Type: new Abstract: This paper introduces a distribution-free framework for constructing post-detection confidence sets for changepoints after stopping a sequential change detection procedure. It is well known that conformal test martingales can be used to sequentially detect changes in distribution, but by themselves provide no inference for the time at which a proclaimed change occurred. Past work on post-dete....
|
|
arXiv:2606.01257v1 Announce Type: new Abstract: Gradient-based algorithms are central to modern statistical estimation, yet their statistical analysis is often restricted to fixed-time behavior, such as convergence to a population target or fluctuations at a prescribed iteration. In many applications, however, uncertainty quantification is needed along the entire optimization path, especially when the stopping time is data-dependent or div....
|
|
Scale-Free Priors and Survival Dynamics: A Bayesian Framework for Conflict Duration
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01328v1 Announce Type: new Abstract: We have developed a fully Bayesian survival-analysis framework that reformulates inference about system lifetimes in terms of hazard and survival functions, and extends this representation to interacting actors. Starting from J.~Richard Gott's Copernican principle, we express the scale-free prior as a baseline hazard $\lambda(t)=1/t$, thereby linking a static prior over lifetimes to the dynam....
|
|
FlowSDR: Sufficient Dimension Reduction via Conditional Normalizing Flows
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01346v1 Announce Type: new Abstract: Sufficient dimension reduction (SDR) seeks a low-dimensional linear projection of predictors that preserves the conditional distribution of the response. Existing methods target this conditional distribution indirectly, via inverse moments, local forward regression, or neural ensemble regression. We propose FlowSDR, a likelihood-based framework that jointly learns the projection and the condi....
|
|
On the Uncertainty Quantification Ability of Tabular Foundation Models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01427v1 Announce Type: new Abstract: Foundation models (FMs) have achieved substantial success in generalizing across tasks without problemspecific training or fine-tuning. However, many critical applications in mechanics and computational science require not only accurate predictions but also reliable uncertainty quantification (UQ). Herein we investigate the UQ capabilities of tabular FMs in regression tasks through a comprehe....
|
|
Quantifying Evidential Rigor in Meta-Analytic Corpora: A Simulation-Characterized, Bias-Robust Bayesian Workflow with a Nutrition Case Study
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01428v1 Announce Type: new Abstract: Conventional meta-analysis summarizes evidence through pooled estimates, intervals, and p-values, but these outputs do not directly measure evidence for an effect, evidence for no effect, or the degree to which conclusions depend on publication selection or small-study effects. We introduce a corpus-scale Bayesian evidential-audit workflow for meta-analytic corpora. The workflow reconstructs ....
|
|
Comb Test: Histogram Uniformity Testing Based on Discrete Total Variation
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01465v1 Announce Type: new Abstract: Histogram uniformity testing is a common statistical task usually performed using Pearson's chi-square test. This paper proposes a new test based on the discrete total variation that is easy to compute and, for comb-like (alternating) deviations, achieves up to 67% higher statistical power than Pearson's chi-square test, making it a complement to standard tests. The exact null distribution is..
|
|
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01468v1 Announce Type: new Abstract: Due to their explicit priors and ability to model uncertainty, Bayesian methods have played a major role in dynamical latent variable modeling of single-cell neural recordings. However, modern-sized datasets have made overparameterized deep networks the preferred methods of choice due to their predictive power and favorable computational scaling. While many posterior approximations exist, all....
|
|
Voronoi-Elitism Genetic Algorithm: A Generic Derivative-Free Routine With Theory and Implementation for Statistical Optimization
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01474v1 Announce Type: new Abstract: In this paper, we propose a generic optimization approach for challenging objective functions that finds applications in various statistical problems. We focus on objective functions with two parameter blocks of one amenable to analytic optimization, and another that is irregular or computationally expensive. To address this setting, we propose the Voronoi-Elitism Genetic Algorithm (VEGA), a ....
|
|
arXiv:2606.01489v1 Announce Type: new Abstract: Regression models and Vector Autoregressive Models (VARs) play crucial roles in econometrics by allowing the analysis of multiple variables simultaneously. Despite their utility, these models face challenges like underfitting and overfitting, especially when determining the optimal model specification, which can lead to significant computational costs. To address these challenges, econometric....
|
|
A flexible and robust approach to univariate Gaussian splitting using parameterised Gaussian mixtures
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01530v1 Announce Type: new Abstract: We consider approximation of a Gaussian distribution with a mixture of homoscedastic Gaussians of smaller variance. The solution is obtained by minimising the $L^2$ norm between the original Gaussian and the mixture, which is parameterised to reduce the complexity of the optimisation problem. The developed technique is straightforward, sufficiently robust and yields Gaussian Mixtures that rap..
|
|
Scalable Counterfactual Risk Estimation for Rare Events in Longitudinal Data
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01539v1 Announce Type: new Abstract: Estimating the causal effect of time-varying treatments on survival outcomes in large observational studies is computationally demanding, particularly when outcomes are rare. While g-formula-based methods such as the iterative conditional expectation (ICE) estimator provide a principled framework for longitudinal causal inference, they become computationally expensive, especially when bootstr....
|
|
Structural Change Detection in High-Dimensional Transformed Factor Models via Canonical Correlation Analysis
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01553v1 Announce Type: new Abstract: This paper develops a canonical-correlation-based method for detecting structural changes in high-dimensional transformed factor models. The proposed approach exploits the low-rank canonical-correlation structure induced by dynamically dependent common factors, while serially uncorrelated idiosyncratic components correspond to a noise subspace with zero canonical correlations. We construct an....
|
|
arXiv:2606.01554v1 Announce Type: new Abstract: This short note proposes a polynomial-time algorithm for near-optimal Euclidean estimation of a signal constrained to lie in the unit ball of a symmetric norm, where the symmetry is with respect to a known basis and the norm is accessible through an evaluation oracle. We further extend the method to a random-design, moderate-dimensional linear regression setting, where the regression paramete..
|
|
arXiv:2606.01645v1 Announce Type: new Abstract: Diffusion models have emerged as a leading framework for deep generative modeling. While the standard Gaussian formulation is theoretically convenient, its suitability for heavy-tailed datasets remains unclear. To address this, heavy-tailed diffusion models (HTDMs) extend the standard formulation by replacing the Gaussian distribution with a Student's t-distribution, thereby improving tail fi..
|
|
Beyond principal ignorability: Nonparametric sensitivity bounds for principal stratification
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01669v1 Announce Type: new Abstract: Principal stratification is an effective framework addressing intermediate variables in causal inference. However, point identification of the principal causal effects (PCEs) often requires the untestable principal ignorability (PI) assumption. This article develops a nonparametric sensitivity analysis framework for evaluating PI violations. We introduce a margin-free bounding factor paramete....
|
|
Higher-Order Efficient Estimators: A Review and Simulation-Based Benchmark Study
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01674v1 Announce Type: new Abstract: Higher-order efficient estimators extend standard first-order semiparametric estimators by replacing second-order residuals with third- or higher-order terms, potentially enabling asymptotic efficiency under slower nuisance function convergence rates and improving finite-sample performance. Existing methods achieve higher-order expansions through structurally different approximation strategie....
|
|
Mapping the Storm: Geospatial Impacts of Severe Weather on LEO Network Performance
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01724v1 Announce Type: new Abstract: LEO satellite constellations, led by deployments such as Starlink, are playing an increasingly pivotal role in enabling global broadband connectivity. However, the reliability and performance of these space-based networks are highly sensitive to environmental dynamics, particularly localized weather phenomena that exhibit strong spatio-temporal variability. In this study, we present a contine....
|
|
LoopPerm-CPD: A Robust Loop Permutation Framework for Automatic Multiple Change-Point Detection in Longitudinal Data
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01796v1 Announce Type: new Abstract: Human viral challenge studies, in which participants are deliberately inoculated with influenza strains such as H1N1 or H3N2 and monitored through longitudinal transcriptomic profiling before and after inoculation, are critical for characterizing dynamic biological immune responses to viral infection. A key analytical goal in such settings is to detect critical transition times, or change poi....
|
|
A Uniform Improvement of the Benjamini-Hochberg Procedure using e-Closure
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01854v1 Announce Type: new Abstract: This paper presents closed BH, a uniform improvement of the False Discovery Rate controlling method of Benjamini and Hochberg (BH). Closed BH is valid under the same assumption of Positive Regression Dependency on a Subset (PRDS) as BH. As a uniform improvement, closed BH never rejects fewer hypotheses than BH, but it may reject quite a few more. An increase in power is observed especially wh..
|
|
Spatial Capture-Recapture With Penalized Regression Splines to Flexibly Model Wildlife Density and Distribution
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.01932v1 Announce Type: new Abstract: Spatial capture-recapture models are routinely used to estimate the abundance and distribution of wild animal populations and involve a latent spatial point process of animal activity centres that describes the spatial distribution of individuals. While traditional spatial capture-recapture models use a Poisson process, the assumption of conditional independence between points is often violat....
|
|
arXiv:2606.01960v1 Announce Type: new Abstract: We consider the problem of detecting a Return to Baseline (RtB) in high-frequency monitoring data preceding and following an intervention, where the aim is to identify the time at which the data-generating distribution realigns with its pre-intervention distribution. We propose a sequential, distribution-free testing procedure that does not rely on specifying a null model and provides anytime....
|
|
arXiv:2606.01990v1 Announce Type: new Abstract: The Admixture Model describes genetic marker data by representing each individual's genome as a mixture of contributions from $K$ ancestral populations, with the individual admixture vector summarizing the corresponding ancestry proportions. In population and forensic genetics, a key question is whether an individual's genome supports a predominantly single-ancestry interpretation or whether ....
|
|
Provable Data Scaling Law for Meta Learning via Complexity Minimization
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02008v1 Announce Type: new Abstract: Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases. However, existing theoretical frameworks for pre-training do not fully explain this phenomenon. In this paper, we introduce complexity minimization, a novel meta-representation learning framewo....
|
|
PliableBVS: A flexible Bayesian variable selection method for modeling interactions with mandatory modifying variables
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02017v1 Announce Type: new Abstract: High-dimensional interaction models are useful for studying, for example, how a large set of variables of interest, such as gene expression or other omics features, interact with a smaller set of modifying variables, such as clinical covariates. In this context, the pliable lasso has recently been proposed as an efficient method for screening large numbers of potential interaction terms under....
|
|
Convex Distance Operator Transport: A Convex and Geometry-Preserving Formulation
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02047v1 Announce Type: new Abstract: We introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence and intrinsic geometric structure. Specifically, CDOT employs an operator-based regularization that aligns aggregated distance structures by introducing distance and conditional expectation oper....
|
|
ICCDesign: An R Package for the Design and Analysis of ICC-Based Reliability Studies with Continuous Responses
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02059v1 Announce Type: new Abstract: The intraclass correlation coefficient (ICC) is among the most widely used statistics in reliability research, playing a central role in medical measurement, psychological assessment, and behavioral science. However, practical application of ICC faces two major obstacles. First, ICC can be organized into multiple forms under the McGraw and Wong (1996) framework -- including six widely reporte....
|
|
Evaluating the role of correlation among markers in prediction models
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02062v1 Announce Type: new Abstract: Different methods have been employed to estimate models maximizing the area under the receiver operating characteristic curve (ROC-AUC). Once a model is developed, integrating novel biomarkers may improve its diagnostic ability. However, the discrimination improvement from adding a new biomarker is not always evident, even if the marker itself has good discriminatory power. The sign and magni....
|
|
arXiv:2606.02065v1 Announce Type: new Abstract: While it is well-known how to compute the cells of a Laguerre tessellation for a given set of weighted generator points, it is not obvious how to invert a Laguerre tessellation. That is, given that one observes a Laguerre tessellation, how can one retrieve the weighted generators corresponding to the observed cells. In this paper, we consider inversion of a class of random Laguerre tessellati....
|
|
Modelling multi-cancer screening data to infer on natural history of disease: when can valid, identifiable and precise inference be obtained?
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02076v1 Announce Type: new Abstract: Background: Multistate models (MSMs) applied to screening data can characterise the natural history of cancer and predict "stage-shifts" from screening. However, inferring parameters like mean sojourn time (MST) is challenging as disease onset is inherently unobserved in these data. This is even more challenging when characterising heterogeneity between cancer types in multicancer early detec....
|
|
It does what it says on the tin: safe synthetic data from coarsened margins
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02101v1 Announce Type: new Abstract: This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to other methods currently available. The first is transparency; unlike other methods, the person in receipt of the SD will know which of the relationships between variables in the original data will be approximately maintained in the SD. The second is a guarantee that th....
|
|
arXiv:2606.02115v1 Announce Type: new Abstract: Parameter estimation in stochastic differential equations is a classical statistical problem of much importance in many scientific fields. Recent work of Tapia Costa et al. (2026) introduced a novel technique for estimating the drift when the diffusion parameter is known, using discrete samples from multiple trajectories. Their method treats drift estimation as a denoising problem, and levera....
|
|
ProbRes: Volatility Learning for Probabilistic Time-Series Forecasting
-
arxiv.org
-
1 month ago
-
eng
arXiv:2606.02117v1 Announce Type: new Abstract: Probabilistic time series forecasting has attracted increasing attention in financial applications due to the need to quantify risk and uncertainty in future observations. We propose ProbRes, a post-hoc probabilistic calibration method that explicitly learns and incorporates volatility dynamics into probabilistic forecasting, enabling effective handling of heteroskedastic data. During trainin....
|