How Robust are Robustness Checks?
Solo-authored · Working paper
Abstract
Robustness checks are routine in empirical work, but there is no standard statistical procedure to formally measure what one can learn from them. I propose a ``robustness radius" measure to quantify the amount by which the robustness checks estimands differ from the main specification estimand. I do so by framing robustness checks as explicitly biased regressions and applying a test from the moment inequalities literature. The resulting robustness radius is a one-sided confidence bound that accounts for sampling uncertainty and correlation across regressions. I also propose relative and standardized robustness radiuses, which provide alternative normalizations that facilitate interpretation and comparison across settings. An application shows that, although assessing overall robustness remains context-specific, the robustness radius guides those judgments and improves transparency.
Functional Classification of Bitcoin Addresses
with Manuel Febrero-Bande, Wenceslao González-Manteiga, and Yuri Saporito (2023)
Computational Statistics & Data Analysis, 107687
Abstract
A classification model for predicting the main activity of bitcoin addresses based on their balances is proposed.
Since the balances are functions of time, methods from functional data analysis are applied; more specifically,
the features of the proposed classification model are the functional principal components of the data.
Classifying bitcoin addresses is a relevant problem for two main reasons: to understand the composition of the
bitcoin market, and to identify addresses used for illicit activities. Although other bitcoin classifiers have
been proposed, they focus primarily on network analysis rather than curve behavior. The proposed approach, on
the other hand, does not require any network information for prediction. Furthermore, functional features have
the advantage of being straightforward to build, unlike expert-built features. Results show improvement when
combining functional features with scalar features, and similar accuracy for the models using those features
separately, which points to the functional model being a good alternative when domain-specific knowledge is
not available.
Externally Valid Selection of Experimental Sites via the k-Median Problem
with José Luis Montiel Olea, Chen Qiu, Jörg Stoye, and Yiwei Sun (2025)
R&R at JPE:Micro
Extended Abstract accepted for publication at Proceedings of the 26th ACM Conference on Economics and Computation
Abstract
We present a decision-theoretic justification for viewing the question of how to best choose where to experiment
in order to optimize external validity as a k-median (clustering) problem, a popular problem in computer science
and operations research. We present conditions under which minimizing the worst-case, welfare-based regret among
all nonrandom schemes that select k sites to experiment is approximately equal—and sometimes exactly equal—to
finding the k most central vectors of baseline site-level covariates. The k-median problem can be formulated as a
linear integer program. Two empirical applications illustrate the theoretical and computational benefits of the
suggested procedure.