arXiv Analytics

Sign in

arXiv:1810.11953 [stat.ML]AbstractReferencesReviewsResources

Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift

Stephan Rabanser, Stephan Günnemann, Zachary C. Lipton

Published 2018-10-29Version 1

We might hope that when faced with unexpected inputs, well-designed software systems would fire off warnings. Machine learning (ML) systems, however, which depend strongly on properties of their inputs (e.g. the i.i.d. assumption), tend to fail silently. This paper explores the problem of building ML systems that fail loudly, investigating methods for detecting dataset shift and identifying exemplars that most typify the shift. We focus on several datasets and various perturbations to both covariates and label distributions with varying magnitudes and fractions of data affected. Interestingly, we show that while classifier-based methods perform well in high-data settings, they perform poorly in low-data settings. Moreover, across the dataset shifts that we explore, a two-sample-testing-based approach, using pretrained classifiers for dimensionality reduction performs best.

Comments: Submitted to the NIPS 2018 Workshop on Security in Machine Learning
Categories: stat.ML, cs.LG
Related articles: Most relevant | Search more
arXiv:1201.1450 [stat.ML] (Published 2012-01-06)
The Interaction of Entropy-Based Discretization and Sample Size: An Empirical Study
arXiv:2004.05007 [stat.ML] (Published 2020-04-10)
An Empirical Study of Invariant Risk Minimization
arXiv:1812.07909 [stat.ML] (Published 2018-12-19)
An Empirical Study of Generative Models with Encoders