Beyond Single-Objective Accuracy: A Longitudinal Evaluation of Tabular Foundation Models under Dataset Shift

João Marcos Tomáz Silva Campos

Evaluations of classification under dataset shift often rely on static benchmarks and single-objective metrics, obscuring how discrimination, calibration, and risk ranking may fail in different ways over time. We present a 19-year longitudinal benchmark of tabular models under real-world clinical shift, using a Brazilian renal registry with 9.6 million records from 1997–2015. We compare dynamically retrained tree-based ensembles with standard and drift-aware Tabular Foundation Models (TabPFN, TabICL) under uniform temporal protocols. Our results show that robustness is not monolithic: models with strong global survival rankings may suffer from calibration and same-year classification failures, while model rankings change across stable, covariate-shift, and structural/observation-process- shift regimes. We further introduce an exploratory Temporal Robustness Index and combine domain-aware SHAP with medication-record perturbations to diagnose model failures. The analysis reveals that standard models often rely on temporally unstable pharmaceutical proxies, whereas temporally conditioned models emphasize more stable physiological signals. Overall, our study provides a multi-objective and shift-aware auditing framework for evaluating tabular foundation models in dynamic clinical environments.


2026/1 - POC2

Orientador: Marcos André Gonçalves, Wagner Meira Junior, Leonardo Rocha

Palavras-chave: Dataset Shift, Calibration, Temporal Robustness, Tabular Foundation Models, Clinical Forecasting

PDF Disponível