Avaliando os Potenciais e as Limitações dos Algoritmos de Imputação em Bases de Dados Clínicos
Missing values are a central challenge in clinical datasets, directly affecting the reliability of predictive models. This work compares traditional imputation strategies (mean, KNN, MICE, and MissForest) with a foundation-model-based approach (TabPFN), evaluating both imputation accuracy and impact on mortality prediction. We introduce a quadrant-based structural analysis of the patient feature space to examine how local data organization influences performance. Results show that imputation effectiveness depends on clinical structure and that TabPFN demonstrates greater robustness in heterogeneous regions with stable computational cost. Overall, we show that imputation is not a neutral preprocessing step: its integration into the predictive pipeline significantly affects outcomes, highlighting the need for adaptive, task-oriented strategies in real-world health-care settings.
2026/1 - POC2
Orientador: Marcos André Gonçalves
Palavras-chave: Imputação de dados clínicos, aprendizado de máquina, modelos de fundação, KNN, TabPFN;
PDF Disponível