ARCHIVES
Beyond Gradient Boosting: A Comparative Deep Learning Framework for Binary Symptom-Based Tuberculosis Screening Using Tabular Neural Architectures and Explainable AI
Published Online: July-August 2026
Pages: 239-255
Cite this article
↗ https://www.doi.org/10.59256/ijrtmr.20260604028Abstract
Tuberculosis (TB) remains one of the leading infectious causes of death worldwide, and low-cost symptom-based pre-screening is a critical first line of defence in resource-constrained settings. A previous study on this dataset demonstrated that classical machine learning — most notably Gradient Boosting — could classify TB status from 13 binary symptom indicators with near-perfect accuracy. The present study asks a different, complementary question: can modern deep-learning architectures purpose-built for tabular data match or improve upon that classical baseline, and what do they reveal about the structure of the problem that tree ensembles do not? We benchmark six deep-learning architectures — a Deep Neural Network (DNN/MLP), a Wide & Deep network, TabTransformer, FT-Transformer, Neural Oblivious Decision Ensembles (NODE), and TabNet — on the same 2,167-patient, 13-symptom Kaggle tuberculosis dataset (51.78% TB-positive, 48.22% TB-negative). Each model was trained with Adam optimisation, early stopping, ReduceLROnPlateau learning- rate scheduling, dropout, and batch/layer normalisation, and evaluated on a stratified 80:20 hold-out split together with stratified 10-fold cross-validation. On the hold-out test set (n = 434), the Wide & Deep network achieved the best overall performance (accuracy 99.08%, F1-score 99.10%, AUC-ROC 0.9963, MCC 0.9817, Cohen's κ 0.9816), closely followed by the DNN (accuracy 98.85%, AUC 0.9964). TabTransformer, FT-Transformer, and NODE achieved strong but slightly lower accuracies (97.70– 98.16%), while TabNet was markedly less stable (holdout accuracy 83.41%; 10-fold accuracy 78.35% ± 12.70%), a finding we analyse in detail. SHAP-based explainability, applied to the two best-performing models, identified swollen lymph nodes, fever for two weeks, persistent productive cough, coughing blood, and chest pain as the dominant predictive symptoms — broadly consistent with, but not identical to, the impurity-based ranking obtained by the earlier Gradient Boosting model. Critically, none of the six deep architectures exceeded the 99.69% accuracy previously reported for tuned Gradient Boosting on this dataset, reinforcing recent evidence that, on small-to-medium, low-dimensional tabular problems, well-regularised deep networks can match but do not automatically surpass strong tree ensembles. We discuss why this is itself a useful and underreported finding for the clinical machinelearning literature, analyse the six architectures' comparative computational complexity, and outline a roadmap for scaling the deep-learning approach to richer, multi-modal TB data where its architectural advantages are more likely to be realised.
Related Articles
2026
A Strategic Framework for Depth-Dependent Hydroelectric Conversion along the Indian Coastline
2026
Reimagining Development in India: A Critical Analysis of the Viksit Bharat Vision
2026
AI-Enabled Image Description: Bridging the Gap for the Visually Impaired
2026
Perceived Occupational Risks of Emergency Medical Services Personnel
2026
Origin, Growth and recent Development of Integrated Reporting (IR): A theoretical Review
2026
Smart Hostel Management System
Share Article
Or copy link
*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.