CORONARY HEART DISEASE screening accuracy improved substantially when tongue and facial images were analysed together using a multimodal deep learning model, according to new research involving 518 participants.
Researchers developed a dual branch artificial intelligence model designed to identify coronary heart disease using standardised tongue and facial photographs collected from 200 patients with coronary heart disease and 318 healthy controls.
Multimodal Coronary Heart Disease Screening Outperforms Single Modalities
The investigators used two deep learning backbones, ResNet50 and EfficientNet-B1, to evaluate whether combining tongue and facial images could improve diagnostic performance compared with analysing either image type alone.
Images underwent standardised preprocessing before being analysed through a dual branch architecture that fused features extracted from both modalities.
On an independent test set of 104 participants, the ResNet50 multimodal model achieved the strongest performance, with an F1 score of 0.9756 and an area under the curve (AUC) of 0.9992.
This exceeded the performance of both the tongue only model (F1: 0.9524; AUC: 0.9867) and the face only model (F1: 0.9250; AUC: 0.9887). The multimodal model correctly classified 102 of 104 test samples.
Consistent Results Across Validation Analyses
To examine model robustness, the researchers conducted fivefold stratified cross validation across all 518 participants.
The multimodal approaches consistently achieved the highest performance and showed minimal variation between folds.
For ResNet50, multimodal fusion achieved an AUC of 0.9980±0.0027 and an F1 score of 0.9827±0.0141. The EfficientNet-B1 multimodal model produced an AUC of 0.9987±0.0014 and an F1 score of 0.9829±0.0138.
These findings indicated strong stability and reproducibility across different data partitions.
Explainable AI Reveals Key Image Regions
The study also incorporated gradient-weighted class activation mapping to explore how the model reached its decisions.
Visualisation analyses showed that the tongue branch consistently focused on the central tongue body and coating, while the facial branch concentrated on the perioral and zygomatic buccal regions.
According to the researchers, these findings demonstrate that tongue and facial images provide complementary information for coronary heart disease detection.
Although external validation is still required, the results suggest that multimodal image analysis could offer an effective and interpretable pathway for future non-invasive coronary heart disease screening.
Reference
Wang Y et al. Explainable multimodal deep learning fusion of tongue and facial images for noninvasive coronary heart disease screening. Sci Rep. 2026;DOI:10.1038/s41598-026-71871-x.