
Hybrid Deep Models for Mental Health Detection with XAI Techniques
Hybrid DenseNet–ViT–GRU architecture with Explainable AI (Grad-CAM, LIME, SHAP) for diagnosing psychological stability from voice spectrograms. Combines local, global, and sequential feature learning with transparent, interpretable predictions.

Workflow at a glance
Prepare voice inputs
Build spectrogram representations.
Explore hybrid learning
Study DenseNet, ViT, and GRU components.
Analyze explanations
Investigate model interpretation methods.
Ongoing evaluation
Research remains in development.
Project facts & reported results
Reported result
Historical project result; evaluation not independently reproduced.
Implementation fact
Implementation fact
Implementation fact
Project Overview
Problem Statement
Mental health diagnostics are often subjective and dependent on self-reports and clinician observation. Existing deep learning models lack transparency in clinical decision-making. There is a need for interpretable, multi-level feature extraction that captures local patterns, global context, and temporal sequences in voice data simultaneously.
Approach & Methodology
Proposed a three-stage hybrid deep learning architecture: DenseNet-style CNN for extracting local time–frequency patterns from log-mel spectrograms, Vision Transformer (ViT) for capturing long-range global dependencies, and GRU for sequential aggregation of learned representations. Applied SMOTE on training spectrogram features for class balancing. Used SpecAugment-style time/frequency masking and random shifts for augmentation. Evaluated with 5-Fold Cross Validation. Integrated three XAI methods — Grad-CAM heatmaps, LIME local explanations, and SHAP global/local attributions — to interpret model predictions on spectrograms.
Outcome & Scope
Achieved ~89.81% combined accuracy across all folds with mean AUC of 0.962. Grad-CAM highlights the most influential time–frequency regions used by the model. LIME identifies positive/negative contributing spectrogram regions. SHAP provides Shapley-value attribution for both stable and unstable predictions. Research carried out under Dr. Md. Taimur Ahad at the 4IR Research Cell, DIU. Manuscript in preparation. Dataset not included due to privacy/ethical constraints.
System Architecture
Hybrid DenseNet–ViT–GRU pipeline with XAI for voice-based mental health detection
Key Components:
Log-mel spectrogram extraction at 48 kHz sample rate, 2-second segments, 128 mel bands
Dense connectivity CNN extracting local time–frequency patterns from spectrograms
Captures long-range global dependencies across the spectrogram using self-attention
Gated Recurrent Unit for sequential aggregation of learned multi-level representations
Grad-CAM heatmaps, LIME local explanations, and SHAP attributions for transparent predictions
Key Features & Capabilities
Log-mel spectrogram generation from audio signals (128 mel bands, 48 kHz)
Hybrid DenseNet → ViT → GRU architecture for multi-level feature extraction
SMOTE applied on training features for class balancing
SpecAugment-style time/frequency masking and random shifts
5-Fold Cross Validation with stable ROC-AUC across folds
Grad-CAM heatmap visualization for model attention
LIME local explanations (positive/negative contributing regions)
SHAP global and local attributions on spectrograms
Current Scope & Limitations
Research use only. These experiments do not establish a clinically validated diagnostic or screening tool.