ResearchIn developmentDaffodil International UniversityWork period · 2025-2026Ongoing Research

Hybrid Deep Models for Mental Health Detection with XAI Techniques

Hybrid DenseNet–ViT–GRU architecture with Explainable AI (Grad-CAM, LIME, SHAP) for diagnosing psychological stability from voice spectrograms. Combines local, global, and sequential feature learning with transparent, interpretable predictions.

Hybrid Deep Models for Mental Health Detection with XAI Techniques project workflow and system architecture
Project workflow and system architecture overview. Open full-size diagram

Workflow at a glance

  1. Prepare voice inputs

    Build spectrogram representations.

  2. Explore hybrid learning

    Study DenseNet, ViT, and GRU components.

  3. Analyze explanations

    Investigate model interpretation methods.

  4. Ongoing evaluation

    Research remains in development.

Project facts & reported results

~89.81%
accuracy

Reported result

Historical project result; evaluation not independently reproduced.

0.962
Mean AUC

Implementation fact

0.95–0.97
AUC Fold Range

Implementation fact

Grad-CAM, LIME, SHAP
XAI Methods

Implementation fact

Project Overview

Problem Statement

Mental health diagnostics are often subjective and dependent on self-reports and clinician observation. Existing deep learning models lack transparency in clinical decision-making. There is a need for interpretable, multi-level feature extraction that captures local patterns, global context, and temporal sequences in voice data simultaneously.

Approach & Methodology

Proposed a three-stage hybrid deep learning architecture: DenseNet-style CNN for extracting local time–frequency patterns from log-mel spectrograms, Vision Transformer (ViT) for capturing long-range global dependencies, and GRU for sequential aggregation of learned representations. Applied SMOTE on training spectrogram features for class balancing. Used SpecAugment-style time/frequency masking and random shifts for augmentation. Evaluated with 5-Fold Cross Validation. Integrated three XAI methods — Grad-CAM heatmaps, LIME local explanations, and SHAP global/local attributions — to interpret model predictions on spectrograms.

Outcome & Scope

Achieved ~89.81% combined accuracy across all folds with mean AUC of 0.962. Grad-CAM highlights the most influential time–frequency regions used by the model. LIME identifies positive/negative contributing spectrogram regions. SHAP provides Shapley-value attribution for both stable and unstable predictions. Research carried out under Dr. Md. Taimur Ahad at the 4IR Research Cell, DIU. Manuscript in preparation. Dataset not included due to privacy/ethical constraints.

System Architecture

Hybrid DenseNet–ViT–GRU pipeline with XAI for voice-based mental health detection

Key Components:

Audio Preprocessing

Log-mel spectrogram extraction at 48 kHz sample rate, 2-second segments, 128 mel bands

DenseNet Feature Extractor

Dense connectivity CNN extracting local time–frequency patterns from spectrograms

Vision Transformer (ViT)

Captures long-range global dependencies across the spectrogram using self-attention

GRU Sequence Aggregator

Gated Recurrent Unit for sequential aggregation of learned multi-level representations

XAI Interpretation Layer

Grad-CAM heatmaps, LIME local explanations, and SHAP attributions for transparent predictions

Key Features & Capabilities

Log-mel spectrogram generation from audio signals (128 mel bands, 48 kHz)

Hybrid DenseNet → ViT → GRU architecture for multi-level feature extraction

SMOTE applied on training features for class balancing

SpecAugment-style time/frequency masking and random shifts

5-Fold Cross Validation with stable ROC-AUC across folds

Grad-CAM heatmap visualization for model attention

LIME local explanations (positive/negative contributing regions)

SHAP global and local attributions on spectrograms

Current Scope & Limitations

Research use only. These experiments do not establish a clinically validated diagnostic or screening tool.

Technologies & Tools

TensorFlow
Keras
DenseNet
Vision Transformer
GRU
LIME
SHAP
Grad-CAM
Librosa
OpenCV
scikit-learn