ARCHIVES

Original Article

Deep Learning-Based Multimodal Emotion Recognition Using Facial Expressions and Physiological Signals with Dynamic Fusion for Real-Time Human–Computer Interaction

Shinde Akshata G1 Shinde S.G2
1 PG Scholar, TPCT’s College of Engineering, Dharashiv, Maharashtra, India. 2 Associate Professor, TPCT’s College of Engineering, Dharashiv, Maharashtra, India.

Published Online: July-August 2026

Pages: 198-207

Abstract

Emotion recognition has become an essential component of intelligent human–computer interaction, healthcare monitoring, driver assistance, and affective computing systems. Conventional facial expression-based emotion recognition methods often exhibit limited robustness because facial expressions may be intentionally suppressed, partially occluded, or inadequately represent an individual's actual emotional state. To address these limitations, this paper presents a deep learning-based multimodal emotion recognition framework that combines facial expression analysis with physiological signal processing to improve the reliability of real-time emotion classification. The proposed system employs a ResNet-18 convolutional neural network trained on the FER-2013 dataset for facial emotion recognition, while physiological information is acquired using a MAX30102 photoplethysmography (PPG) sensor and a Galvanic Skin Response (GSR) sensor interfaced with an STM32F411 microcontroller. The acquired physiological signals are processed to extract heart rate, heart rate variability (HRV), and skin conductance features that characterize the user's emotional arousal. A dynamic decision-level fusion strategy integrates facial prediction probabilities with physiological indicators using individualized baseline calibration, enabling effective recognition of concealed or ambiguous emotional states. The proposed framework operates in real time by combining embedded hardware, deep learning inference, and adaptive multimodal fusion within a unified architecture. Experimental investigations demonstrate that the multimodal approach provides more reliable emotion recognition than facial analysis alone, particularly for emotions associated with elevated physiological responses such as fear, stress, and surprise. Representative experimental results confirm stable real-time operation under different emotional conditions while maintaining computational efficiency suitable for practical deployment. Owing to its low-cost implementation, adaptive fusion mechanism, and real-time performance, the proposed framework is a promising solution for next-generation affective computing, intelligent healthcare, smart surveillance, wearable monitoring, and human–machine interaction applications.

Related Articles

2026

A Strategic Framework for Depth-Dependent Hydroelectric Conversion along the Indian Coastline

2026

Reimagining Development in India: A Critical Analysis of the Viksit Bharat Vision

2026

AI-Enabled Image Description: Bridging the Gap for the Visually Impaired

2026

Perceived Occupational Risks of Emergency Medical Services Personnel

2026

Origin, Growth and recent Development of Integrated Reporting (IR): A theoretical Review

2026

Smart Hostel Management System

Share Article

X
LinkedIn
Facebook
WhatsApp

Or copy link

https://www.ijrtmr.com/archives/deep-learning-based-multimodal-emotion-recognition-using-facial-expressions-and-physiological-signals-with-dynamic-fusion-for-real-time-human-computer-interaction

*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.