Power, Control, and Data Processing Systems

Power, Control, and Data Processing Systems

A New Method for Detecting Emotions in Video Using a Hybrid CNN-RNN-RBM Network

Document Type : Original Research

Authors
1 M.Sc. Student, Electrical and Computer Engineering Department, Hamedan University of Technology, Hamedan, Iran
2 Electrical and Computer Engineering Department, Hamedan University of Technology, Hamedan, Iran
Abstract
Facial expressions are one of the most important techniques of non-verbal communication that humans use to display their emotional states and establish non-verbal communication. In recent research in the field of emotion recognition based on facial expression changes, deep learning networks are used. In this regard, an appropriate network is selected according to the type of data. In the present research, video data is used, and for recognizing facial expressions, recurrent networks are employed. Subsequently, to achieve better results, a hybrid network consisting of Convolutional, Recurrent, and Restricted Boltzmann Machine networks has also been used. Finally, in this paper, using deep learning techniques and the Python language, an algorithm based on a hybrid CNN-LSTM-RBM neural network for emotion recognition through facial expressions in a video is presented. In this method, seven emotional states (happy, sad, disgust, excitement, contempt, fear, and anger) are detected. The CK+48 dataset has been used to train and test the proposed model. The experimental results demonstrate that the proposed Hybrid CNN-LSTM-GRBM model achieves a training accuracy of 99.06%. Crucially, the model attains 91.16% test accuracy and an F1-score of 89.27% on the test set, outperforming existing baseline methods and confirming its robustness in handling spatiotemporal facial features.
Keywords
Subjects

[1]    D. Joshi et al., “Aesthetics and Emotions in Images,” IEEE Signal Processing Magazine, vol. 28, no. 5, pp. 94–115, Sep. 2011, doi: 10.1109/msp.2011.941851.
[2]    D. K. Jain, P. Shamsolmoali, and P. Sehdev, “Extended deep neural network for facial emotion recognition,” Pattern Recognition Letters, vol. 120, pp. 69–74, Apr. 2019, doi: 10.1016/j.patrec.2019.01.008.
[3]    N. Jain, S. Kumar, A. Kumar, P. Shamsolmoali, and M. Zareapoor, “Hybrid deep neural networks for face emotion recognition,” Pattern Recognition Letters, vol. 115, pp. 101–106, Nov. 2018, doi: 10.1016/j.patrec.2018.04.010.
[4]    M. Nemanja. “Convolutions and convolutional neural networks,” Introduction to Convolutional Neural Networks vol. 11, pp. 435-479, Aug. 2020. doi: 10.1007/978-1-4842-5648-0_12.
[5]    J.C. Baird, “Information theory and information processing,” Information processing & management, vol. 20, pp.373-381, Jun. 1984. doi: 10.1016/0306-4573(84)90068-2.
[6]    G. E. Hinton, “Training Products of Experts by Minimizing Contrastive Divergence,” Neural Computation, vol. 14, no. 8, pp. 1771–1800, Aug. 2002, doi: 10.1162/089976602760128018.
[7]    Y. Muneki, and X. Zhongren, “New learning algorithm of gaussian–bernoulli restricted boltzmann machine and its application in feature extraction," IEICE Proceedings Series, 76 (A3L-42). doi:10.34385/proc.76.A3L-42.
[8]    G. E. Hinton and R. R. Salakhutdinov, “Reducing the Dimensionality of Data with Neural Networks,” Science, vol. 313, no. 5786, pp. 504–507, Jul. 2006, doi: 10.1126/science.1127647.
[9]    L. F. Barrett, R. Adolphs, S. Marsella, A. M. Martinez, and S. D. Pollak, “Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements,” Psychological Science in the Public Interest, vol. 20, no. 1, pp. 1–68, Jul. 2019, doi: 10.1177/1529100619832930.
[10]    R. M. D and H. H. Kenchannavar, “Hybrid Deep Optimal Network for Recognizing Emotions Using Facial Expressions at Real Time,” International Journal of Intelligent Systems and Applications, vol. 16, no. 3, pp. 47–58, Jun. 2024, doi: 10.5815/ijisa.2024.03.04.
[11]    M. N. Ab Wahab, A. Nazir, A. T. Zhen Ren, M. H. Mohd Noor, M. F. Akbar, and A. S. A. Mohamed, “Efficientnet-Lite and Hybrid CNN-KNN Implementation for Facial Expression Recognition on Raspberry Pi,” IEEE Access, vol. 9, pp. 134065–134080, 2021, doi: 10.1109/access.2021.3113337.
[12]    X. Li, T. Pfister, X. Huang, G. Zhao, M. Pietikäinen, “A spontaneous micro-expression database: Inducement, collection and baseline,” In 10th IEEE International Conference and Workshops on Automatic face and gesture recognition (fg) Apr. 2013, pp. 1-6. doi: 10.1109/fg.2013.6553717.
[13]    I. Lakshan, L. Wickramasinghe, S. Disala, S. Chandrasegar, P. Haddela, “Real time deception detection for criminal investigation,” In2019 National information technology conference (NITC) Oct. 2019, pp. 90-96. doi: 10.1109/nitc48475.2019.9114422.
[14]    H. Karimi, J. Tang , Y. Li,  “Toward end-to-end deception detection in videos,” In2018 IEEE international conference on big data (Big Data) Dec. 2018, pp. 1278-1283. doi: 10.1109/bigdata.2018.8621909.
[15]    B. Martinez, M. F. Valstar, B. Jiang, and M. Pantic, “Automatic Analysis of Facial Actions: A Survey,” IEEE Transactions on Affective Computing, pp. 17–25, 2017, doi: 10.1109/taffc.2017.2731763.
[16]    E. Sariyanidi, H. Gunes, and A. Cavallaro, “Automatic Analysis of Facial Affect: A Survey of Registration, Representation, and Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 6, pp. 1113–1133, Jun. 2015, doi: 10.1109/tpami.2014.2366127.
[17]    S. Zhang, X. Pan, Y. Cui, X. Zhao, and L. Liu, “Learning Affective Video Features for Facial Expression Recognition via Hybrid Deep Learning,” IEEE Access, vol. 7, pp. 32297–32304, Mar. 2019, doi: 10.1109/access.2019.2901521.
[18]    D. Tran, L. Bourdev, R. Fergus, L. Torresani, M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” InProceedings of the IEEE international conference on computer vision, 2015, pp. 4489-4497. doi: 10.1109/iccv.2015.510.
[19]    B. Hasani, M. H. Mahoor,  “Facial expression recognition using enhanced deep 3D convolutional neural networks,” In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2017, pp. 30-40. doi: 10.1109/cvprw.2017.282.
[20]    S. Zhang, S. Zhang, T. Huang, W. Gao, and Q. Tian, “Learning Affective Features With a Hybrid Deep Model for Audio–Visual Emotion Recognition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 10, pp. 3030–3043, Oct. 2018, doi: 10.1109/tcsvt.2017.2719043.
[21]    H. V. Manalu, and A. P. Rifai,  “Detection of human emotions through facial expressions using hybrid convolutional neural network-recurrent neural network algorithm,” Intelligent systems with applications, 2024, pp. 330-339. doi: 10.1016/j.iswa.2024.200339.
[22]    A. Chaudhari, C. Bhatt, A. Krishna, and P. L. Mazzeo, “ViTFER: Facial Emotion Recognition with Vision Transformers,” Applied System Innovation, vol. 5, no. 4, p. 80, Aug. 2022, doi: 10.3390/asi5040080.
[23]    P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, I. Matthews, “The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,” In2010 ieee computer society conference on computer vision and pattern recognition-workshops, Jun. 2010 pp. 94-101. doi: 10.1109/cvprw.2010.5543262.
[24]    J. Zhang, H. Wang, J. Chu, S. Huang, T. Li, and Q. Zhao, “Improved Gaussian–Bernoulli restricted Boltzmann machine for learning discriminative representations,” Knowledge-Based Systems, vol. 185, p. 104911, Dec. 2019, doi: 10.1016/j.knosys.2019.104911.
[25]    O. Prathwini and L. Prathyakshini, “DeepEmoVision: Unveiling Emotion Dynamics in Video Through Deep Learning Algorithms,” International Journal of Advanced Computer Science and Applications, vol. 15, no. 3, 2024, doi: 10.14569/ijacsa.2024.0150388.
[26]    G. E. Hinton, “A practical guide to training restricted Boltzmann machines,” InNeural Networks: Tricks of the Trade, Jan 2018, pp. 599-619. doi: 10.1007/978-3-642-35289-8_32.
[27]    K. Cho, T. Raiko, A. T. Ihler, “Enhanced gradient and adaptive learning rate for training restricted Boltzmann machines,” InProceedings of the 28th international conference on machine learning (ICML-11), 2011, pp. 105-112. doi: 10.1162/neco_a_00397.
[28]    N. Yalçin, M. Alisawi, “Introducing a novel dataset for facial emotion recognition and demonstrating significant enhancements in deep learning performance through pre-processing techniques,” Heliyon. Vol. 30, no. 10, Oct. 2024. doi: 10.1016/j.heliyon.2024.e38913.
[29]    J. R. Stevens, R. Venkatesan, S. Dai, B. Khailany and A. Raghunathan, “Softermax: Hardware/software co-design of an efficient softmax for transformers,” In 2021 58th ACM/IEEE Design Automation Conference (DAC), Dec. 2021, pp. 469-474. doi: 10.1109/dac18074.2021.9586134.
[30]    N. Gupta, R. V. Priya, and C. K. Verma, “Video-based emotion recognition using motion-aware deep hybrid learning,” Cluster Computing, vol. 29, no. 1, Nov. 2025, doi: 10.1007/s10586-025-05807-x.
[31]    B. V. Gokulnath et al., “Empowering emotional intelligence through deep learning techniques,” Scientific Reports, vol. 16, no. 1, Dec. 2025, doi: 10.1038/s41598-025-29073-4.
[32]    Q. Zhang, Y. Liu, B. Zhu, X. Han, R. Zhang, J. Xiao, Z. Wang, “Deep multi-modal fusion transformer for emotion recognition,” Proceeding in Engineering Applications of Artificial Intelligence. Mar. 2026, pp. 116-129. doi: 10.1016/j.engappai.2026.113967.
[33]    Z.Y. Huang, C. C. Chiang, J. H. Chen, “A study on computer vision for facial emotion recognition,” Proceeding in neural networks Sep. 2023, pp.258-270. doi: 10.1038/s41598-023-35446-4.
[34]    M. Luo, “Machine learning for time series analysis and forecasting,”, Master's thesis, Northeastern University, 2023, pp.62-63. doi: 10.17760/d20487630.
Volume 3, Issue 3
Summer 2026
Pages 18-26

  • Receive Date 29 April 2026
  • Revise Date 29 June 2026
  • Accept Date 01 July 2026
  • First Publish Date 01 July 2026
  • Publish Date 01 September 2026