The invention relates to a multi-
modal emotion recognition fusion and communication method oriented to real-time human-computer interaction, and aims to solve the problems of asynchronous
modal time sequence, non-uniform
feature dimension, single fusion mode, high communication
delay and the like of an existing
emotion recognition system and improve the recognition precision and real-time performance of the
system. According to the method, voice, images and physiological signals are synchronously collected through a multi-
modal input module, feature alignment is achieved through
time sequence interpolation,
dynamic time warping and space
coordinate mapping, double-domain
feature fusion is conducted in combination with a
time path network and a space
path network, emotional state judgment is completed through a lightweight neural network, and the emotional
state recognition accuracy is improved. And outputting six types of basic emotions and confidence coefficients. Meanwhile, a low-
delay communication protocol based on UDP clipping extension realizes rapid feedback of emotion data, and in combination with modal
priority scheduling,
data compression and bandwidth sensing mechanisms, high-efficiency and low-
delay transmission is ensured, and the real-time interaction requirement in a weak network environment is met. The method has the advantages of high accuracy, low
power consumption, low time delay and flexible deployment, is suitable for various real-time interaction application scenes such as voice assistants, virtual customer service, emotion accompanying and telemedicine, and has wide application value and market prospect.