The invention discloses an
emotion recognition method,
system and device based on multi-
modal adaptive fusion and a storage medium, and relates to the technical field of
artificial intelligence, and the method comprises the steps: selecting a pre-training model, respectively extracting the original features of an audio and a video, carrying out the preliminary extraction of the audio through a
convolution layer, carrying out the multi-module
processing of the video, and keeping the
time sequence information. Constructing an attention module to generate an attention matrix and interaction features, and adjusting the original features by using the matrix; and inputting the weighted and fused features into a convolutional network to extract advanced
time sequence features, performing
pooling compression on the advanced
time sequence features in a time dimension, splicing audio and video features, and finally sending the spliced audio and video features into a full-connection layer classifier to obtain an
emotion classification result. According to the method, the weights of different features can be dynamically adjusted, so that the audio and visual features are effectively fused, the accuracy and robustness of
emotion recognition are improved, the weighted
recall rate and the unweighted
recall rate are remarkably improved, and the method has high calculation efficiency and expandability.