The invention relates to a multi-mode user
emotion recognition method, device and equipment and a medium, the method is applied to sound equipment, and the method comprises the steps that physiological data,
voice data and image data of a target user are acquired; extracting physiological features from the physiological data; extracting voice features and semantic features from the
voice data; extracting facial features from the image data; and determining the emotion of the target user according to the physiological features, the voice features, the semantic features and the facial features. Therefore, according to the technical scheme of the invention, the accurate judgment of the emotion of the target user is realized by integrating the four types of
modal data of the physiological features, the voice features, the semantic features and the facial features. The four types of features respectively map the emotional state of the user from four dimensions of
physiological reaction, acoustic expression, semantic connotation and visual expression to form a multi-dimensional complementary
emotion recognition system. The mechanism can recommend appropriate music content to the user, the auditory demands of the user under different emotions are met, and the use experience of the user is remarkably improved.