Contrast learning audio and video emotion recognition method based on canonical correlation analysis

Through a comparison learning method based on typical correlation analysis, deep semantic alignment of audio and video features in the shared feature space is achieved, and the problem of difficulty in aligning audio and video features in the prior art is solved, and the accuracy of emotion recognition is improved.

CN120337153APending Publication Date: 2025-07-18XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510504280.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve deep semantic alignment of audio features and video features, resulting in low accuracy of emotion recognition.

Method used

Using a comparison learning method based on typical correlation analysis, the embedded feature normalization of the graph structure enhanced by the logarithmic Mel spectrogram and standard image frame set output by the graph neural network module, and combined with the InfoNCE comparison loss, the semantic distance between the same sample audio features and video features is narrowed, and the semantic distance between different sample audio features and video features is pushed away, so as to achieve deep semantic alignment of audio and video features in the shared feature space.

Benefits of technology

It effectively improves the accuracy of audio and video emotions recognition, realizes deep semantic alignment of audio and video features, and improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337153A_ABST
    Figure CN120337153A_ABST
Patent Text Reader

Abstract

The invention provides a canonical correlation analysis-based comparative learning audio and video emotion recognition method. The implementation steps are as follows: obtaining a training sample set and a test sample set; constructing a canonical correlation analysis-based comparative learning audio and video emotion recognition network model and carrying out iterative training on the model; and obtaining an audio and video emotion recognition result. According to the invention, the CCA module normalizes the logarithm Mel spectrogram output by the graph neural network module and four groups of embedded features of the graph structure after standard image frame set enhancement, realizes semantic alignment of audio and video features in a shared feature space, and combines with contrast learning based on InfoNCE contrast loss, so as to realize the semantic alignment of the audio and video features in the shared feature space. The semantic distance between the audio features and the video features of the same sample is shortened, and the semantic distance between the audio features and the video features of different samples is increased at the same time, so that deep semantic alignment of the audio features and the video features is realized, and the accuracy of audio and video emotion recognition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and relates to a method for audio-visual emotion recognition, specifically to a contrastive learning audio-visual emotion recognition method based on canonical correlation analysis, which can be used in fields such as human-computer interaction, online education, and mental health monitoring. Background Art

[0002] Emotion recognition refers to the process of identifying and understanding the emotional states of individuals or groups by analyzing various forms of human expression such as speech, facial expressions, and body movements. It is a method that combines technologies such as artificial intelligence, machine learning, computer vision, and natural language processing, aiming to achieve automatic detection and analysis of human emotions. With the continuous in-depth research on multi-modal emotion recognition, a single modality can no longer fully express complex human emotions. Considering the intuitive forms of facial expressions and speech features, multi-modal emotion recognition methods based on audio-visual have gradually become a research hotspot.

[0003] Machine learning or deep learning audio-visual emotion recognition methods first extract features from audio and video signals, then fuse the extracted features and input them into a classification model, and finally output the predicted emotion category labels. However, in the field of emotion recognition, the existing technology is far from reaching the accuracy level of humans. There are several main factors affecting the accuracy of existing audio-visual emotion recognition methods: (1) the sufficiency of feature extraction from audio and video signals; (2) the effective fusion of audio features and video features; (3) the semantic alignment of audio features and video features.

[0004] Since audio and video belong to different types of data, how to achieve semantic alignment between audio features and video features has become the key to improving the accuracy of emotion recognition. For example, a patent application with the application number CN202311702397.4, the publication number CN117708754A, and the title "A Multimodal Emotion Recognition Method and Device" discloses an audio-visual emotion recognition method. The invention preprocesses the original multimodal data samples, obtains the primary features of each unimodal and inputs them into the unimodal feature extraction network respectively, extracts the high-level features of each unimodal, and after one-dimensionalization, inputs them into each unimodal recognition network respectively, and trains the unimodal models composed of each unimodal feature extraction network and unimodal recognition network; obtains the weights of each unimodal feature extraction network, and extracts the activation degrees of all neurons in each unimodal recognition network; determines the type of each neuron according to the activation degree of each neuron, and obtains the unique connection constraints between different neurons; based on the weights of each unimodal feature extraction network obtained and the connection constraints extracted, establishes a complete multimodal emotion recognition model and completes the training, and tests the trained multimodal emotion recognition model to obtain the emotion recognition result. The invention realizes cross-modal interaction and can effectively recognize audio-visual emotions. However, because it only relies on the activation degree and type of neurons to form connection constraints to achieve cross-modal interaction, it is essentially an implicit association between activation features, and it is difficult to achieve deep semantic alignment between audio features and video features, which affects the further improvement of the recognition accuracy. Summary of the Invention

[0005] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and propose a contrastive learning audio-visual emotion recognition method based on canonical correlation analysis, which is used to solve the technical problem of low emotion recognition accuracy existing in the prior art due to the difficulty in achieving deep semantic alignment between audio features and video features.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] (1) Obtain a training sample set and a test sample set:

[0008] Obtain M audio and video signals including multiple emotion categories, preprocess each audio signal and the video signal with the same emotion category as it, then form a data pair from the corresponding logarithmic Mel spectrogram and standard image frame set of the mth preprocessed audio and video signal, and then randomly select more than half of the data pairs and the emotion category labels announced in the file names of each audio or video signal to form a training sample set, and form the remaining data pairs into a test sample set, where M>1000;

[0009] (2) Construct a contrastive learning audio-visual emotion recognition network model based on canonical correlation analysis:

[0010] Construct an audio - video emotion recognition network model O including a contrastive learning network and an emotion classification module cascaded with it. Among them, the contrastive learning network includes audio processing modules U arranged in parallel and with the same structure a and video processing modules U v , U a and U v All include a cascaded graph construction and enhancement module, a graph neural network module, and a canonical correlation analysis CCA module;

[0011] (3) Iteratively train the audio - video emotion recognition network model:

[0012] Randomly select N training samples from the training sample set as the input of the audio - video emotion recognition network model O and perform iterative training on it to obtain the trained audio - video emotion recognition network model O*;

[0013] (4) Obtain the audio - video emotion recognition result:

[0014] Use the test sample set as the input of the trained audio - video emotion recognition network model O* for forward propagation to obtain the emotion recognition result of each test sample.

[0015] Compared with the prior art, the present invention has the following advantages:

[0016] In the process of training the audio - video emotion recognition network model and obtaining the audio - video emotion recognition result, the CCA module normalizes four groups of embedded features of the logarithmic mel - spectrogram output by the graph neural network module and the graph structure after enhancing the standard image frame set, realizes the semantic alignment of audio and video features in the shared feature space, and combines contrastive learning based on the InfoNCE contrast loss to narrow the semantic distance between the audio features and video features of the same sample, while pushing away the semantic distance between the audio features and video features of different samples, thereby realizing the deep semantic alignment of audio features and video features, avoiding the defect that it is difficult to achieve the deep semantic alignment of audio features and video features in the prior art, and effectively improving the accuracy of audio - video emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the implementation process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0019] Refer to Figure 1 , the present invention includes the following steps:

[0020] Step 1) Obtain a training sample set and a test sample set:

[0021] Obtain M audio and video signals including multiple emotion categories, preprocess each audio signal and the video signal with the same emotion category as it, then form a data pair from the logarithmic Mel spectrogram and the standard image frame set corresponding to the preprocessed m-th audio and video signals, and then randomly select more than half of the data pairs and the emotion category labels announced in the file names of each audio or video signal to form a training sample set, and form the remaining data pairs into a test sample set, where M > 1000; in this embodiment, M = 1440;

[0022] Convert each audio signal into a Mel spectrogram through short-time Fourier transform and Mel filter bank, standardize the number of audio frames in the Mel spectrogram of each audio signal to Q by zero-padding or truncation, and take the logarithm of each frequency value in the Mel spectrogram containing Q audio frames after standardization to obtain the logarithmic Mel spectrogram of each audio signal; at the same time, crop each frame of the video signal after frame division, and extract the local binary pattern LBP features of each frame of the image that only retains the facial area after cropping, and then select P cropped image frames with the smallest Pearson correlation with the LBP features of the first frame image, and standardize the P image frames to a fixed size to form the standard image frame set of each video signal; in this embodiment, the short-time Fourier transform parameters are: the sampling rate is 22050Hz, the Fourier window length is 2048, the frame step size is 132, the number of Mel filters is 22, the number of audio frames Q = 600, the number of image frames P = 60, and the fixed size of the image frame is 224×224.

[0023] Step 2) Construct a contrastive learning audio-visual emotion recognition network model based on canonical correlation analysis:

[0024] Construct an audio-visual emotion recognition network model O including a contrastive learning network and an emotion classification module cascaded with it, where the contrastive learning network includes audio processing modules U arranged in parallel and with the same structure a and video processing module U v ,U a and U v both include a cascaded graph construction and enhancement module, a graph neural network module, and a canonical correlation analysis CCA module;

[0025] The graph construction and enhancement module includes a graph structure establishment module and a first graph enhancement module and a second graph enhancement module with the same structure and arranged in parallel cascaded with it;

[0026] The graph neural network module includes two stacked graph convolutional layers and a ReLU activation function loaded between the two graph convolutional layers;

[0027] The emotion classification module includes a stacked feature splicing layer, a fully connected layer, and a softmax activation function.

[0028] Step 3) Iteratively train the audio-visual emotion recognition network model:

[0029] Randomly select N training samples from the training sample set as the input to the audio-visual emotion recognition network model O and perform iterative training on it to obtain the trained audio-visual emotion recognition network model O*; in this embodiment, N = 32.

[0030] (3a) Initialize the iteration number as t, the maximum iteration number as T, T≥100, and the weight parameter of the current audio-visual emotion recognition network model O t is θ t , and let t = 1; in this embodiment, T = 100.

[0031] (3b) In the contrastive learning network, the audio processing module U a and the video processing module U v respectively perform feature encoding on each log Mel spectrogram and the standard image frame set; the emotion classification module predicts the embedding representations z1 and z2 obtained by feature encoding each log Mel spectrogram, and the embedding representations z3 and z4 obtained by feature encoding each standard image frame set to obtain the emotion prediction soft label of each training sample

[0032] The audio processing module U a and the video processing module U v The graph construction and enhancement modules in respectively construct graph structures and graph enhancements for each log Mel spectrogram and the standard image frame set; the graph neural network module respectively processes the enhanced structures G a1 and G a2 , and the enhanced structure G v1 and G v2 of the standard image frame set perform feature extraction; the CCA module normalizes the four groups of extracted embedding features, that is, through the canonical correlation analysis method, projects the four groups of embedding features into a shared feature space respectively, so that in this space, the four groups of projected embedding features have the maximum correlation, realizing the semantic alignment of audio and video features in the shared feature space, thereby enhancing the collaborative expression ability of audio and video embedding features; after being processed by this module, four groups of embedding representations z1, z2, z3, and z4 of each training sample are obtained, and the formula for normalizing the embedding features is:

[0033]

[0034] Among them, h∈[h1, h2, h3, h4] represents the embedded features extracted by the graph neural network module, z∈[z1, z2, z3, z4] represents the normalized embedded representation, std() represents the standard deviation operation, and mean() represents the mean operation;

[0035] In this embodiment, in the graph neural network module that processes the logarithmic Mel-spectrogram, the output dimensions of the two graph convolutional layers are set to 64, and in the graph neural network module that processes the standard image frame set, the output dimensions of the first and second graph convolutional layers are set to 512 and 64 respectively;

[0036] The graph structure building module in the graph building and enhancement module is based on the frame-level graph building method of temporal connection and distance weight decay, and builds graph structures for the logarithmic Mel spectrum graph and the standard image frame set respectively; this module uses the Q audio frames in the logarithmic Mel spectrum graph and the P image frames in the standard image frame set as graph nodes, and q With Y q-1 To Y q-(q-1) The audio frames are connected to form edges, and each edge is assigned a weight W to form a graph structure G of the logarithmic Mel spectrum graph. a , and for each image frame S p With S p-1 To S p-(p-1) The image frames are connected to form edges, and each edge is assigned a weight K to form a graph structure G of a standard image frame set. v ; The first and second image enhancement modules are respectively a , G v After random edge discarding, random feature masking is performed to obtain the first image enhancement module pair G a , G v The enhanced two sets of graph structures G a1 and G v1 , and the second image enhancement module for G a , G v The enhanced two sets of graph structures G a2 and G v2 ; In this embodiment, the edge discarding rate is 0.1 and the feature masking rate is 0.5;

[0037] Graph structure G a The edge weight W is calculated as follows:

[0038] W(Y b ,Y c )=QD(Y b ,Y c )

[0039] Among them, D(Y b ,Y c ) represents audio frame Yb and Y c The distance between them, that is, the absolute value of the frame index difference, b≠c, and 1≤b,c≤Q;

[0040] The graph structure G v The edge weight K of is calculated as follows:

[0041] K(S b′ ,S c′ ) = P - D(S b′ ,S c′ )

[0042] Where D(S b′ ,S c′ ) represents the distance between the image frames S b′ and S c′ , that is, the absolute value of the frame index difference, b′≠c′, and 1≤b′,c′≤P;

[0043] The edge weight decreases as the frame distance increases, thereby enhancing the influence between neighboring frames;

[0044] The feature concatenation layer in the emotion classification module concatenates the four groups of embedded representations z1, z2, z3, and z4; the fully connected layer linearly maps the concatenated feature vector to the emotion category space to obtain the prediction scores corresponding to each emotion category, and the softmax activation function normalizes the prediction scores to obtain the emotion prediction soft label of each training sample

[0045] (3c) Calculate the loss L of the audio-visual emotion recognition network model O through the true label y and the emotion prediction soft label of each training sample and adopt the stochastic gradient descent method to update the weight parameter θ total_loss through L total_loss to obtain the audio-visual emotion recognition network model O of this iteration t ; t ;

[0046] The loss L of the described audio-visual emotion recognition network model O total_loss , the calculation formula is:

[0047] L total_loss = L total_nce + L CE

[0048] L total_nce = αL NCE (z1,z2) + βL NCE (z3,z4) + γL NCE (z i ,z j )

[0049] L NCE (z i ,z j )=[L NCE (z1,z3)+L NCE (z1,z4)+L NCE (z2,z3)+L NCE (z2,z4)]

[0050]

[0051]

[0052] Among them, L total_nce Represents the loss value of the contrastive learning network, L CE Represents the true label y and the sentiment prediction soft label The binary cross entropy loss, L NCE (z1,z2),L NCE (z3,z4),L NCE (z i ,z j ) represent the InfoNCE contrast loss of a batch of log-Mel spectrograms, standard image frame sets, and embedding representations corresponding to log-Mel spectrograms and standard image frame sets, respectively. L NCE (z i ,z j ) indicates that z i and z j is the InfoNCE contrastive loss for a batch of positive pairs, i, j∈[1,2,3,4], denote the embedding representation of the nth and fth samples respectively, and sim(·,·) represents the cosine similarity operation, exp() represents the natural exponential function operation, τ represents the temperature hyperparameter, α, β, and γ represent L NCE (z1,z2),L NCE (z3,z4),L NCE (z i ,z j ) weights. In this embodiment, α=1.0, β=1.0, γ=0.16, τ=0.5;

[0053] The loss of the contrastive learning network in the audio and video emotion recognition network model is calculated through the InfoNCE contrastive loss, which shortens the semantic distance between the audio features and video features of the same sample, and at the same time pushes the semantic distance between the audio features and video features of different samples, thus achieving deep semantic alignment of audio features and video features, and facilitating the efficient fusion of subsequent audio and video features, thereby further improving the accuracy of emotion recognition;

[0054] Update the weight parameter θ t using the following update formula:

[0055] θ t+1 = θ t - η·g

[0056]

[0057] where θ t+1 represents the updated result of θ, η represents the learning rate, t and represents the partial derivative operation of L with respect to θ; in this embodiment, η = 0.001. total_loss with respect to θ t

[0058] (3d) Determine whether t = T holds. If so, obtain the trained audio-visual emotion recognition network model O*, otherwise set t = t + 1, O = O(3d) Determine whether t = T holds. If so, obtain the trained audio-visual emotion recognition network model O*, otherwise set t = t + 1, O = O t , and execute step (3b).

[0059] Step 4) Obtain the audio-visual emotion recognition result:

[0060] Use the test sample set as the input of the trained audio-visual emotion recognition network model O* for forward propagation to obtain the emotion recognition result of each test sample.

Claims

1. A comparative learning audio-visual emotion recognition method based on canonical correlation analysis, characterized in that It includes the following steps: (1) Obtain a training sample set and a test sample set: Obtain M audio and video signals including multiple emotion categories, preprocess each audio signal and the video signal with the same emotion category as it, then form a data pair consisting of the logarithmic Mel spectrogram and the standard image frame set corresponding to the preprocessed m-th audio and video signals, and then randomly select more than half of the data pairs and the emotion category labels announced in the names of each audio or video signal file to form a training sample set, and form the remaining data pairs into a test sample set, where M > 1000; (2) Construct a contrastive learning audio-visual emotion recognition network model based on canonical correlation analysis: Construct an audio-visual emotion recognition network model O including a contrastive learning network and an emotion classification module cascaded with it. Among them, the contrastive learning network includes audio processing modules U arranged in parallel and having the same structure a and video processing modules U v , U a and U v both include a cascaded graph construction and enhancement module, a graph neural network module, and a canonical correlation analysis CCA module; (3) Iteratively train the audio-visual emotion recognition network model: Randomly select N training samples from the training sample set as the input to the audio-visual emotion recognition network model O and perform iterative training on it to obtain the trained audio-visual emotion recognition network model O*; (4) Obtain the audio-visual emotion recognition result: Use the test sample set as the input to the trained audio-visual emotion recognition network model O* for forward propagation to obtain the emotion recognition result of each test sample.

2. The method according to claim 1, wherein The preprocessing of each audio signal and the video signal with the same emotion category as it described in step (1) is realized as follows: Normalize the Mel spectrogram of each audio signal, and transform the normalized Mel spectrogram containing Q audio frames to obtain the logarithmic Mel spectrogram of each audio signal; at the same time, crop each frame image after frame division of each video signal, and extract the local binary pattern LBP features of each frame image that only retains the facial area after cropping, and then select P cropped image frames with the smallest Pearson correlation with the LBP features of the first frame image for normalization processing to form the standard image frame set of each video signal.

3. The method according to claim 1, wherein The audio-visual emotion recognition network model O described in step (2), where: The graph construction and enhancement module includes a graph structure establishment module and a first graph enhancement module and a second graph enhancement module with the same structure and arranged in parallel cascaded with it; The graph neural network module includes two stacked graph convolutional layers and a ReLU activation function loaded between the two graph convolutional layers; The emotion classification module includes a stacked feature splicing layer, a fully connected layer, and a softmax activation function.

4. The method according to claim 3, wherein The random selection of N training samples from the training sample set as the input to the audio-visual emotion recognition network model O and performing iterative training on it described in step (3) is realized as follows: (3a) Initialize the number of iterations as t, the maximum number of iterations as T, where T ≥ 100, and the current weight parameters of the audio-visual emotion recognition network model O t are θ t , and set t = 1; (3b) Audio processing module U in the contrastive learning network a and video processing module U v respectively perform feature encoding on the logarithmic Mel spectrograms and standard image frame sets in each data pair; emotion The classification module makes predictions on the embedding representations z1 and z2 obtained by encoding each log Mel spectrogram feature, as well as the embedding representations z3 and z4 obtained by encoding each standard image frame set feature, to obtain the emotion prediction soft labels for each training sample (3c) Calculate the loss L of the audio-visual emotion recognition network model O through the true label y and the emotion prediction soft label of each training sample Calculate the loss L of the audio-visual emotion recognition network model O total_loss , and adopt the stochastic gradient descent method. Through L total_loss Update the weight parameter θ t To obtain the audio-visual emotion recognition network model O of this iteration t ; (3d) Determine whether t = T holds. If so, obtain the trained audio-visual emotion recognition network model O*. Otherwise, set t = t + 1, O = O t , and execute step (3b).

5. The method according to claim 4, wherein The audio processing module U in the contrastive learning network described in step (3b) a and the video processing module U v respectively perform feature encoding on each log Mel spectrogram and the set of standard image frames, and the implementation steps are as follows: Audio processing module U a and video processing module U v The graph construction and enhancement modules in them respectively construct graph structures and perform graph enhancement on each log Mel spectrogram and standard image frame set; The graph neural network module respectively processes the structures G a1 and G a2 after the enhancement of the log Mel spectrogram, and the structures G v1 and G v2 after the enhancement of the standard image frame set for feature extraction; CCA The module normalizes the four groups of extracted embedding features to obtain four groups of embedding representations z1, z2, z3, and z4 of each training sample.

6. The method according to claim 4, wherein The graph construction and enhancement module described in step (3b) respectively constructs a graph structure and performs graph enhancement on each logarithmic Mel spectrogram and standard image frame set, and the realization steps are as follows: The graph structure building module in the graph construction and enhancement module uses Q audio frames in the logarithmic mel spectrogram and P image frames in the standard image frame set as graph nodes respectively, and for each audio frame Y q connects to the Y q-1 th to the Y q-(q-1) th audio frames to form the graph structure G a of the logarithmic mel spectrogram. At the same time, for each image frame S p connects to the S p-1 th to the S p-(p-1) th image frames to form the graph structure G v of the standard image frame set; the first and second graph enhancement modules respectively perform random edge dropping on G a , G v and then perform random feature masking to obtain two sets of graph structures G a , G v enhanced by the first graph enhancement module, and two sets of graph structures G a1 and G v1 enhanced by the second graph enhancement module, and two sets of graph structures G a , G v enhanced by the second graph enhancement module, and two sets of graph structures G a2 and G v2 .

7. The method according to claim 4, characterized in that The emotion classification module described in step (3b) predicts the embedding representations z1 and z2 obtained by feature encoding of each logarithmic Mel spectrogram, and the embedding representations z3 and z4 obtained by feature encoding of each standard image frame set, and the realization steps are as follows: The feature concatenation layer in the emotion classification module concatenates the four groups of embedded representations z1, z2, z3, and z4; the fully connected layer performs a linear mapping on the concatenated feature vector, and the softmax activation function normalizes the linearly mapped feature vector to obtain the emotion prediction soft label for each training sample 8. The method according to claim 4, characterized in that, The loss L of the audio-visual emotion recognition network model O described in step (3c) total_loss , and the calculation formula is: L total_loss = L total_nce + L CE L total_nce = αL NCE (z1, z2) + βL NCE (z3, z4) + γL NCE (z i , z j ) L NCE (z i ,z j ) = [L NCE (z1,z3) + L NCE (z1,z4) + L NCE (z2,z3) + L NCE (z2,z4)] Among them, L total_nce represents the loss value of the contrastive learning network, and L CE represents the binary cross-entropy loss between the true label y and the soft label of emotion prediction . L NCE (z1,z2), L NCE (z3,z4), L NCE (z i ,z j ) respectively represent the InfoNCE contrastive loss of a batch of the log Mel spectrogram, the set of standard image frames, the log Mel spectrogram, and the embedding representations corresponding to the set of standard image frames. L NCE (z i ,z j ) represents the InfoNCE contrastive loss of a batch with z i and z j as positive sample pairs. i, j ∈ [1,2,3,4], respectively represent the embedding representations of the nth and fth samples, and sim(·,·) represents the cosine similarity operation, exp(·) represents the natural exponential function operation, τ represents the temperature hyperparameter, and α, β, γ respectively represent the weights of L NCE (z1,z2), L NCE (z3,z4), L NCE (z i ,z j ).

9. The method according to claim 4, wherein The update of the weight parameter θ described in step (3c) t is performed according to the following update formula: θ t+1 = θ t - η·g Among them, θ t+1 represents the updated result of θ t , η represents the learning rate, represents L total_loss with respect to θ t partial derivative operation.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and device

    CN117708754A

  • A multimodal emotion recognition method and device

    CN117708754B