Sentiment Category Recognition Method, Device, Equipment and Medium Based on Transformer Architecture

The multimodal data is processed through the encoder and decoder of the Transformer architecture, and the cosine similarity and weight coefficients are used to fuse the eigenvectors, solving the problem of emotional category recognition of modes in multimodal data, improving the reliability and efficiency of recognition.

CN119622492BActive Publication Date: 2025-07-08湖南工商大学
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510162239.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-08
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In multimodal data, the difficulty in identifying emotional categories due to the lack of modality, and the prior art is difficult to effectively deal with the problem of missing modal data.

Method used

Using the Transformer architecture, the normal and missing modal data are feature extracted and reconstructed through encoder and decoder, and the missing modal feature vectors are fused with cosine similarity and weight coefficients, and the vector machine is used for identification.

Benefits of technology

Improve the reliability and efficiency of emotional category identification in the case of missing modal data, and reduce the recognition time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622492B_ABST
    Figure CN119622492B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of artificial intelligence technology and emotion recognition, and discloses an emotion category recognition method, device, equipment and medium based on the Transformer architecture. The method includes: obtaining a current encoded feature vector through an encoder; reconstructing the current encoded feature vector through a decoder to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modality feature vector and a reconstructed vector corresponding to the current normal modality feature vector; determining a first weight coefficient and a second weight coefficient based on the cosine similarity; fusing the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector based on the first weight coefficient and the second weight coefficient to obtain a fused feature vector, and recognizing the fused feature vector to obtain the emotion category of the current multimodal data. The present invention is beneficial to improving the recognition efficiency of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and emotion recognition, and particularly to a method, device, equipment and medium for emotion category recognition based on the Transformer architecture. Background Art

[0002] In the field of emotion recognition, traditional research methods often rely on complete modal datasets, such as video clips that simultaneously contain visual, auditory, and text information. However, in real-world applications, due to environmental noise, sensor failures, and data transmission errors, current multi-modal data often encounters the situation of missing modal data.

[0003] For example, in current multi-modal data, speech may be lost due to background noise or sensor problems, text may be unavailable due to automatic speech recognition errors or unknown vocabulary, and facial expressions may be undetectable due to lighting, motion blur, or occlusion. Since the current missing modal data in current multi-modal data may contain key emotion information, this poses a challenge to emotion category recognition. Therefore, in the case of current missing modal data in current multi-modal data, how to recognize the emotion category of current multi-modal data is an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a method, device, computer equipment and storage medium for emotion category recognition based on the Transformer architecture to solve the technical problem of how to recognize the emotion category of current multi-modal data in the case of current missing modal data in current multi-modal data.

[0005] In a first aspect, a method for emotion category recognition based on the Transformer architecture is provided, including:

[0006] Obtain the current normal modal data and the current missing modal data in the current multi-modal data, and use zero vectors to fill and process the current missing modal data to obtain the processed current missing modal data;

[0007] Extract features from the current normal modal data and the processed current missing modal data through the trained encoder in the Transformer architecture to obtain a current encoded feature vector;

[0008] Reconstruct the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modal feature vector and a reconstructed vector corresponding to the current normal modal feature vector;

[0009] Obtain the cosine similarity between the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector;

[0010] Based on the cosine similarity and a predefined determination method, determine the first weight coefficient and the second weight coefficient;

[0011] Based on the first weight coefficient and the second weight coefficient, fuse the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and through a vector machine, identify the fused feature vector to obtain the emotion category of the current multi-modal data.

[0012] Further, the feature extraction of the current normal modality data and the processed current missing modality data through the trained encoder in the Transformer architecture to obtain a current encoded feature vector includes:

[0013] Through the trained encoder in the Transformer architecture, perform feature extraction on the current normal modality data to obtain a current normal modality feature vector;

[0014] Perform feature extraction on the processed current missing modality data to obtain a current missing modality feature vector, and splice the current normal modality feature vector and the current missing modality feature vector to obtain a current encoded feature vector.

[0015] Further, the reconstruction of the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, and the current decoded feature vector includes the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector, includes:

[0016] Call the input interface, and through the input interface, input the current encoded feature vector into the trained decoder in the Transformer architecture;

[0017] Through the trained decoder in the Transformer architecture, reconstruct the current encoded feature vector to obtain a current decoded feature vector, and the current decoded feature vector includes the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector.

[0018] Further, the obtaining of the cosine similarity between the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector includes:

[0019] Call a preset cosine similarity algorithm;

[0020] By using the cosine similarity algorithm, obtain the cosine similarity between the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector.

[0021] Further, based on the cosine similarity and a predefined determination method, determining the first weight coefficient and the second weight coefficient includes:

[0022] When the cosine similarity is less than a preset similarity, determine the reconstruction vector corresponding to the current missing modality feature vector as the key feature vector;

[0023] Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality, and determine the first weight coefficient and the second weight coefficient according to the energy score of the input modality.

[0024] Further, based on the first weight coefficient and the second weight coefficient, fusing the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and through a vector machine, identifying the fused feature vector to obtain the emotion category of the current multimodal data, includes:

[0025] Using the first weight coefficient and the second weight coefficient, fuse the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector to obtain a fused feature vector;

[0026] Input the fused feature vector into a vector machine, and through the vector machine, identify the fused feature vector to obtain the emotion category of the current multimodal data.

[0027] Further, before obtaining the current normal modality data and the current missing modality data in the current multimodal data and using a zero vector to fill and process the current missing modality data to obtain the processed current missing modality data, the emotion category recognition method includes:

[0028] Obtain the preset normal modality data and the preset missing modality data in the preset multimodal data, and use a zero vector to fill and process the preset missing modality data to obtain the processed preset missing modality data;

[0029] Through the encoder in the Transformer architecture, perform feature extraction on the preset normal modality data and the processed preset missing modality data to obtain a preset encoded feature vector; through the decoder in the Transformer architecture, reconstruct the preset encoded feature vector to obtain a preset decoded feature vector;

[0030] Obtain a first loss value between the preset encoded feature vector and the sample encoded feature vector through a first loss function, and obtain a second loss value between the preset decoded feature vector and the preset encoded feature vector through a second loss function;

[0031] Train the encoder and decoder in the Transformer architecture in a manner that minimizes the first loss value and the second loss value, to obtain the trained encoder and the trained decoder in the Transformer architecture.

[0032] In a second aspect, there is provided an emotion category recognition device based on a Transformer architecture, including:

[0033] A first acquisition module, configured to acquire current normal modality data and current missing modality data in current multimodal data, and use a zero vector to perform padding processing on the current missing modality data to obtain the processed current missing modality data;

[0034] An extraction module, configured to perform feature extraction on the current normal modality data and the processed current missing modality data through the trained encoder in the Transformer architecture to obtain a current encoded feature vector;

[0035] A reconstruction module, configured to reconstruct the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modality feature vector and a reconstructed vector corresponding to the current normal modality feature vector;

[0036] A second acquisition module, configured to obtain the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector;

[0037] A determination module, configured to determine a first weight coefficient and a second weight coefficient based on the cosine similarity and a predefined determination method;

[0038] An identification module, configured to fuse the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector based on the first weight coefficient and the second weight coefficient to obtain a fused feature vector, and perform identification on the fused feature vector through a vector machine to obtain the emotion category of the current multimodal data.

[0039] In a third aspect, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above emotion category recognition method when executing the computer program.

[0040] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned emotion category recognition method are implemented.

[0041] This application provides an emotion category recognition method, device, computer device, and storage medium based on the Transformer architecture. The current normal modality data and the current missing modality data in the current multimodal data are obtained, and the current missing modality data is filled with zero vectors to obtain the processed current missing modality data. The current normal modality data and the processed current missing modality data are subjected to feature extraction through the trained encoder in the Transformer architecture to obtain the current encoded feature vector. The current encoded feature vector is reconstructed through the trained decoder in the Transformer architecture to obtain the current decoded feature vector, and the current decoded feature vector includes the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector. The cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector is obtained. Based on the cosine similarity and a predefined determination method, a first weight coefficient and a second weight coefficient are determined. Based on the first weight coefficient and the second weight coefficient, the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector, and the fused feature vector is recognized through a vector machine to obtain the emotion category of the current multimodal data. The beneficial effects are in two aspects. On the one hand, based on the first weight coefficient and the second weight coefficient, the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector, and the fused feature vector is recognized through a vector machine to obtain the emotion category of the current multimodal data. In the case where the current multimodal data has current missing modality data, the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector. Since there is an effective feature reconstruction mechanism to make up for the current missing modality data, the vector machine can more comprehensively understand the features of the current missing modality data, so it is beneficial to improve the reliability of the recognized emotion category. On the other hand, since the vector machine can automatically recognize the emotion category of the current multimodal data, the recognition time of the emotion category of the current multimodal data is reduced, which is beneficial to improving the recognition efficiency of the emotion category. Description of the Drawings

[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 It is a schematic diagram of an application environment of an emotion category recognition method in an embodiment of the present invention;

[0044] Figure 2 It is a schematic flowchart of an emotion category recognition method provided by an embodiment of the present invention;

[0045] Figure 3 is Figure 2 a schematic flowchart of a specific implementation manner of step S23 in;

[0046] Figure 4 is Figure 2 a schematic flowchart of a specific implementation manner of step S25 in;

[0047] Figure 5 is Figure 2 a schematic flowchart of a specific implementation manner of step S26 in;

[0048] Figure 6 It is a schematic structural diagram of an emotion category recognition device in an embodiment of the present invention;

[0049] Figure 7 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Specific implementation manner

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment of an emotion category recognition method in an embodiment of the present invention. The emotion category recognition method provided by the embodiment of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through the network.

[0052] The server obtains the current normal modal data and the current missing modal data in the current multimodal data through the client, and uses a zero vector to fill and process the current missing modal data to obtain the processed current missing modal data;

[0053] Through the trained encoder in the Transformer architecture, feature extraction is performed on the current normal modal data and the processed current missing modal data to obtain a current encoded feature vector;

[0054] Through the trained decoder in the Transformer architecture, the current encoded feature vector is reconstructed to obtain a current decoded feature vector, and the current decoded feature vector includes a reconstructed vector corresponding to the current missing modal feature vector and a reconstructed vector corresponding to the current normal modal feature vector;

[0055] Obtain the cosine similarity between the reconstructed vector corresponding to the current missing modal feature vector and the reconstructed vector corresponding to the current normal modal feature vector;

[0056] Based on the cosine similarity and a predefined determination method, determine a first weight coefficient and a second weight coefficient;

[0057] Based on the first weight coefficient and the second weight coefficient, fuse the reconstructed vector corresponding to the current missing modal feature vector and the reconstructed vector corresponding to the current normal modal feature vector to obtain a fused feature vector, and through a vector machine, identify the fused feature vector to obtain the emotion category of the current multimodal data.

[0058] In the solution implemented by the above emotion category recognition method, device, equipment and medium, the beneficial effects are in two aspects. On the one hand, based on the first weight coefficient and the second weight coefficient, the reconstructed vector corresponding to the current missing modal feature vector and the reconstructed vector corresponding to the current normal modal feature vector are fused to obtain a fused feature vector, and through a vector machine, the fused feature vector is identified to obtain the emotion category of the current multimodal data. In the case where the current multimodal data has current missing modal data, the reconstructed vector corresponding to the current missing modal feature vector and the reconstructed vector corresponding to the current normal modal feature vector are fused to obtain a fused feature vector. Since there is an effective feature reconstruction mechanism to make up for the current missing modal data, the vector machine can more comprehensively understand the features of the current missing modal data, so it is beneficial to improve the reliability of the recognized emotion category; on the other hand, since the vector machine can automatically identify the emotion category of the current multimodal data, the recognition time of the emotion category of the current multimodal data is reduced, which is beneficial to improving the recognition efficiency of the emotion category.

[0059] Among them, the device running the client is simply referred to as: client device.

[0060] Among them, the device running the server is simply referred to as: server device.

[0061] Among them, client devices include but are not limited to smart phones, personal computers, vehicle networking terminals, tablet computers, and portable wearable devices.

[0062] Among them, the server device can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an emotion category recognition method provided by an embodiment of the present invention, including the following steps:

[0063] S21, obtain the current normal modality data and the current missing modality data in the current multi-modal data, and use a zero vector to fill and process the current missing modality data to obtain the processed current missing modality data;

[0064] Among them, the current multi-modal data is multi-modal data that does not have a complete modality.

[0065] Exemplarily, obtaining the current normal modality data and the current missing modality data in the current multi-modal data, and using a zero vector to fill and process the current missing modality data includes:

[0066] Select the modality data that is not missing in the current multi-modal data as the current normal modality data, and select the modality data that is missing in the current multi-modal data as the current missing modality data;

[0067] Use a zero vector to fill and process the current missing modality data to obtain the processed current missing modality data.

[0068] Among them, the process of selecting the modality data that is not missing in the current multi-modal data as the current normal modality data and selecting the modality data that is missing in the current multi-modal data as the current missing modality data is illustrated by the following example for easy explanation:

[0069] For example, when the modality data that is not missing in the current multi-modal data is the current video data and the current audio data, and the modality data that is missing in the current multi-modal data is the current text data, then select the current video data and the current audio data as the current normal modality data, and select the missing current text data as the current missing modality data;

[0070] For example, when the modality data that is not missing in the current multi-modal data is the current video data and the current text data, and the missing modality data in the current multi-modal data is the current audio data, the current video data and the current text data are selected as the current normal modality data, and the missing current audio data is selected as the current missing modality data;

[0071] For example, when the modality data that is not missing in the current multi-modal data is the current text data and the current audio data, and the missing modality data in the current multi-modal data is the current video data, the current text data and the current audio data are selected as the current normal modality data, and the missing current video data is selected as the current missing modality data.

[0072] The process of using a zero vector to fill the current missing modality data to obtain the processed current missing modality data is illustrated by the following example for ease of explanation:

[0073] For example, the current multi-modal data is a data matrix that contains the current missing modality data. To process the current missing modality data, a zero vector is used to fill the current missing modality data, and the positions of the current missing modality data are all replaced with 0. The purpose of filling is to make the dimensions of the current multi-modal data complete, facilitating subsequent data processing and analysis.

[0074] S22, using the trained encoder in the Transformer architecture, extract features from the current normal modality data and the processed current missing modality data to obtain the current encoded feature vector;

[0075] Among them, the Chinese name of the Transformer architecture is: Transformer architecture.

[0076] Among them, the process of using the trained encoder in the Transformer architecture to extract features from the current normal modality data and the processed current missing modality data to obtain the current encoded feature vector includes:

[0077] Using the trained encoder in the Transformer architecture, extract features from the current normal modality data to obtain the current normal modality feature vector;

[0078] Extract features from the processed current missing modality data to obtain the current missing modality feature vector, and concatenate the current normal modality feature vector and the current missing modality feature vector to obtain the current encoded feature vector.

[0079] Among them, the core of the Transformer architecture lies in its self-attention mechanism and multi-layer perceptron, which enable the model to capture the dependencies between any two positions in the sequence without relying on the sequential processing in traditional recurrent neural networks or convolutional neural networks.

[0080] S23. Through the trained decoder in the Transformer architecture, reconstruct the current encoded feature vector to obtain the current decoded feature vector, where the current decoded feature vector includes the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector;

[0081] S24. Obtain the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector;

[0082] Among them, obtaining the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector includes:

[0083] Call a preset cosine similarity algorithm;

[0084] Through the cosine similarity algorithm, obtain the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector.

[0085] Among them, the cosine similarity algorithm is a mathematical method for calculating the similarity between the reconstructed vectors corresponding to two current missing modality feature vectors and the reconstructed vectors corresponding to the current normal modality feature vectors.

[0086] Among them, if the cosine similarity is larger, it indicates that the reconstructed vector corresponding to the current missing modality feature vector is more similar to the reconstructed vector corresponding to the current normal modality feature vector, indicating that the reconstructed vector corresponding to the current missing modality feature vector does not play a key role in the sentiment category.

[0087] Among them, if the cosine similarity is smaller, it indicates that the reconstructed vector corresponding to the current missing modality feature vector is less similar to the reconstructed vector corresponding to the current normal modality feature vector, indicating that the reconstructed vector corresponding to the current missing modality feature vector plays a key role in the sentiment category, and the first weight coefficient and the second weight coefficient should be adjusted.

[0088] S25. Based on the cosine similarity and a predefined determination method, determine the first weight coefficient and the second weight coefficient;

[0089] Exemplarily, the first weight coefficient is the weight coefficient of the reconstructed vector corresponding to the current missing modality feature vector, and the second weight coefficient is the weight coefficient of the reconstructed vector corresponding to the current normal modality feature vector.

[0090] S26. Based on the first weight coefficient and the second weight coefficient, fuse the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and use a vector machine to identify the fused feature vector to obtain the emotion category of the current multi-modal data.

[0091] Among them, the support vector machine (SVM for short) is a powerful supervised learning algorithm that can effectively process the fused feature vector and find the optimal classification boundary by maximizing the margin.

[0092] Among them, before obtaining the current normal modality data and the current missing modality data in the current multi-modal data and using a zero vector to fill and process the current missing modality data to obtain the processed current missing modality data, the emotion category recognition method includes:

[0093] Obtain the preset normal modality data and the preset missing modality data in the preset multi-modal data, and use a zero vector to fill and process the preset missing modality data to obtain the processed preset missing modality data;

[0094] Extract features from the preset normal modality data and the processed preset missing modality data through the encoder in the Transformer architecture to obtain a preset encoded feature vector;

[0095] Reconstruct the preset encoded feature vector through the decoder in the Transformer architecture to obtain a preset decoded feature vector;

[0096] Obtain the first loss value between the preset encoded feature vector and the sample encoded feature vector through the first loss function, and obtain the second loss value between the preset decoded feature vector and the preset encoded feature vector through the second loss function;

[0097] Train the encoder and decoder in the Transformer architecture in a way that minimizes the first loss value and the second loss value to obtain the trained encoder and the trained decoder in the Transformer architecture.

[0098] Among them, the first loss function is:

[0099] ;

[0100] is the first loss value, represents the divergence loss function, represents the preset encoded feature vector output by the encoder, Represents the sample encoded feature vector output by the pre-trained network.

[0101] Among them, the first loss value is used to describe the difference between the preset encoded feature vector and the sample encoded feature vector. The larger the first loss value, the greater the difference between the preset encoded feature vector and the sample encoded feature vector; the smaller the first loss value, the smaller the difference between the preset encoded feature vector and the sample encoded feature vector.

[0102] Among them, the preset encoded feature vector is the feature vector of the preset normal modality data and the processed preset missing modality data.

[0103] Among them, the preset decoded feature vector is the feature vector after the reconstruction of the preset encoded feature vector.

[0104] Among them, forward represents forward propagation. After the forward propagation ends, the first loss value is obtained.

[0105] Among them, by minimizing the first loss value, it is beneficial to help the encoder learn high-quality feature representations similar to the pre-trained network. Because the pre-trained network has usually been fully trained on a large dataset, the sample encoded feature vector output by the pre-trained network often contains rich and discriminative feature information. By approximating the preset encoded feature vector output by the encoder to the sample encoded feature vector, the encoder can learn the feature extraction ability of the pre-trained network, thereby accelerating the training process of the encoder and improving the accuracy and effectiveness of the feature representation output by the encoder.

[0106] Among them, the second loss function is:

[0107] ;

[0108] is the second loss value, represents the divergence loss function, represents the preset decoded feature vector output by the decoder, represents the preset encoded feature vector output by the encoder.

[0109] Among them, the second loss value is used for the difference between the preset decoded feature vector and the preset encoded feature vector. The larger the second loss value, the greater the difference between the preset decoded feature vector and the preset encoded feature vector; the smaller the second loss value, the smaller the difference between the preset decoded feature vector and the preset encoded feature vector.

[0110] Among them, backward refers to backward propagation. After the backward propagation ends, the second loss value is obtained.

[0111] Among them, by minimizing the second loss value, the decoder can restore the preset encoded feature vector input by the encoder as accurately as possible, which can reduce the loss of the preset encoded feature vector and retain the original structure and features of the preset encoded feature vector.

[0112] Exemplarily, obtaining the preset normal modal data and the preset missing modal data in the preset multimodal data, and using a zero vector to fill the preset missing modal data to obtain the processed preset missing modal data includes:

[0113] Selecting the non-missing modal data in the preset multimodal data as the preset normal modal data, and selecting the missing modal data in the preset multimodal data as the preset missing modal data;

[0114] Using a zero vector to fill the preset missing modal data to obtain the processed preset missing modal data.

[0115] Among them, the process of selecting the non-missing modal data in the preset multimodal data as the preset normal modal data and selecting the missing modal data in the preset multimodal data as the preset missing modal data is exemplified as follows for easy explanation:

[0116] For example, when the non-missing modal data in the preset multimodal data are the preset video data and the preset audio data, and the missing modal data in the preset multimodal data is the preset text data, then select the preset video data and the preset audio data as the preset normal modal data, and select the missing preset text data as the preset missing modal data;

[0117] For example, when the non-missing modal data in the preset multimodal data are the preset video data and the preset text data, and the missing modal data in the preset multimodal data is the preset audio data, then select the preset video data and the preset text data as the preset normal modal data, and select the missing preset audio data as the preset missing modal data;

[0118] For example, when the non-missing modal data in the preset multimodal data are the preset text data and the preset audio data, and the missing modal data in the preset multimodal data is the preset video data, then select the preset text data and the preset audio data as the preset normal modal data, and select the missing preset video data as the preset missing modal data.

[0119] Exemplarily, the encoder in the Transformer architecture is used to extract features from the preset normal modal data and the processed preset missing modal data to obtain a preset encoded feature vector, including:

[0120] Extract features from the preset normal modal data through the encoder in the preset Transformer architecture to obtain a preset normal modal feature vector, extract features from the processed preset missing modal data to obtain a preset missing modal feature vector, and concatenate the preset normal modal feature vector and the preset missing modal feature vector to obtain a preset encoded feature vector.

[0121] Exemplarily, obtaining a first loss value between the preset encoded feature vector and the sample encoded feature vector through a first loss function, and obtaining a second loss value between the preset decoded feature vector and the preset encoded feature vector through a second loss function includes:

[0122] Obtain sample multimodal data corresponding to the preset multimodal data, where the sample multimodal data includes sample video data, sample audio data, and sample text data;

[0123] Extract video features from the sample video data through a pre-trained network to obtain a feature vector of the sample video data, extract audio features from the sample audio data to obtain a feature vector of the sample audio data, and extract text features from the sample text data to obtain a feature vector of the sample text data;

[0124] Concatenate the feature vector of the sample video data, the feature vector of the sample audio data, and the feature vector of the sample text data to obtain a sample encoded feature vector;

[0125] Obtain a first loss value between the preset encoded feature vector and the sample encoded feature vector through a first loss function, and obtain a second loss value between the preset decoded feature vector and the preset encoded feature vector through a second loss function.

[0126] Among them, the sample multimodal data is multimodal data with complete modalities.

[0127] Among them, the sample video data is the video data in the sample multimodal data.

[0128] Among them, the sample audio data is the audio data in the sample multimodal data.

[0129] Among them, the sample text data is the text data in the sample multimodal data.

[0130] In the embodiments of the present invention, the beneficial effects are in two aspects. On the one hand, based on the first weight coefficient and the second weight coefficient, the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector. Through a vector machine, the fused feature vector is recognized to obtain the emotion category of the current multi-modal data. In the case where the current multi-modal data has current missing modality data, the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector. Since there is an effective feature reconstruction mechanism to make up for the current missing modality data, the vector machine can more comprehensively understand the features of the current missing modality data. Therefore, it is beneficial to improve the reliability of the recognized emotion category. On the other hand, since the vector machine can automatically recognize the emotion category of the current multi-modal data, the recognition time of the emotion category of the current multi-modal data is reduced, which is beneficial to improving the recognition efficiency of the emotion category.

[0131] Please refer to Figure 3 , Figure 3 is Figure 2 a schematic flowchart of a specific implementation manner of step S23 in

[0132] S31, call the input interface, and through the input interface, input the current encoded feature vector into the trained decoder in the Transformer architecture;

[0133] S32, through the trained decoder in the Transformer architecture, reconstruct the current encoded feature vector to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstruction vector corresponding to the current missing modality feature vector and a reconstruction vector corresponding to the current normal modality feature vector.

[0134] In the embodiments of the present invention, through the trained decoder in the Transformer architecture, the current missing modality data is inferred to obtain a reconstruction vector corresponding to the current missing modality feature vector. Through the reconstruction vector corresponding to the current missing modality feature vector, the comprehensiveness of the current decoded feature vector is improved.

[0135] Please refer to Figure 4 , Figure 4 is Figure 2 a schematic flowchart of a specific implementation manner of step S25 in

[0136] S41, when the cosine similarity is less than the preset similarity, determine the reconstruction vector corresponding to the current missing modality feature vector as the key feature vector;

[0137] S42. Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality, and determine the first weight coefficient and the second weight coefficient according to the energy score of the input modality.

[0138] Among them, obtaining the input modality corresponding to the key feature vector, obtaining the energy score of the input modality, and determining the first weight coefficient and the second weight coefficient according to the energy score of the input modality includes:

[0139] Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality through an energy score generation model, select the weight coefficient corresponding to the energy score of the input modality as the first weight coefficient, and select the difference obtained by subtracting the first weight coefficient from 1 as the second weight coefficient.

[0140] Among them, the energy score generation model:

[0141] ;

[0142] Among them, is the energy score of the m-th input modality, is the m-th input modality, m is the serial number of the input modality, Energy is the energy function, is the temperature parameter, log is the abbreviation of logarithm, ∑ represents the summation operation, the summation range is from k to G, k is the serial number of the label, G is the total number of labels, e represents the exponential function, and SVM is a single-modal classifier, is the output logical value of the single-modal classifier of the m-th input modality corresponding to the k-th class label;

[0143] Among them, represents the result of taking the natural exponential function after dividing the output logical value of the single-modal classifier of the m-th input modality corresponding to the k-th class label by the temperature parameter.

[0144] Among them, through , the output logical value of the single-modal classifier of the m-th input modality corresponding to the k-th class label can be converted into a parameter that is easy to process.

[0145] Among them, the weight coefficients corresponding to the energy scores of different input modalities are different.

[0146] For the convenience of explanation, the following is an example:

[0147] For example, if the energy score of the input modality is in the first range, the coefficient 1 is used as the first weight coefficient, and if the energy score of the input modality is in the second range, the coefficient 2 is used as the first weight coefficient. Among them, the coefficient 1 is greater than the coefficient 2.

[0148] Among them, the input modality is evaluated by an energy - fraction - based method to achieve the effect of dynamically adjusting the first weight coefficient. The higher the influence of the input modality on the prediction result, the higher the energy fraction. For multimodal data, the density functions of different modalities can be estimated through corresponding energy functions.

[0149] Among them, the input modality refers to the type used when data is input into a single - modality classifier.

[0150] Among them, the input modalities include video modality, audio modality, and text modality. Video modality, audio modality, and text modality are three basic input forms in multimedia data processing.

[0151] Among them, the video modality refers to the data form containing a dynamic image sequence. The dynamic image sequence is arranged in a certain time order and together constitutes a visual scene that changes over time. The video modality not only contains static image information but also contains dynamic information in the time dimension.

[0152] Among them, the audio modality refers to the data form containing sound signals. The audio modality contains frequency characteristics, amplitude characteristics, and timbre characteristics of sound. The frequency characteristics, amplitude characteristics, and timbre characteristics of sound are of great significance for sound recognition, classification, and analysis.

[0153] Among them, the text modality refers to the data form containing characters, words, and sentences. The text modality is the basic data type in the field of natural language processing.

[0154] Among them, the process of obtaining the input modality corresponding to the key feature vector is described in detail as follows for the convenience of explanation:

[0155] For example, when the missing modality data in the current multimodal data is the current video data, the video modality is selected as the input modality corresponding to the key feature vector.

[0156] For example, when the missing modality data in the current multimodal data is the current audio data, the audio modality is selected as the input modality corresponding to the key feature vector.

[0157] For example, when the missing modality data in the current multimodal data is the current text data, the text modality is selected as the input modality corresponding to the key feature vector.

[0158] In the embodiments of the present invention, the input modality corresponding to the key feature vector is obtained, the energy fraction of the input modality is obtained, and the first weight coefficient and the second weight coefficient are determined according to the energy fraction of the input modality. In this way, the first weight coefficient can be dynamically adjusted. This dynamic adjustment strategy enables the trained decoder to flexibly adjust according to the importance of the reconstruction vector corresponding to the current missing - modality feature vector, so as to achieve more accurate emotion recognition.

[0159] Please refer to Figure 5 , Figure 5 which Figure 2 is a schematic flowchart of a specific implementation manner of step S26 in

[0160] S51, using the first weight coefficient and the second weight coefficient, fuse the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector to obtain a fused feature vector;

[0161] S52, input the fused feature vector into a vector machine, and through the vector machine, identify the fused feature vector to obtain the emotion category of the current multimodal data.

[0162] In the embodiment of the present invention, in the case where there is current missing modality data in the current multimodal data, the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector. Since there is an effective feature reconstruction mechanism to make up for the current missing modality data, the vector machine can more comprehensively understand the features of the current missing modality data, so it is beneficial to improve the reliability of the identified emotion category.

[0163] Please refer to Figure 6 , Figure 6 which Figure 6 is a schematic structural diagram of an emotion category recognition device in an embodiment of the present invention. As

[0164] shown, the emotion category recognition device includes a first acquisition module 101, an extraction module 102, a reconstruction module 103, a second acquisition module 104, a determination module 105, and an identification module 106. The detailed descriptions of each functional module are as follows:

[0165] The first acquisition module 101 is configured to acquire the current normal modality data and the current missing modality data in the current multimodal data, and use a zero vector to fill and process the current missing modality data to obtain the processed current missing modality data;

[0166] The extraction module 102 is configured to extract features from the current normal modality data and the processed current missing modality data through a trained encoder in the Transformer architecture to obtain a current encoded feature vector;

[0167] The second acquisition module 104 is configured to acquire the cosine similarity between the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector;

[0168] The determination module 105 is configured to determine a first weight coefficient and a second weight coefficient based on the cosine similarity and a predefined determination method;

[0169] The recognition module 106 is configured to fuse the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector based on the first weight coefficient and the second weight coefficient to obtain a fused feature vector, and identify the fused feature vector through a vector machine to obtain the emotion category of the current multimodal data.

[0170] In the embodiment of the present invention, the beneficial effects are in two aspects. On the one hand, based on the first weight coefficient and the second weight coefficient, the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector, and the fused feature vector is identified through a vector machine to obtain the emotion category of the current multimodal data. When there is current missing modality data in the current multimodal data, the reconstruction vector corresponding to the current missing modality feature vector and the reconstruction vector corresponding to the current normal modality feature vector are fused to obtain a fused feature vector. Since there is an effective feature reconstruction mechanism to make up for the current missing modality data, the vector machine can more comprehensively understand the features of the current missing modality data, so it is beneficial to improve the reliability of the identified emotion category. On the other hand, since the vector machine can automatically identify the emotion category of the current multimodal data, the recognition time of the emotion category of the current multimodal data is reduced, which is beneficial to improving the recognition efficiency of the emotion category.

[0171] For the specific limitations of the emotion category recognition device, reference can be made to the limitations of the emotion category recognition method in the above text, which will not be elaborated here.

[0172] Each module in the above emotion category recognition device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0173] Please refer to Figure 7 , Figure 7 which is another structural schematic diagram of a computer device in an embodiment of the present invention. In one embodiment, a computer device is provided. The computer device is a server device or a client device, and its internal structural diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external database. When the computer program is executed by the processor, it can implement the functions or steps of a sentiment category recognition method based on the Transformer architecture.

[0174] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor.

[0175] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions of the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0176] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU for short), a graphics processing unit (GPU for short), and a network processor (NP for short); it can also be a digital signal processor (DSP for short), an application-specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0177] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. In this article, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the methods and products disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts can refer to the description of the method part.

[0178] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0179] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A method for sentiment category recognition based on the Transformer architecture, characterized in that Including: Select the modality data that is not missing in the preset multi-modal data as the preset normal modality data, and select the modality data that is missing in the preset multi-modal data as the preset missing modality data. Use a zero vector to fill and process the preset missing modality data to obtain the processed preset missing modality data. The preset missing modality data is one of preset text data, preset audio data, and preset video data; Extract features from the current normal modality data and the processed current missing modality data through the trained encoder in the Transformer architecture to obtain the current encoded feature vector; Reconstruct the current encoded feature vector through the trained decoder in the Transformer architecture to obtain the current decoded feature vector. The current decoded feature vector includes the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector; Obtain the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector; When the cosine similarity is less than the preset similarity, determine the reconstructed vector corresponding to the current missing modality feature vector as the key feature vector; Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality, and determine the first weight coefficient and the second weight coefficient according to the energy score of the input modality; Based on the first weight coefficient and the second weight coefficient, fuse the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and identify the fused feature vector through a vector machine to obtain the emotion category of the current multi-modal data; Among them, obtaining the input modality corresponding to the key feature vector, obtaining the energy score of the input modality, and determining the first weight coefficient and the second weight coefficient according to the energy score of the input modality includes: Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality through an energy score generation model, select the weight coefficient corresponding to the energy score of the input modality as the first weight coefficient, and select the difference obtained by subtracting the first weight coefficient from 1 as the second weight coefficient; Among them, the energy score generation model: where Energy(x m ) is the energy score of the m-th input modality, x m is the m-th input modality, m is the sequence number of the input modality, Energy is the energy function, T m is the temperature parameter, log is the abbreviation of logarithm, ∑ represents the summation operation, the summation range is from k to G, k is the sequence number of the label, G is the total number of labels, e represents the exponential function, SVM is the single-modal classifier, is the output logical value corresponding to the k-th class label of the single-modal classifier of the m-th input modality.

2. The emotional category recognition method according to claim 1, characterized in that, The step of extracting features from the current normal modality data and the processed current missing modality data through the trained encoder in the Transformer architecture to obtain the current encoded feature vector includes: Extract features from the current normal modality data through the trained encoder in the Transformer architecture to obtain the current normal modality feature vector; Extract features from the processed current missing modality data to obtain the current missing modality feature vector, and splice the current normal modality feature vector and the current missing modality feature vector to obtain the current encoded feature vector.

3. The emotional category recognition method according to claim 1, characterized in that, Reconstructing the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modality feature vector and a reconstructed vector corresponding to the current normal modality feature vector, including: Invoking an input interface, and inputting the current encoded feature vector into the trained decoder in the Transformer architecture through the input interface; Reconstructing the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modality feature vector and a reconstructed vector corresponding to the current normal modality feature vector.

4. The emotional category recognition method according to claim 1, wherein Obtaining the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector, including: Invoking a preset cosine similarity algorithm; Obtaining the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector through the cosine similarity algorithm.

5. The emotional category recognition method according to claim 1, wherein Based on the first weight coefficient and the second weight coefficient, fusing the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and identifying the sentiment category of the current multimodal data through a vector machine, including: Using the first weight coefficient and the second weight coefficient to fuse the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector to obtain a fused feature vector; Inputting the fused feature vector into a vector machine, and identifying the fused feature vector through the vector machine to obtain the sentiment category of the current multimodal data.

6. The emotional category recognition method according to claim 1, wherein Before selecting the modality data that is not missing in the preset multimodal data as the preset normal modality data, selecting the modality data that is missing in the preset multimodal data as the preset missing modality data, and using a zero vector to fill and process the preset missing modality data to obtain the processed preset missing modality data, where the preset missing modality data is one of the preset text data, preset audio data, and preset video data, the sentiment category recognition method includes: Obtaining the preset normal modality data and the preset missing modality data in the preset multimodal data, and using a zero vector to fill and process the preset missing modality data to obtain the processed preset missing modality data; Extracting features from the preset normal modality data and the processed preset missing modality data through the encoder in the Transformer architecture to obtain a preset encoded feature vector; Reconstructing the preset encoded feature vector through the decoder in the Transformer architecture to obtain a preset decoded feature vector; Obtain a first loss value between the preset encoded feature vector and the sample encoded feature vector through a first loss function, and obtain a second loss value between the preset decoded feature vector and the preset encoded feature vector through a second loss function; Train the encoder and decoder in the Transformer architecture in a manner that minimizes the first loss value and the second loss value, to obtain the trained encoder and the trained decoder in the Transformer architecture.

7. An emotional category recognition device based on the Transformer architecture, characterized in that, Comprising: A first acquisition module, configured to select the modality data without missing values in the preset multimodal data as the preset normal modality data, select the modality data with missing values in the preset multimodal data as the preset missing modality data, and use a zero vector to perform padding processing on the preset missing modality data to obtain the processed preset missing modality data, where the preset missing modality data is one of preset text data, preset audio data, and preset video data; An extraction module, configured to perform feature extraction on the current normal modality data and the processed current missing modality data through the trained encoder in the Transformer architecture to obtain a current encoded feature vector; A reconstruction module, configured to reconstruct the current encoded feature vector through the trained decoder in the Transformer architecture to obtain a current decoded feature vector, where the current decoded feature vector includes a reconstructed vector corresponding to the current missing modality feature vector and a reconstructed vector corresponding to the current normal modality feature vector; A second acquisition module, configured to obtain the cosine similarity between the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector; A determination module, configured to, when the cosine similarity is less than a preset similarity, determine the reconstructed vector corresponding to the current missing modality feature vector as a key feature vector; obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality, and determine a first weight coefficient and a second weight coefficient according to the energy score of the input modality; An identification module, configured to, based on the first weight coefficient and the second weight coefficient, fuse the reconstructed vector corresponding to the current missing modality feature vector and the reconstructed vector corresponding to the current normal modality feature vector to obtain a fused feature vector, and identify the fused feature vector through a vector machine to obtain the emotion category of the current multimodal data; wherein, obtaining the input modality corresponding to the key feature vector, obtaining the energy score of the input modality, and determining a first weight coefficient and a second weight coefficient according to the energy score of the input modality includes: Obtain the input modality corresponding to the key feature vector, obtain the energy score of the input modality through an energy score generation model, select the weight coefficient corresponding to the energy score of the input modality as the first weight coefficient, and select the difference obtained by subtracting the first weight coefficient from 1 as the second weight coefficient; Wherein, the energy score generation model: Among them, Energy(x m ) is the energy score of the m-th input modality, x m is the m-th input modality, m is the serial number of the input modality, Energy is the energy function, T m is the temperature parameter, log is the abbreviation of logarithm, ∑ represents the summation operation, the summation range is from k to G, k is the serial number of the label, G is the total number of labels, e represents the exponential function, SVM is the single-modal classifier, is the output logical value corresponding to the k-th class label of the single-modal classifier of the m-th input modality.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the emotion category recognition method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the emotion category recognition method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method and system for uncertain modal missing

    CN115983280A

  • Multi-modal emotion recognition method and system for modal missing scene

    CN116933051A

  • Balanced multi-modal video analysis method and system based on cross-modal reconstruction

    CN117671559A