Emotion recognition method, device, equipment and storage medium

By combining video frames and emotion labels to generate text data and similarity data in the emotion recognition model, the problems of low accuracy and lack of migration ability in dynamic faces in the prior art are solved, and higher recognition accuracy and model universality are achieved.

CN115050077BActive Publication Date: 2025-05-16LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210760941.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-05-16
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing dynamic face emotion recognition methods have low recognition accuracy and lack the transfer ability of zero-sample learning.

Method used

By obtaining the video and audio to be tested, determining the video frames and combining them with the emotion tags to generate text data, inputting the emotion recognition model to generate similarity data, and determining the emotion tag corresponding to the maximum similarity is used as the recognition result.

Benefits of technology

It improves the accuracy of emotion recognition and gives the model a certain zero-sample learning ability, enhancing the universality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050077B_ABST
    Figure CN115050077B_ABST
Patent Text Reader

Abstract

The present application discloses an emotion recognition method, device, equipment and storage medium, which are applied to the field of neural network technology. The emotion recognition model training method comprises: obtaining a video to be tested and an audio to be tested; determining a plurality of video frames to be tested in the video to be tested, and using each emotion label in a label set to respectively splice with a text template to be tested to generate text data to be tested corresponding to each emotion label; inputting the video frame to be tested, the text data to be tested and the audio to be tested into an emotion recognition model to obtain non-text encoded data to be tested and each text encoded data to be tested corresponding to each text data to be tested; using the non-text encoded data to be tested and each text encoded data to be tested to generate similarity data to be tested; determining the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested; the method introduces the semantic information contained in the label itself to improve the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of neural network technology, and in particular to an emotion recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] With the maturity of current face recognition technology, it is a relatively mature technology to find the face of the focus person from a picture or video. Therefore, the current research on emotion recognition focuses on the research on facial emotion recognition. Researchers usually divide facial emotion recognition into static facial emotion recognition and dynamic facial emotion recognition. The former identifies people's emotions through a single face picture, and the latter identifies people's emotions through dynamic images or videos. Since facial emotion recognition is a dynamic process, it is sometimes difficult to define the current person's true emotions based on just one picture. However, the recognition accuracy of the current dynamic facial emotion recognition method is poor, and it does not have the transfer ability of zero-sample learning. Summary of the invention

[0003] In view of this, the purpose of the present application is to provide an emotion recognition method, device, electronic device and computer-readable storage medium to improve the accuracy of emotion recognition and the versatility of the model.

[0004] To solve the above technical problems, the present application provides an emotion recognition model training method, comprising:

[0005] Get the video and audio to be tested;

[0006] Determine a plurality of video frames to be tested in the video to be tested, and use each emotion label in the label set to respectively splice with the text template to be tested to generate text data to be tested corresponding to each emotion label;

[0007] Inputting the video frame to be tested, the text data to be tested and the audio to be tested into an emotion recognition model to obtain the non-text coded data to be tested and the text coded data to be tested corresponding to each of the text data to be tested;

[0008] Generate similarity data to be tested using the non-text encoded data to be tested and each of the text encoded data to be tested;

[0009] The emotion label corresponding to the maximum similarity data to be tested is determined as the emotion recognition result corresponding to the video to be tested.

[0010] Optionally, the step of splicing each emotion tag in the tag set with the text template to be tested to generate the text data to be tested corresponding to each emotion tag includes:

[0011] Selecting the text template to be tested from a preset template library;

[0012] Performing vector mapping processing on the text template to be tested and each of the emotion labels respectively to obtain a template vector to be tested and each label vector;

[0013] The template vector is concatenated with each of the label vectors to obtain the text data to be tested.

[0014] Optionally, the training process of the emotion recognition model includes:

[0015] Get training videos, training audios, and emotion labels;

[0016] Determining a plurality of training video frames in the training video, and generating training text data using the emotion tags;

[0017] Input the training video frame, the training text data and the training audio into an initial model to obtain training text encoding data and training non-text encoding data;

[0018] Generate similarity data using the training text encoding data and the training non-text encoding data;

[0019] Generating a loss value using the similarity data, and adjusting parameters of the initial model based on the loss value;

[0020] If it is detected that the training completion condition is met, the initial model after parameter adjustment is determined as the emotion recognition model.

[0021] Optionally, the initial model includes a text encoder, an image encoder and an audio encoder, and also includes a pooling network module and a time recursive network module, the output of the text encoder is the input of the pooling network module, and the output of the image encoder is the input of the time recursive network module.

[0022] Optionally, the step of inputting the training video frame, the training text data and the training audio into an initial model to obtain training text encoding data and training non-text encoding data includes:

[0023] Inputting the training text into the text encoder to obtain a plurality of initial text encodings;

[0024] Inputting the multiple initial text encodings into the pooling network module to obtain the training text encoding data;

[0025] Input the training video frame into the image encoder to obtain a plurality of initial image codes, and input the training audio into the audio encoder to obtain initial audio codes;

[0026] Inputting the multiple initial image codes into the time recursive network module to obtain intermediate image codes;

[0027] The intermediate image code and the initial audio code are concatenated to obtain the training non-text code data.

[0028] Optionally, the text encoder and the image encoder belong to a language-image contrast learning pre-trained model, and the audio encoder is pre-trained.

[0029] Optionally, adjusting parameters of the initial model based on the loss value includes:

[0030] Parameters of the pooling network module and the time recursive network module in the initial model are adjusted based on the loss value.

[0031] Optionally, the generating training text data using the emotion tag includes:

[0032] Select a target text template from the preset template library;

[0033] Performing vector mapping processing on the target text template and the emotion label to obtain a template vector and a label vector;

[0034] The template vector and the label vector are concatenated to obtain the training text data.

[0035] Optionally, the detecting that a training completion condition is met includes:

[0036] Using the test data, the accuracy of the initial model after parameter adjustment is tested to obtain a test result;

[0037] If the test result is greater than a preset threshold, it is determined that the training completion condition is met.

[0038] Optionally, the test data includes multiple groups of test sub-data, including target test sub-data, and the target test sub-data includes target test video, target test audio and target test label.

[0039] Optionally, the using of the test data to perform an accuracy test on the initial model after parameter adjustment to obtain a test result includes:

[0040] Determine a plurality of target test video frames in the target test video, and generate a plurality of target test text data using each emotion tag in the tag set; wherein the target test text data corresponds to at least one text template;

[0041] The target test video frame, the target test text data and the initial model after adjusting the target test audio input parameters are obtained to obtain target non-text encoding data and multiple target text encoding data;

[0042] Calculating test similarity data between the target non-text encoded data and each target text encoded data, and using the test similarity data to determine at least one maximum similarity data corresponding to the at least one text template;

[0043] Determine the emotion label corresponding to the at least one maximum similarity data as the initial prediction result corresponding to the target test video, and perform maximum number screening on the initial prediction results to obtain a prediction result;

[0044] Determine a test sub-result corresponding to the target test sub-data based on the prediction result and the target test label;

[0045] All test sub-results corresponding to the test data are counted to obtain the test result.

[0046] Optionally, the detecting that a training completion condition is met includes:

[0047] When it is detected that the training duration reaches a preset duration limit, determining that the training completion condition is met;

[0048] Or when it is detected that the number of training rounds reaches a preset number of training rounds, it is determined that the training completion condition is met.

[0049] The present application also provides an emotion recognition device, comprising:

[0050] A test acquisition module is used to acquire the video and audio to be tested;

[0051] The test data processing module is used to determine a plurality of test video frames in the test video, and use each emotion label in the label set to splice with the test text template to generate the test text data corresponding to each emotion label;

[0052] The test input module is used to input the test video frame, the test text data and the test audio into the emotion recognition model to obtain the test non-text code data and the test text code data corresponding to each test text data;

[0053] A similarity generation module to be tested, used for generating similarity data to be tested by using the non-text encoded data to be tested and each of the text encoded data to be tested;

[0054] The recognition result determination module is used to determine the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested.

[0055] The present application also provides an electronic device, including a memory and a processor, wherein:

[0056] The memory is used to store the computer program;

[0057] The processor is used to execute the computer program to implement the above-mentioned emotion recognition method.

[0058] The present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the above-mentioned emotion recognition method.

[0059] The emotion recognition model training method provided in the present application obtains a video to be tested and an audio to be tested; determines a plurality of video frames to be tested in the video to be tested, and uses each emotion label in the label set to respectively splice with a text template to be tested to generate text data to be tested corresponding to each emotion label; inputs the video frame to be tested, the text data to be tested and the audio to be tested into the emotion recognition model to obtain non-text encoded data to be tested and each text encoded data to be tested corresponding to each text data to be tested; uses the non-text encoded data to be tested and each text encoded data to be tested to generate similarity data to be tested; and determines the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested.

[0060] It can be seen that this method transforms the emotion recognition process from the original probability prediction problem to a similarity matching problem, and at the same time introduces the semantic information contained in the label itself, which improves the accuracy while enabling the model to have a certain zero-shot learning migration capability. Specifically, when identifying emotions, the present application uses various emotion labels and the same text template to generate multiple text data to be tested. The emotion recognition model is trained to learn the semantic information carried by the emotion label, and the similarity between the non-text encoded data to be tested of the video to be tested and the text encoded data to be tested corresponding to each emotion label is generated to select the maximum similarity data to be tested and determine the most similar emotion label, thereby improving the accuracy of emotion recognition. At the same time, even if an emotion label that is not involved in the training of the emotion recognition model is added during application, the emotion recognition model can distinguish it from other emotion labels based on the semantic information of the emotion label, and has a certain zero-sample learning capability, which improves the versatility of the model.

[0061] In addition, the present application also provides a device, an electronic device and a computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0063] Figure 1 A flow chart of an emotion recognition model training method provided in an embodiment of the present application;

[0064] Figure 2 A flow chart of an emotion recognition method provided in an embodiment of the present application;

[0065] Figure 3 A specific data processing flow chart provided for an embodiment of the present application;

[0066] Figure 4 A schematic diagram of the structure of an identification terminal provided by an embodiment of the present invention;

[0067] Figure 5 A schematic diagram of the structure of an emotion recognition device provided in an embodiment of the present application;

[0068] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0070] At present, the main solution for dynamic facial emotion recognition is to use the multimodal fusion information of vision and sound to realize emotion recognition. That is, the visual image and sound audio in the video are extracted by feature extractors respectively, and then fused using a feature fusion network, and finally a set of fixed pre-defined emotion categories are predicted. However, this solution completely ignores the semantic information contained in the emotion label itself, but directly maps the emotion label to a fixed number of category indexes (numbers). This solution not only limits the versatility of the model, but also does not have the migration / prediction capabilities of zero-shot learning. It requires additional training data to migrate the model application to new scenarios, and also leads to low accuracy of emotion recognition.

[0071] In this application, we draw on the way humans recognize emotions. When people see a video, they can associate and correspond the features of the images in the video (whether they have seen it before or not) with the features of the natural language in their minds, rather than with numbers / indexes. Therefore, this application uses an unconventional training method to mine the semantic information of the label text during training and associate it with the corresponding video features, which not only enhances the semantics of the video representation and improves the recognition accuracy, but also enables the model to have a certain zero-shot learning migration capability.

[0072] For details, please refer to Figure 1 , Figure 1 A flow chart of an emotion recognition model training method provided in an embodiment of the present application. The method includes:

[0073] S101: Obtain training videos, training audios, and emotion labels.

[0074] It should be noted that each step in the present application can be completed by a designated electronic device. The executing electronic device can be in any form such as a server, a computer, etc. The number of electronic devices can be one or more, that is, all steps can be executed by one electronic device, or multiple electronic devices can execute part of the steps respectively, and work together to complete the model training and / or emotion recognition process.

[0075] The training video, training audio, and emotion labels correspond to each other. The training video refers to a video recording the emotional changes of a person's face, and the training audio refers to the audio corresponding to the training video, which usually records the sounds corresponding to the emotional changes of the person's face recorded in the training video, such as crying, laughing, etc. The emotion label refers to the text name corresponding to the emotion expressed by the training video and training audio, such as happy, angry, sad, fear, etc.

[0076] S102: Determine a plurality of training video frames in a training video, and generate training text data using emotion tags.

[0077] The training video frame can be any video frame in the training video, and the number of training video frames can be multiple, for example, M, where M is a fixed positive number. Using multiple training video frames, the emotional changes of the face in the training video can be represented in the time sequence direction. The method for determining the training video frame is not limited. In one embodiment, the training video frame can be extracted from the first frame of the training video at a preset time interval; in another embodiment, the number of training video frames can be determined, and based on the number, the training video is averagely spaced to obtain the training video frame.

[0078] Training text data refers to data used to represent the semantic information of emotion tags, and its specific form is not limited, for example, it can be in text form, or it can be in vector form. In one embodiment, the emotion tag can be directly used as the training text data, or the emotion tag can be mapped from text to vector to obtain a corresponding tag vector, and the tag vector is determined as the training text data. In another embodiment, a preset text template (prompt) can be obtained, and the text template and the emotion tag are used to jointly generate training text data to further provide more semantic information. The specific content of the text template is not limited, for example, it can be "The person seems to express the feeling of the [CLASS]", "From this video, we can see that the person is [CLASS]", where the [CLASS] position is used to insert the emotion tag.

[0079] In another embodiment, since different prompt sentence patterns may cause the semantic information learned by the model to be different, in order to avoid the influence of the text template on the model training effect, multiple text templates can be preset to form a preset template library. When generating training text data, a target text template can be selected from the preset template library, which can be selected randomly or in sequence. The target text template and the emotion label are respectively vector mapped to obtain a template vector and a label vector, and the specific vector mapping method is not limited. After the mapping is completed, the template vector and the label vector are spliced ​​to obtain the training text data. This method enables the model to adapt to various prompt sentence patterns.

[0080] S103: Input the training video frames, training text data and training audio into the initial model to obtain training text encoding data and training non-text encoding data.

[0081] After obtaining the training video frames and training text data, they are input into the initial model together with the training audio, and the initial model encodes them to obtain the training text encoding data representing text features and the training non-text encoding data representing non-text features. The training text encoding data is obtained based on the training text data, which can represent the emotional semantic features of the emotional label. The non-text features are obtained based on the training video frames and the training audio, which can represent the emotional features represented by the image and sound.

[0082] The initial model refers to an emotion recognition model that has not been trained yet. After iterative training and parameter adjustment, the feature extraction capability is improved and then transformed into an emotion recognition model. The specific type of the initial model is not limited, and any feasible neural network architecture can be used. In a feasible implementation, the initial model includes a text encoder, an image encoder, and an audio encoder. The text encoder is used to process training text data to obtain training text encoding data. The image encoder and the audio encoder are used to process training video frames and training audio, respectively, and the two cooperate to obtain training non-text encoding data. In another implementation, in order to extract timing information and thus improve recognition accuracy, a pooling network module and a time recursive network module can also be included in the initial model. Among them, the output of the text encoder is the input of the pooling network module, and the output of the image encoder is the input of the time recursive network module. The time recursive network module can specifically be an LSTM (Long Short-Term Memory, long short-term memory network) network, and the pooling network module is specifically used to perform a time-series pooling operation on the output of the text encoder.

[0083] This embodiment does not limit the way in which the initial model obtains training text encoding data and training non-text encoding data, and the specific generation method is related to the model structure of the initial model. In one embodiment, if the initial model is the above-mentioned structure including a text encoder, an image encoder, an audio encoder, a pooling network module and a time recursive network module, the training text can be input into the text encoder to obtain multiple initial text encodings, and the number of initial text encodings is the same as the number of training video frames. Then, multiple initial text encodings are input into the pooling network module to obtain training text encoding data. In addition, the training video frame can be input into the image encoder to obtain multiple initial image encodings, and the training audio can be input into the audio encoder to obtain the initial audio encoding, and then multiple initial image encodings can be input into the time recursive network module to obtain intermediate image encodings, and finally the intermediate image encodings and the initial audio encodings are spliced ​​to obtain training non-text encoding data. The specific method of splicing is not limited, and the initial audio encoding can be in front, or the intermediate image encoding can be in front.

[0084] S104: Generate similarity data using the training text encoding data and the training non-text encoding data.

[0085] S105: Generate a loss value using the similarity data, and adjust parameters of the initial model based on the loss value.

[0086] For the sake of convenience, steps S104 and S105 are combined for description.

[0087] This application transforms the emotion recognition process from the original probability prediction problem to a similarity matching problem. Therefore, during training, similarity data is generated by using training text encoding data and training non-text encoding data, and the similarity data is used to characterize the gap between the training text encoding data and the training non-text encoding data. Since the emotion label and the training video and training audio represent the same emotion, the gap can characterize the defect of the initial model in feature extraction, that is, the loss value, and then the parameters of the initial model can be adjusted based on the loss value, so that the initial model learns how to accurately extract text-type emotion features and non-text-type emotion features.

[0088] The calculation method of similarity data can be set as needed. For example, in one embodiment, both the training text encoding data and the training non-text encoding data are vector images. In this case, cosine similarity can be calculated as similarity data. The specific type of loss value is also not limited, for example, it can be a cross entropy loss value.

[0089] When adjusting parameters, the entire initial model can be adjusted as needed, or some of them can be adjusted. For example, in one embodiment, if the initial model is the above-mentioned structure including a text encoder, an image encoder, an audio encoder, a pooling network module and a time recursive network module, the text encoder and the image encoder can belong to a language image contrast learning pre-training model, and the audio encoder is also pre-trained. At this time, when adjusting parameters, the pooling network module and the time recursive network module in the initial model can be adjusted based on the loss value. The language image contrast learning pre-training model is the CLIP (Contrastive Language-Image Pre-Training) model. After large-scale pre-training, it already has better model parameters and does not need to continue to adjust parameters. The audio encoder (or sound encoder) can use the YAMNET model, which is an audio event classifier trained on the AudioSet data set. The YAMNET overall network architecture uses MobileNet v1, and the feature dimension of the extracted sound is 1024 dimensions.

[0090] After the parameter adjustment is completed, it can be detected whether the training completion condition is met. The detection can be performed periodically, for example, once after completing several rounds of iterative training. If not, continue to execute step S101 and continue training, otherwise execute step S106.

[0091] S106: If it is detected that the training completion condition is met, the initial model after parameter adjustment is determined as the emotion recognition model.

[0092] The training completion conditions refer to the conditions indicating that the training of the initial model can be terminated. The number and content of the conditions are not limited. For example, the conditions may be conditions for limiting the training duration, or conditions for limiting the number of training rounds, or conditions for limiting the detection accuracy of the initial model. When one, some or all of the training completion conditions are met, the initial model after parameter adjustment can be determined as the emotion recognition model, indicating that the training is completed.

[0093] It is understandable that the method of detecting whether the training completion condition is satisfied is different depending on the content of the training completion condition. For example, when the training completion condition is a condition for limiting the training duration, it can be determined that the training completion condition is satisfied when it is detected that the training duration reaches the preset duration limit; when the training completion condition is a condition for limiting the number of training rounds, it can be determined that the training completion condition is satisfied when it is detected that the number of training rounds reaches the preset number of training rounds; when the training completion condition is an accuracy condition, the accuracy test of the initial model after parameter adjustment can be performed using test data to obtain a test result. If the test result is greater than a preset threshold, it is determined that the training completion condition is satisfied.

[0094] Specifically, the test data may include multiple groups of test sub-data, including target test sub-data, which may be any group of test sub-data, and the target test sub-data may include target test video, target test audio, and target test label. When performing the test, multiple target test video frames are determined in the target test video, and multiple target test text data are generated using each emotion label in the label set. It should be noted that the target test text data corresponds to at least one text template. That is, when the number of text templates is multiple, each emotion label can be used to match each text template to generate corresponding target test text data. The target test video frame, the target test text data, and the target test audio input parameter adjusted initial model are used to obtain target non-text encoded data and multiple target text encoded data, wherein each target text encoded data corresponds to each target test text data one by one. The test similarity data between the target non-text encoded data and each target text encoded data is calculated.

[0095] The larger the test similarity data, the more similar it is. Since the maximum similarity data indicates that the two are most similar, the test similarity data is used to determine at least one maximum similarity data corresponding to at least one text template. Each maximum similarity data represents the most reliable prediction result obtained when the text template is used for emotion recognition. The emotion label corresponding to at least one maximum similarity data is determined as the initial prediction result corresponding to the target test video, and the initial prediction results are screened to the maximum number to obtain the prediction result, that is, the result with the largest number among the initial prediction results corresponding to multiple text templates is used as the prediction result. Based on the prediction result and the target test label, the test sub-result corresponding to the target test sub-data is determined. If the two are the same, the test sub-result indicates that the prediction is correct, otherwise it is wrong. All the test sub-results corresponding to the test data are counted to obtain the test result.

[0096] After obtaining the emotion recognition model, you can use it to perform emotion recognition. Please refer to Figure 2 , Figure 2 A flow chart of an emotion recognition method provided in an embodiment of the present application includes:

[0097] S201: Acquire the video and audio to be tested.

[0098] S202: determining a plurality of video frames to be tested in the video to be tested, and using each emotion tag in the tag set to concatenate with the text template to be tested to generate text data to be tested corresponding to each emotion tag.

[0099] S203: Input the video frame to be tested, the text data to be tested and the audio to be tested into the emotion recognition model to obtain the non-text coded data to be tested and the text coded data to be tested corresponding to each text data to be tested.

[0100] S204: Generate similarity data to be tested by using the non-text encoded data to be tested and each text encoded data to be tested.

[0101] S205: Determine the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested.

[0102] Among them, the emotion recognition model is obtained based on any of the above-mentioned emotion recognition model training methods. In practical applications, the label set includes various emotion labels, which may include some or all of the emotion labels used in the training process, and may also include emotion labels that have not been used in the training process. Since it is not possible to determine the emotions specifically represented by the video to be tested when performing emotion recognition, each emotion label can be used to generate a corresponding text data to be tested. Among them, if a text template is used to generate the text data to be tested, each text data to be tested can use the same or different text templates. Specifically, the process of generating the text data to be tested can be: selecting a text template to be tested from a preset template library; performing vector mapping processing on the text template to be tested and each emotion label respectively to obtain a template vector to be tested and each label vector; splicing the template vector with each label vector respectively to obtain the text data to be tested. The specific generation process is similar to the training process, which will not be repeated here.

[0103] After processing using the emotion recognition model, non-text coded data to be tested corresponding to the video frame to be tested and the audio to be tested, as well as text coded data to be tested corresponding to each text data to be tested can be obtained. The non-text coded data to be tested and each text coded data to be tested are used to generate similarity data to be tested. The obtained multiple similarity data to be tested respectively represent the similarity between the features represented by the video to be tested and each emotion label, and the closest one is selected, that is, the maximum similarity data to be tested, and the corresponding emotion label is used as the emotion recognition result corresponding to the video to be tested.

[0104] Please refer to Figure 3 , Figure 3 A specific data processing flow chart provided for an embodiment of the present application. During the training process, a target text template and an emotion label are obtained, and they are respectively mapped into a prompt embedding vector and a label embedding vector by means of text preprocessing, and a generalized text vector, i.e., training text data, is generated by vector splicing. The generalized text vector is input into a text encoder in a CLIP model constructed based on CLIP pre-training weights to obtain training text encoded data. In addition, the video is frame-extracted to obtain training video frames, which are then input into a visual encoder, and the training audio is input into a sound encoder at the same time, and the data vectors of the visual encoder and the sound encoder are spliced ​​to obtain training non-text encoded data. The similarity between the training text encoded data and the training non-text encoded data is calculated, and a cross entropy loss is generated based on the similarity.

[0105] In this application, y can be used to represent the tag set of emotion tags, and x can be used to represent the training video or the video to be tested. Then the emotion tag corresponding to the maximum similarity data to be tested can be expressed as y pred , specifically:

[0106]

[0107] argmax represents the maximum value, p represents the target text template, and f vid represents the encoder of the video end. Here, the sound encoder, visual encoder and LSTM timing module are combined together as the encoder of the video end, so f vid (E1(x)) represents the non-text encoded data to be tested, f txt represents a text encoder, so f txt ([E T (p); E T (y i )]). C represents the number of emotion categories in the label set. E1 and ET represent video preprocessing (i.e., frame extraction) and text preprocessing (i.e., vector mapping), respectively.

[0108] During training, cross entropy loss can be used, expressed as Loss, specifically:

[0109]

[0110] The entire training process includes the following steps:

[0111] a. Input the face video. After the video is preprocessed, M frames of images are selected.

[0112] b. Sample the corresponding prompt from the artificially prepared prompt set, denoted as p.

[0113] c. The label vector y (specifically the vector of the emotion label corresponding to the training video) and the vector p are respectively pre-processed, and then the text embedding vector t is synthesized by vector splicing.

[0114] d. Input the text embedding vector t and M frames of images into the text encoder and visual encoder to obtain M temporal text features and M temporal image features. The text encoder and visual encoder load VIT-CLIP large-scale pre-trained weights.

[0115] e. M temporal text features are pooled in time series to obtain the final text encoding vector final_t.

[0116] f. M time series image features pass through the LSTM model, and the feature of the last node is used as the final image encoding feature final_img.

[0117] g. The sound features are output through the sound encoder as a sound coding vector, which is concatenated with the final_img obtained in step f to obtain the final video coding vector final_vid.

[0118] h. Calculate the cosine similarity of the text encoding vectors final_t ​​and final_vid, calculate the cross entropy loss, and use the loss to adjust the parameters of the pooling network module and LSTM model used in pooling.

[0119] During the test, you can perform the following steps:

[0120] a. Input the face video. After the video is preprocessed, M frames of images are selected.

[0121] b. Let the set of manually created prompts be denoted as P, where each prompt is denoted as p, and each p executes steps c to h.

[0122] c. The vectors corresponding to each emotion label in the label vector set y are respectively pre-processed with the vector p, and then the text embedding vector t is synthesized by vector splicing.

[0123] d. Input the text embedding vector t and M frames of images into the text encoder and visual encoder to obtain M temporal text features and M temporal image features. The text encoder and visual encoder load VIT-CLIP large-scale pre-trained weights.

[0124] e. M temporal text features are pooled in time series to obtain the final text encoding vector final_t.

[0125] f. M time series image features pass through the LSTM model, and the feature of the last node is used as the final image encoding feature final_img.

[0126] g. The sound features are output through the sound encoder as a sound coding vector, which is concatenated with the final_img obtained in step f to obtain the final video coding vector final_vid.

[0127] h. Select the emotion category corresponding to the video for each p according to the following formula:

[0128]

[0129] Among them, f vid (E1(x)) represents final_vid, f txt ([E T (p); E T (y i )]) means final_t.

[0130] i. According to the votes corresponding to each p, the corresponding final sentiment category is obtained.

[0131] During the application process, the following steps can be performed:

[0132] a. Input the face video. After the video is preprocessed, M frames of images are selected.

[0133] b. Let the artificially created prompt set be denoted as P, each prompt in it be denoted as p, and select the target template p0 from P.

[0134] c. The vectors corresponding to each emotion label in the label vector set y are respectively pre-processed with the vector p0, and then the text embedding vector t0 is synthesized by vector splicing.

[0135] d. Input the text embedding vector t0 and M frames of images into the text encoder and visual encoder to obtain M temporal text features and M temporal image features. The text encoder and visual encoder are loaded with VIT-CLIP large-scale pre-trained weights.

[0136] e. M temporal text features are pooled in time series to obtain the final text encoding vector final_t0.

[0137] f. M time series image features pass through the LSTM model, and the feature of the last node is used as the final image encoding feature final_img.

[0138] g. The sound features are output through the sound encoder as a sound coding vector, which is then concatenated with the final_img obtained in step f to obtain the final video coding vector final_vid.

[0139] h. Select the emotion category corresponding to the video for p0 according to the following formula:

[0140]

[0141] Among them, f vid (E1(x)) represents final_vid, f txt ([E T (p); E T (y i )]) represents final_t0.

[0142] By applying the emotion recognition model training and emotion recognition method provided by the embodiment of the present application, the emotion recognition process is converted from the original probability prediction problem to a similarity matching problem, and the semantic information contained in the label itself is introduced. While improving the accuracy, the model can also have a certain zero-shot learning migration ability. Specifically, when training the emotion recognition model, the present application generates training text data using emotion labels, and uses it to train the initial model, so that the initial model can learn the semantic information carried by the emotion label. After the encoding is completed, the loss value is calculated and the parameters are adjusted through the similarity data, so that the encoding process of the initial model focuses on reflecting the similarity between text and non-text. When applied, the most similar emotion label is determined by the similarity between the non-text encoded data of the video to be tested and the text encoded data to be tested corresponding to each emotion label, thereby improving the accuracy of emotion recognition. At the same time, even if an emotion label that is not involved in the training of the emotion recognition model is added during application, the emotion recognition model can distinguish it from other emotion labels based on the semantic information of the emotion label, and has a certain zero-sample learning ability, which improves the versatility of the model.

[0143] In addition, in practical applications, the trained emotion recognition model can be applied to the recognition terminal. The recognition terminal may include a processor, a detection component and a display screen, and of course, an input component. The processor is connected to the detection component, the input component and the display screen respectively, and the processor can obtain the video to be tested and the audio to be tested; determine multiple video frames to be tested in the video to be tested, and use each emotion label in the label set to respectively splice with the text template to be tested to generate the text data to be tested corresponding to each emotion label; input the video frame to be tested, the text data to be tested and the audio to be tested into the emotion recognition model to obtain the non-text encoding data to be tested and the text encoding data to be tested corresponding to each text data to be tested; use the non-text encoding data to be tested and each text encoding data to be tested to generate the similarity data to be tested; determine the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested. After obtaining the emotion recognition result, the emotion recognition result can be displayed on the display screen.

[0144] In practical applications, the detection component may include a detection interface and a collection component (such as a camera and a microphone). The input component may include an input interface and an input keyboard, and the input keyboard may facilitate the user to input relevant instructions or data to the identification terminal. In order to reduce the wiring difficulty and meet the data transmission requirements, a wireless transmission module may also be provided on the identification terminal. Among them, the wireless transmission module may be a Bluetooth module or a wifi module.

[0145] Figure 4The structural diagram of an identification terminal provided in an embodiment of the present application may include a processor, a display screen 41, an input interface 42, an input keyboard 43, a detection interface 44, a camera 45, a microphone 46, and a wireless transmission module 47. When the display screen 41 is a touch screen, the input keyboard 43 may be a soft keyboard presented on the display screen 41. The input interface 42 may be used to achieve connection with an external device. There may be multiple input interfaces, Figure 3 In the example of an input interface, the detection interface 44 is connected to the collection component 45. The processor is embedded in the identification terminal, so it is not in the Figure 3 Shown in.

[0146] The identification terminal may be a smart phone, a tablet computer, a laptop computer, or a desktop computer, etc. In the embodiment of the present application, the form of the identification terminal is not limited. When the identification terminal is a smart phone or a tablet computer, the input interface 42 may be connected to an external device via a data cable, and the input keyboard 43 may be a soft keyboard presented on the display interface. When the identification terminal is a laptop computer or a desktop computer, the input interface 42 may be a USB interface for connecting an external device such as a USB flash drive, and the input keyboard 43 may be a hard keyboard.

[0147] Taking a desktop computer as an example, in actual application, the user can import the video to be tested and the audio to be tested into a U disk, and insert the U disk into the input interface 52 of the recognition terminal. After acquiring the video to be tested and the audio to be tested, the recognition terminal determines multiple video frames to be tested in the video to be tested, and uses each emotion label in the label set to respectively splice with the text template to be tested to generate text data to be tested corresponding to each emotion label, and inputs the video frame to be tested, the text data to be tested and the audio to be tested into the emotion recognition model to obtain the non-text coding data to be tested and the text coding data to be tested corresponding to each text data to be tested, and uses the non-text coding data to be tested to respectively generate the similarity data to be tested with each text coding data to be tested, and determines the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested, and displays the recognition result through the display screen 41. It should be noted that Figure 5 The functional modules included in the identification terminal, such as the display screen 41, input interface 42, input keyboard 43, detection interface 44, camera 45, microphone 46, wireless transmission module 47, etc., are only examples. In actual applications, the question and answer terminal may also include more or fewer functional modules based on actual needs, and there is no limitation on this.

[0148] The emotion recognition method provided in the embodiment of the present application can be deployed in a software platform based on a neural network acceleration application or an AI (Artificial Intelligence) acceleration chip based on an FPGA (Field Programmable Gate Array). It should be noted that the embodiment of the present application compresses the neural network model based on the offset, which can be applied to time series data processing based on LSTM (Long Short-Term Memory) in addition to determining text answers, such as multi-target tracking and other scenarios.

[0149] The emotion recognition device provided in the embodiment of the present application is introduced below. The emotion recognition device described below and the emotion recognition model training method described above can be referenced to each other.

[0150] Please refer to Figure 5 , Figure 5 A schematic diagram of the structure of an emotion recognition device provided in an embodiment of the present application includes:

[0151] The test acquisition module 51 is used to acquire the video and audio to be tested;

[0152] The test data processing module 52 is used to determine a plurality of test video frames in the test video, and use each emotion label in the label set to splice with the test text template to generate the test text data corresponding to each emotion label;

[0153] The test input module 53 is used to input the test video frame, the test text data and the test audio into the emotion recognition model to obtain the test non-text code data and the test text code data corresponding to each test text data;

[0154] A similarity generation module 54 is used to generate similarity data to be tested by using the non-text coded data to be tested and each of the text coded data to be tested;

[0155] The recognition result determination module 55 is used to determine the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested.

[0156] Optionally, the test data processing module 52 includes:

[0157] A test template determination unit, used for selecting the test text template from a preset template library;

[0158] A test vector mapping unit is used to perform vector mapping processing on the test text template and each of the emotion labels to obtain a test template vector and each label vector;

[0159] The test splicing unit is used to splice the template vector with each of the label vectors to obtain the test text data.

[0160] Optionally, it also includes:

[0161] The training acquisition module is used to obtain training videos, training audios and emotion labels;

[0162] A training data processing module, used to determine a plurality of training video frames in the training video, and generate training text data using the emotion tags;

[0163] A training input module, used for inputting the training video frame, the training text data and the training audio into an initial model to obtain training text encoding data and training non-text encoding data;

[0164] A training similarity generation module, used to generate similarity data using the training text encoding data and the training non-text encoding data;

[0165] A parameter adjustment module, used to generate a loss value using the similarity data, and adjust the parameters of the initial model based on the loss value;

[0166] The model determination module is used to determine the initial model after parameter adjustment as the emotion recognition model if it is detected that the training completion condition is met.

[0167] Optionally, the initial model includes a text encoder, an image encoder and an audio encoder, and also includes a pooling network module and a time recursive network module, the output of the text encoder is the input of the pooling network module, and the output of the image encoder is the input of the time recursive network module.

[0168] Optionally, a training input module includes:

[0169] A training text encoding unit, used for inputting the training text into the text encoder to obtain a plurality of initial text encodings;

[0170] A training pooling processing unit, used for inputting the multiple initial text encodings into the pooling network module to obtain the training text encoding data;

[0171] A training audio encoding unit, used for inputting the training video frame into the image encoder to obtain a plurality of initial image codes, and inputting the training audio into the audio encoder to obtain an initial audio code;

[0172] A training image coding unit, used for inputting the multiple initial image codes into the time recursive network module to obtain intermediate image codes;

[0173] A training splicing unit is used to splice the intermediate image code and the initial audio code to obtain the training non-text code data.

[0174] Optionally, the text encoder and the image encoder belong to a language-image contrast learning pre-training model, and the audio encoder is pre-trained;

[0175] Parameter adjustment module, including:

[0176] A partial adjustment unit is used to adjust parameters of the pooling network module and the time recursive network module in the initial model based on the loss value.

[0177] Optionally, the training data processing module includes:

[0178] A target template selection unit, used to select a target text template from a preset template library;

[0179] A vector mapping unit, used for performing vector mapping processing on the target text template and the emotion label to obtain a template vector and a label vector;

[0180] The text vector concatenation unit is used to concatenate the template vector and the label vector to obtain the training text data.

[0181] Optionally, the model determination module includes:

[0182] A testing unit, used to perform an accuracy test on the initial model after parameter adjustment using test data to obtain a test result;

[0183] A determination unit is used to determine whether the training completion condition is met if the test result is greater than a preset threshold.

[0184] Optionally, the test data includes multiple groups of test sub-data, including target test sub-data, and the target test sub-data includes target test video, target test audio and target test label;

[0185] Test unit, including:

[0186] A test data processing subunit, configured to determine a plurality of target test video frames in the target test video, and generate a plurality of target test text data using each emotion tag in the tag set; wherein the target test text data corresponds to at least one text template;

[0187] A test input subunit, used to obtain target non-text encoded data and multiple target text encoded data by inputting the target test video frame, the target test text data and the target test audio into the initial model after parameter adjustment;

[0188] A test calculation subunit, used for calculating test similarity data between the target non-text encoded data and each target text encoded data, and determining at least one maximum similarity data corresponding to the at least one text template respectively by using the test similarity data;

[0189] A prediction result determination subunit, used to determine the emotion tag corresponding to the at least one maximum similarity data as the initial prediction result corresponding to the target test video, and perform maximum number screening on the initial prediction result to obtain a prediction result;

[0190] a sub-result determination sub-unit, configured to determine a test sub-result corresponding to the target test sub-data based on the prediction result and the target test label;

[0191] The statistical subunit is used to count all test sub-results corresponding to the test data to obtain the test result.

[0192] The electronic device provided in the embodiment of the present application is introduced below. The electronic device described below and the emotion recognition model training method and / or the emotion recognition method described above can be referenced to each other.

[0193] Please refer to Figure 6 , Figure 6 The electronic device 100 may include a processor 101 and a memory 102 , and may further include one or more of a multimedia component 103 , an information input / information output (I / O) interface 104 , and a communication component 105 .

[0194] The processor 101 is used to control the overall operation of the electronic device 100 to complete the above-mentioned emotion recognition model training method and / or all or part of the steps in the emotion recognition method; the memory 102 is used to store various types of data to support the operation of the electronic device 100, and these data may include, for example, instructions for any application or method operated on the electronic device 100, and application-related data. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, EPROM), programmable read-only memory (Programmable Read-Only Memory, PROM), read-only memory (Read-Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disk. One or more.

[0195] The multimedia component 103 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, which is used to receive external audio signals. The received audio signal may be further stored in the memory 102 or sent through the communication component 105. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 104 provides an interface between the processor 101 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 105 is used for wired or wireless communication between the electronic device 100 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 105 may include: Wi-Fi components, Bluetooth components, NFC components.

[0196] The electronic device 100 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the emotion recognition model training method and / or emotion recognition method given in the above embodiments.

[0197] The computer-readable storage medium provided in an embodiment of the present application is introduced below. The computer-readable storage medium described below and the emotion recognition model training method and / or the emotion recognition method described above can be referenced to each other.

[0198] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned emotion recognition model training method and / or the steps of the emotion recognition method are implemented.

[0199] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0200] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0201] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0202] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0203] Finally, it should be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms include, include or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0204] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An emotion recognition method, characterized in that: include: Get the video and audio to be tested; Determine a plurality of video frames to be tested in the video to be tested, and use each emotion label in the label set to respectively splice with the text template to be tested to generate text data to be tested corresponding to each emotion label; Inputting the video frame to be tested, the text data to be tested and the audio to be tested into an emotion recognition model to obtain the non-text coded data to be tested and the text coded data to be tested corresponding to each of the text data to be tested; Generate similarity data to be tested using the non-text encoded data to be tested and each of the text encoded data to be tested; The emotion label corresponding to the maximum similarity data to be tested is determined as the emotion recognition result corresponding to the video to be tested; wherein: The step of splicing each emotion tag in the tag set with the text template to be tested to generate the text data to be tested corresponding to each emotion tag includes: Selecting the text template to be tested from a preset template library; Performing vector mapping processing on the text template to be tested and each of the emotion labels respectively to obtain a template vector to be tested and each label vector; The template vector is concatenated with each of the label vectors to obtain the text data to be tested; The training process of the emotion recognition model includes: Get training videos, training audios, and emotion labels; Determine a plurality of training video frames in the training video, select a target text template from a preset template library, perform vector mapping processing on the target text template and the emotion label to obtain a template vector and a label vector, and splice the template vector and the label vector to obtain training text data; The training video frame, the training text data and the training audio are input into an initial model to obtain training text encoding data and training non-text encoding data; the training text encoding data is obtained based on the training text data and represents the emotional semantic features of the emotional label; the training non-text encoding data is obtained based on the training video frame and the training audio and represents the emotional features of the image and sound representation; Generate similarity data using the training text encoding data and the training non-text encoding data; Generating a loss value using the similarity data, and adjusting parameters of the initial model based on the loss value; If it is detected that the training completion condition is met, the initial model after parameter adjustment is determined as the emotion recognition model; The initial model includes a text encoder, an image encoder and an audio encoder, and also includes a pooling network module and a time recursive network module, the output of the text encoder is the input of the pooling network module, and the output of the image encoder is the input of the time recursive network module; The step of inputting the training video frame, the training text data and the training audio into an initial model to obtain training text encoding data and training non-text encoding data includes: Inputting the training text into the text encoder to obtain a plurality of initial text encodings; Inputting the multiple initial text encodings into the pooling network module to obtain the training text encoding data; Input the training video frame into the image encoder to obtain a plurality of initial image codes, and input the training audio into the audio encoder to obtain initial audio codes; Inputting the multiple initial image codes into the time recursive network module to obtain intermediate image codes; The intermediate image code and the initial audio code are concatenated to obtain the training non-text code data.

2. The emotion recognition model training method according to claim 1, characterized in that: The text encoder and the image encoder belong to a language-image contrast learning pre-training model, and the audio encoder has been pre-trained.

3. The emotion recognition model training method according to claim 2, characterized in that: The step of adjusting parameters of the initial model based on the loss value includes: Parameters of the pooling network module and the time recursive network module in the initial model are adjusted based on the loss value.

4. The emotion recognition model training method according to claim 1, characterized in that: The detecting that the training completion condition is met includes: Using the test data, the accuracy of the initial model after parameter adjustment is tested to obtain a test result; If the test result is greater than a preset threshold, it is determined that the training completion condition is met.

5. The emotion recognition model training method according to claim 4, characterized in that: The test data includes multiple groups of test sub-data, including target test sub-data, and the target test sub-data includes target test video, target test audio and target test label.

6. The emotion recognition model training method according to claim 5, characterized in that: The using of the test data to perform an accuracy test on the initial model after parameter adjustment to obtain a test result includes: Determine a plurality of target test video frames in the target test video, and generate a plurality of target test text data using each emotion tag in the tag set; wherein the target test text data corresponds to at least one text template; The target test video frame, the target test text data and the initial model after adjusting the target test audio input parameters are obtained to obtain target non-text encoding data and multiple target text encoding data; Calculating test similarity data between the target non-text encoded data and each target text encoded data, and using the test similarity data to determine at least one maximum similarity data corresponding to the at least one text template; Determine the emotion label corresponding to the at least one maximum similarity data as the initial prediction result corresponding to the target test video, and perform maximum number screening on the initial prediction results to obtain a prediction result; Determine a test sub-result corresponding to the target test sub-data based on the prediction result and the target test label; All test sub-results corresponding to the test data are counted to obtain the test result.

7. The emotion recognition model training method according to claim 1, characterized in that: The detecting that the training completion condition is met includes: When it is detected that the training duration reaches a preset duration limit, determining that the training completion condition is met; Or when it is detected that the number of training rounds reaches a preset number of training rounds, it is determined that the training completion condition is met.

8. An emotion recognition device, characterized in that: include: A test acquisition module is used to acquire the video and audio to be tested; The test data processing module is used to determine a plurality of test video frames in the test video, and use each emotion label in the label set to splice with the test text template to generate the test text data corresponding to each emotion label; The test input module is used to input the test video frame, the test text data and the test audio into the emotion recognition model to obtain the test non-text code data and the test text code data corresponding to each test text data; A similarity generation module to be tested, used for generating similarity data to be tested by using the non-text encoded data to be tested and each of the text encoded data to be tested; The recognition result determination module is used to determine the emotion label corresponding to the maximum similarity data to be tested as the emotion recognition result corresponding to the video to be tested; wherein: The test data processing module comprises: A test template determination unit, used for selecting the test text template from a preset template library; A test vector mapping unit is used to perform vector mapping processing on the test text template and each of the emotion labels to obtain a test template vector and each label vector; A test splicing unit, used for splicing the template vector with each of the label vectors to obtain the test text data; Also includes: The training acquisition module is used to obtain training videos, training audios and emotion labels; A training data processing module is used to determine a plurality of training video frames in the training video, and respectively perform vector mapping processing on the target text template and the emotion label to obtain a template vector and a label vector, and splice the template vector and the label vector to obtain training text data; A training input module is used to input the training video frame, the training text data and the training audio into an initial model to obtain training text encoding data and training non-text encoding data; the training text encoding data is obtained based on the training text data and represents the emotional semantic features of the emotional label; the training non-text encoding data is obtained based on the training video frame and the training audio and represents the emotional features of the image and sound representation; A training similarity generation module, used to generate similarity data using the training text encoding data and the training non-text encoding data; A parameter adjustment module, used to generate a loss value using the similarity data, and adjust the parameters of the initial model based on the loss value; A model determination module, configured to determine the initial model after parameter adjustment as an emotion recognition model if it is detected that a training completion condition is met; The initial model includes a text encoder, an image encoder and an audio encoder, and also includes a pooling network module and a time recursive network module, the output of the text encoder is the input of the pooling network module, and the output of the image encoder is the input of the time recursive network module; The training input module comprises: A training text encoding unit, used for inputting the training text into the text encoder to obtain a plurality of initial text encodings; A training pooling processing unit, used for inputting the multiple initial text encodings into the pooling network module to obtain the training text encoding data; A training audio encoding unit, used for inputting the training video frame into the image encoder to obtain a plurality of initial image codes, and inputting the training audio into the audio encoder to obtain an initial audio code; A training image coding unit, used for inputting the multiple initial image codes into the time recursive network module to obtain intermediate image codes; A training splicing unit, used for splicing the intermediate image code and the initial audio code to obtain the training non-text code data; The training data processing module comprises: A target template selection unit, used to select a target text template from a preset template library; A vector mapping unit, used for performing vector mapping processing on the target text template and the emotion label to obtain a template vector and a label vector; The text vector concatenation unit is used to concatenate the template vector and the label vector to obtain the training text data.

9. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store the computer program; The processor is used to execute the computer program to implement the emotion recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the emotion recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Facial expression recognition method and device based on zero sample learning

    CN113920561A

  • Video emotion recognition method based on multi-modal representation learning

    CN114550057A