A multimodal emotion recognition method

By performing the fusion processing of speech and text features and self-attention-weighted calculation in the multimodal emotion recognition method, the problem of low emotional recognition accuracy in the prior art is solved, and higher emotional recognition accuracy is achieved.

CN119475252BActive Publication Date: 2025-05-13CETC NEW SMART CITY RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510055932.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing multimodal emotion recognition method lacks correlation analysis between speech and text and the extraction of key information in the context of the sentence, resulting in low accuracy of emotion recognition.

Method used

By obtaining the initial modal features of the multimodal data in the video data to be tested, fusion splicing and timing feature processing are performed, vocabulary-level multimodal fusion features are obtained, and video-level multimodal fusion features are obtained through self-attention weighting calculation, and finally emotional recognition processing is performed to obtain the emotional prediction results.

Benefits of technology

By making full use of the correlation between pronunciation and text and sentence context information, the accuracy of emotion recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119475252B_ABST
    Figure CN119475252B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of emotion recognition technology, and provides a multimodal emotion recognition method, which includes: obtaining the initial modal features of multiple modal data contained in the video data to be tested; fusing and splicing the initial modal features of the multiple modal data and processing the time series features to obtain multiple vocabulary-level multimodal fusion features; then, performing self-attention weighted calculation on the multiple vocabulary-level multimodal fusion features to obtain the video-level multimodal fusion features of the video data to be tested; performing emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain the emotion prediction results corresponding to the video data to be tested. The accuracy of emotion recognition is improved by fusing the initial modal features of the audio modal data, the image modal data, and the text modal data, and assigning more weight information to the key information of the vocabulary-level multimodal fusion features through self-attention weighted calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of emotion recognition technology, and in particular, relates to a multimodal emotion recognition method. Background Art

[0002] Human emotions can be expressed in multiple modes, such as language, facial expressions, and text, and these modes complement each other emotionally. Multimodal emotion recognition combines multiple sensors and signals, such as video, audio, and physiological data, and uses algorithms to convert human emotions into machine-recognizable features to more comprehensively capture emotional expressions.

[0003] In the current multimodal emotion recognition methods, due to the complementarity between modalities, the fusion strategy of inter-modal features is very important. There are three main types of multimodal fusion methods: early fusion, mid-term fusion, and late fusion. Early fusion is to fuse the features of different modalities through feature splicing in the early stage of the multimodal emotion recognition model; late fusion is to establish a separate model for each modality, and then use majority voting or weighted average method to integrate the output; mid-term fusion is to convert the data of different modalities into high-dimensional feature expressions first, and then fuse them through the middle layer of the model.

[0004] In sentiment analysis, there is a strong correlation between speech and text. There are strong temporal characteristics between expressions, speech, and text. The correlation between word-level and sentence-level features is also very large. However, the fusion strategy of current sentiment recognition methods does not fully utilize the correlation between speech and text, and lacks the extraction of key information in the sentence context, resulting in low accuracy of sentiment recognition. Summary of the invention

[0005] The embodiment of the present application provides a multimodal emotion recognition method, which can solve the problem of low accuracy of emotion recognition due to the lack of analysis of the correlation between speech and text and the lack of extraction of key information in the sentence context in existing emotion recognition methods.

[0006] In a first aspect, an embodiment of the present application provides a multimodal emotion recognition method, the method comprising:

[0007] Acquire initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data and text modal data;

[0008] The initial modal features of the plurality of modal data are fused and spliced ​​and subjected to temporal feature processing to obtain a plurality of vocabulary-level multimodal fusion features;

[0009] Performing self-attention weighted calculation on a plurality of the vocabulary-level multimodal fusion features to obtain video-level multimodal fusion features of the video data to be tested;

[0010] Emotion recognition processing is performed on the video data to be tested according to the video-level multimodal fusion features to obtain an emotion prediction result corresponding to the video data to be tested.

[0011] In a possible implementation manner of the first aspect, the initial modality features include: a plurality of vocabulary-level text features of the text modality data, a plurality of vocabulary-level audio features of the audio modality data, and a plurality of vocabulary-level expression features of the image modality data.

[0012] The fusing, splicing and time series feature processing of the initial modal features of the plurality of modal data to obtain a plurality of vocabulary-level multimodal fusion features includes:

[0013] Fusing the plurality of vocabulary-level text features and the plurality of vocabulary-level audio features to obtain a plurality of vocabulary-level audio-text fusion features;

[0014] splicing the plurality of the lexical level audio text fusion features with the plurality of the lexical level expression features to obtain a plurality of initial lexical level multimodal fusion features;

[0015] Performing time series feature processing on the plurality of the initial vocabulary level multimodal fusion features to obtain the plurality of the vocabulary level multimodal fusion features.

[0016] In a possible implementation manner of the first aspect, the fusing the plurality of vocabulary-level text features and the plurality of vocabulary-level audio features to obtain a plurality of vocabulary-level audio-text fusion features includes:

[0017] The contextual relationship of the plurality of vocabulary-level text features is captured by a self-attention mechanism to obtain a plurality of vocabulary-level text modal temporal features corresponding to the plurality of vocabulary-level text features;

[0018] The contextual relationship of the plurality of the lexical level audio features is captured by the self-attention mechanism to obtain a plurality of lexical level audio modal time series features corresponding to the plurality of the lexical level audio features;

[0019] Performing smooth window movement according to a preset number of vocabulary windows and a vocabulary order in the text modality data to obtain contextual vocabulary level text features corresponding to a plurality of the vocabulary level text features and contextual vocabulary level audio features corresponding to a plurality of the vocabulary level audio features;

[0020] Calculating audio coordination features of multiple text modalities and text coordination features of multiple audio modalities based on the multiple vocabulary-level text features, the multiple vocabulary-level audio features, the multiple context vocabulary-level text features, and the multiple context vocabulary-level audio features;

[0021] A plurality of the vocabulary-level text modality time series features, a plurality of the vocabulary-level audio modality time series features, a plurality of the audio coordination features of the text modality, and a plurality of the text coordination features of the audio modality are added and averaged to obtain a plurality of the vocabulary-level audio-text fusion features.

[0022] In a possible implementation manner of the first aspect, the step of calculating, according to the plurality of vocabulary-level text features, the plurality of vocabulary-level audio features, the plurality of context vocabulary-level text features, and the plurality of context vocabulary-level audio features, obtaining audio collaborative features of the plurality of text modalities and text collaborative features of the plurality of audio modalities includes:

[0023] Determining a first similarity between the plurality of vocabulary-level text features and the plurality of context vocabulary-level audio features according to the plurality of vocabulary-level text features, the plurality of context vocabulary-level audio features, and a preset similarity formula;

[0024] Calculating and obtaining multiple audio collaborative features of the text modalities according to the multiple contextual vocabulary-level audio features and the first similarity;

[0025] Determining a second similarity between the plurality of vocabulary-level audio features and the plurality of context vocabulary-level text features according to the plurality of vocabulary-level audio features, the plurality of context vocabulary-level text features, and a preset similarity formula;

[0026] Based on the multiple contextual vocabulary level text features and the second similarity, the multiple text collaborative features of the audio modalities are calculated.

[0027] In a possible implementation manner of the first aspect, performing temporal feature processing on the multiple initial vocabulary level multimodal fusion features to obtain the multiple vocabulary level multimodal fusion features includes:

[0028] Through a bidirectional long short-term memory network Bi-LSTM model, context feature extraction is performed on the initial vocabulary level multimodal fusion feature at each moment to obtain forward hidden layer features and backward hidden layer features corresponding to the initial vocabulary level multimodal fusion feature at that moment;

[0029] The forward hidden layer features and the backward hidden layer features are concatenated to obtain the vocabulary-level multimodal fusion features corresponding to the initial vocabulary-level multimodal fusion features at that moment.

[0030] In a possible implementation manner of the first aspect, the obtaining initial modal features of multiple modal data contained in the video data to be tested includes:

[0031] Decomposing the video data to be tested to obtain the text modality data, the audio modality data and the image modality data in the video data to be tested;

[0032] Feature extraction is performed on the text modality data, the audio modality data, and the image modality data, respectively, to obtain a plurality of vocabulary-level text features of the text modality data, a plurality of vocabulary-level audio features of the audio modality data, and a plurality of vocabulary-level expression features of the image modality data.

[0033] In a possible implementation manner of the first aspect, the extracting features from the text modality data to obtain the plurality of vocabulary-level text features in the text modality data includes:

[0034] Performing text segmentation processing on the text modality data through a word vector dictionary to obtain a plurality of words in the text modality data;

[0035] The word embedding feature processing is performed on multiple words in the text modality data through a pre-trained language representation model to obtain multiple word-level text features.

[0036] In a possible implementation manner of the first aspect, the extracting features from the audio modality data to obtain the plurality of vocabulary-level audio features in the audio modality data includes:

[0037] Marking the start time and the end time of the plurality of words in the text modal data in the audio modal data to obtain first time marking information of the plurality of words in the audio modal data;

[0038] Segmenting the audio modality data according to the first time tag information to obtain a plurality of audio segment data corresponding to a plurality of words;

[0039] The multiple audio segment data are subjected to feature extraction through an automatic speech recognition model to obtain multiple vocabulary-level audio features.

[0040] In a possible implementation manner of the first aspect, the extracting features of the image modality data to obtain the plurality of vocabulary-level expression features in the image modality data includes:

[0041] Marking the start time and the end time of the plurality of words in the text modality data in the image modality data to obtain second time marking information of the plurality of words in the image modality data;

[0042] Segmenting the image modality data according to the second time tag information and a preset acquisition frequency to obtain a plurality of segmented images corresponding to a plurality of words;

[0043] Performing face recognition on the plurality of segmented images through a neural network model to obtain a plurality of face regions in the plurality of segmented images;

[0044] Extracting image features of the plurality of face regions to obtain a plurality of image features in the plurality of segmented images;

[0045] The plurality of image features corresponding to the plurality of words are summed up and averaged to obtain a plurality of word-level expression features.

[0046] In a possible implementation manner of the first aspect, performing self-attention weighted calculation on the plurality of vocabulary-level multimodal fusion features to obtain the video-level multimodal fusion features of the video data to be tested includes:

[0047] Performing self-attention calculation on the plurality of lexical-level multimodal fusion features in the plurality of sentences through a self-attention mechanism to obtain lexical weight information corresponding to the plurality of lexical items in the plurality of sentences;

[0048] According to the vocabulary weight information corresponding to the multiple words in the multiple sentences, weighted averaging the multiple vocabulary-level multimodal fusion features in the multiple sentences is performed to obtain multiple sentence-level multimodal fusion features corresponding to the multiple sentences;

[0049] A plurality of the sentence-level multimodal fusion features corresponding to a plurality of sentences in the video data to be tested are added and averaged to obtain the video-level multimodal fusion features of the video data to be tested.

[0050] In a possible implementation manner of the first aspect, performing emotion recognition processing on the video data to be tested according to the video-level multimodal fusion feature to obtain an emotion prediction result corresponding to the video data to be tested includes:

[0051] Performing a first linear processing on the video-level multimodal fusion feature through a first linear layer of a multimodal sentiment analysis model to obtain a first linear feature corresponding to the video-level multimodal fusion feature;

[0052] Performing nonlinear correction on the first linear feature by using a first activation function of the multimodal sentiment analysis model to obtain a first corrected feature;

[0053] Performing a second linear processing on the first corrected feature through the second linear layer of the multimodal sentiment analysis model to obtain an initial sentiment prediction result corresponding to the video data to be tested;

[0054] The initial emotion prediction result is normalized by the second activation function of the multimodal emotion analysis model to obtain the emotion prediction result corresponding to the video data to be tested.

[0055] In a second aspect, an embodiment of the present application provides a multimodal emotion recognition device, the device comprising:

[0056] A feature acquisition module, used to acquire initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes at least one of the following: audio modal data, image modal data and text modal data;

[0057] A first fusion module is used to fuse and splice the initial modal features of the plurality of modal data and perform temporal feature processing to obtain a plurality of vocabulary-level multimodal fusion features;

[0058] A second fusion module is used to perform self-attention weighted calculation on the plurality of vocabulary-level multimodal fusion features to obtain video-level multimodal fusion features of the video data to be tested;

[0059] The emotion recognition module is used to perform emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain emotion prediction results corresponding to the video data to be tested.

[0060] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the multimodal emotion recognition methods described above when executing the computer program.

[0061] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multimodal emotion recognition method described in any one of the above items is implemented.

[0062] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the multimodal emotion recognition method described in any one of the above-mentioned first aspects.

[0063] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0064] The embodiment of the present application provides a multimodal emotion recognition method, including: obtaining the initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data and text modal data; fusing and splicing the initial modal features of the multiple modal data and processing the temporal features to obtain multiple vocabulary-level multimodal fusion features; then, performing self-attention weighted calculation on the multiple vocabulary-level multimodal fusion features to obtain the video-level multimodal fusion features of the video data to be tested; performing emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain the emotion prediction results corresponding to the video data to be tested. The accuracy of emotion recognition is improved by fusing the initial modal features of the audio modal data, image modal data and text modal data, and assigning more weight information to the key information of the vocabulary-level multimodal fusion features through self-attention weighted calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0066] Figure 1 It is a flowchart of a multimodal emotion recognition method provided in one embodiment of the present application;

[0067] Figure 2 is a flowchart of a multimodal emotion recognition method provided by another embodiment of the present application;

[0068] Figure 3 is a structural schematic diagram of a multimodal emotion recognition device provided in one embodiment of the present application;

[0069] Figure 4 It is a structural diagram of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0070] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0071] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0072] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0073] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0074] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0075] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0076] See also Figure 1 , Figure 1 1 is a flow chart of a multimodal emotion recognition method provided in an embodiment of the present application. The multimodal emotion recognition method includes:

[0077] S11, obtaining initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data and text modal data;

[0078] S12, fusing and splicing the initial modal features of multiple modal data and performing temporal feature processing to obtain multiple vocabulary-level multimodal fusion features;

[0079] S13, performing self-attention weighted calculation on multiple vocabulary-level multimodal fusion features to obtain video-level multimodal fusion features of the video data to be tested;

[0080] S14. Perform emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain emotion prediction results corresponding to the video data to be tested.

[0081] It should be noted that, in this embodiment, the execution subject may be a terminal device such as a server, and no specific limitation is made to this.

[0082] The video data to be tested is the video data that needs to be subjected to emotion recognition. The video data to be tested includes information of multiple modalities, such as audio, image and text. By analyzing the information of multiple modalities of audio, image and text in the video data to be tested, the emotional content in the video data to be tested can be identified. Modal data is media data in different forms in video data, namely audio modal data (voice, background music, etc.), image modal data (character expression, action, etc.) and text modal data (subtitles, dialogue, etc.). The initial modal features are preliminary features at the lexical level extracted from the multiple modal data, namely audio features, expression features and text features. The lexical level multimodal fusion features are features obtained by fusing and splicing the initial modal features of different modalities and processing the temporal features. The lexical level multimodal fusion features reflect the emotional information at the lexical level. The video level multimodal fusion features are fusion features obtained by further performing self-attention weighted calculation on the lexical level multimodal fusion features. The video level multimodal fusion features integrate multiple modal information in the video and reflect the overall emotional information of the entire video. The emotion prediction result is the prediction result of the emotion to be expressed by the video data to be tested. The emotion prediction result can be a discrete emotion category (such as happiness, sadness, anger, etc.) or a continuous emotion distribution.

[0083] Fusion and splicing is the process of combining and splicing the initial modal features of different modalities. Through fusion and splicing, the initial modal features of different modalities can be formed into a data containing multiple information features. Time series feature processing is to perform time series analysis on features and extract time-related information. Time series feature processing can help capture dynamic emotional changes in the video data to be tested. Self-attention weighted calculation is a calculation method used to emphasize important features and weaken secondary features. Through self-attention weighted calculation, the importance weight of each vocabulary-level multimodal fusion feature can be calculated, which helps to obtain more accurate and effective video-level multimodal fusion features.

[0084] Specifically, first, the initial modal features of audio modal data, image modal data and text modal data are extracted from the video data to be tested. Secondly, the initial modal features of different modalities are fused and spliced, and the fused and spliced ​​features are processed by time series features to capture the dynamic changes in time, thereby obtaining vocabulary-level multimodal fusion features. Then, the self-attention mechanism is applied to perform weighted calculation on the multimodal fusion features of different vocabulary levels to obtain accurate and effective video-level multimodal fusion features. Among them, the self-attention mechanism can learn the correlation between different features and highlight the features that are important for emotion recognition. Finally, the emotion recognition model is used to perform emotion prediction on the video-level multimodal fusion features to obtain the final emotion prediction result. Among them, the emotion recognition model is an emotion classifier that can identify different emotion categories based on the video-level multimodal fusion features.

[0085] It can be understood that the present embodiment provides a multimodal emotion recognition method, including: obtaining the initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data and text modal data; fusing and splicing the initial modal features of the multiple modal data and processing the temporal features to obtain multiple vocabulary-level multimodal fusion features; then, performing self-attention weighted calculation on the multiple vocabulary-level multimodal fusion features to obtain the video-level multimodal fusion features of the video data to be tested; performing emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain the emotion prediction result corresponding to the video data to be tested. The accuracy of emotion recognition is improved by fusing the initial modal features of the audio modal data, image modal data and text modal data, and assigning more weight information to the key information of the vocabulary-level multimodal fusion features through self-attention weighted calculation.

[0086] In a possible implementation, the initial modality features include: a plurality of vocabulary-level text features of text modality data, a plurality of vocabulary-level audio features of audio modality data, and a plurality of vocabulary-level expression features of image modality data.

[0087] The initial modal features of multiple modal data are fused and spliced ​​and the time series features are processed to obtain multiple vocabulary-level multimodal fusion features, including:

[0088] Fusing multiple vocabulary-level text features and multiple vocabulary-level audio features to obtain multiple vocabulary-level audio-text fusion features;

[0089] Multiple lexical-level audio-text fusion features are concatenated with multiple lexical-level expression features to obtain multiple initial lexical-level multimodal fusion features;

[0090] The multiple initial lexical level multimodal fusion features are processed with time series features to obtain multiple lexical level multimodal fusion features.

[0091] It should be noted that the initial modal features include: multiple lexical-level text features of text modal data, multiple lexical-level audio features of audio modal data, and multiple lexical-level expression features of image modal data. Among them, lexical-level text features are features related to a single word extracted from text modal data. These lexical-level text features may include the semantic meaning, emotional tendency, position in the sentence, etc. of the word, and are used to analyze the language structure and content in the text. Lexical-level audio features are relevant speech features corresponding to the words in the text extracted from audio modal data. These lexical-level audio features usually correspond to the voice changes when the speaker says a certain word in the text, and may include the pitch, volume, speech speed, etc. of the voice corresponding to the word, and are used to analyze the emotional expression in the voice. Lexical-level expression features are relevant expression features corresponding to the words in the text or audio extracted from image modal data. These lexical-level expression features usually correspond to the expression changes when the speaker says a certain word, and may include eye movements, mouth shapes, eyebrow positions, etc., and are used to analyze the non-verbal emotional expression of the character.

[0092] Specifically, first, multiple vocabulary-level text features and multiple vocabulary-level audio features are fused, such as weighted summation, and the fused features are called vocabulary-level audio-text fusion features; then, the vocabulary-level audio-text fusion features are concatenated with the vocabulary-level expression features to form initial vocabulary-level multimodal fusion features; finally, the initial vocabulary-level multimodal fusion features are processed with time series features to obtain vocabulary-level multimodal fusion features by capturing dynamic changes and dependencies over time.

[0093] It can be understood that by processing the vocabulary-level text features, vocabulary-level audio features and vocabulary-level expression features, the vocabulary-level multimodal fusion features are obtained, which integrate the information of the three modalities of text, audio and image, and provide comprehensive input for subsequent emotion recognition.

[0094] In a possible implementation, multiple vocabulary-level text features and multiple vocabulary-level audio features are fused to obtain multiple vocabulary-level audio-text fusion features, including:

[0095] The contextual relationship of multiple lexical-level text features is captured through the self-attention mechanism, and multiple lexical-level text modal temporal features corresponding to the multiple lexical-level text features are obtained;

[0096] The contextual relationship of multiple lexical-level audio features is captured through the self-attention mechanism to obtain multiple lexical-level audio modal temporal features corresponding to the multiple lexical-level audio features;

[0097] The windows are smoothly moved according to the preset number of vocabulary windows and the vocabulary order in the text modal data to obtain contextual vocabulary level text features corresponding to the multiple vocabulary level text features and contextual vocabulary level audio features corresponding to the multiple vocabulary level audio features;

[0098] According to multiple vocabulary-level text features, multiple vocabulary-level audio features, multiple contextual vocabulary-level text features, and multiple contextual vocabulary-level audio features, audio collaborative features of multiple text modalities and text collaborative features of multiple audio modalities are calculated;

[0099] Multiple lexical-level text modal temporal features, multiple lexical-level audio modal temporal features, multiple audio collaborative features of text modalities, and multiple audio collaborative features of text modalities are added and averaged to obtain multiple lexical-level audio-text fusion features.

[0100] It should be noted that the self-attention mechanism is a mechanism for capturing the relationship between elements in the input sequence. It allows the model to focus on other elements in the sequence when processing each element, thereby capturing global contextual information.

[0101] The word-level text modality temporal features are features that reflect the information of the words themselves and the contextual relationship between words in the text modality data after processing the word-level text features through the self-attention mechanism. The word-level audio modality temporal features are features that reflect the contextual relationship between the word audio information itself and the word audio in the audio modality data after processing the word-level audio features through the self-attention mechanism. The context word-level text features are text features that reflect the context words in the text modality data. The context word-level audio features are audio features that reflect the context word audio in the audio modality data. The audio synergy features of the text modality are features that reflect the synergy relationship between words in the text modality data and the corresponding context audio in the audio modality data, calculated by combining the word-level text features and the context word-level audio features. The text synergy features of the audio modality are features that reflect the synergy relationship between audio in the audio modality data and the corresponding context words in the text modality data, calculated by combining the audio-level text features and the context word-level text features.

[0102] Specifically, (1) the self-attention mechanism is used to obtain vocabulary-level text features. The query vector , key vector , value vector :

[0103] ;

[0104] ;

[0105] ;

[0106] in, , , is the weight matrix corresponding to the vocabulary level text features. In this embodiment, the weight matrix , , No specific limitation.

[0107] (2) Using the attention mechanism to obtain vocabulary-level audio features The query vector , key vector , value vector :

[0108] ;

[0109] ;

[0110] ;

[0111] in, , , is the weight matrix corresponding to the vocabulary level audio feature. In this embodiment, the weight matrix , , No specific restrictions. Vocabulary level audio features Dimensions and vocabulary-level text features The dimensions are consistent.

[0112] (3) For text modal data, the contextual relationship of multiple word-level text features is captured through the self-attention mechanism, that is, the contextual relationship between words is obtained, and the attention vector of the word in the sentence is calculated as the word-level text modal temporal feature. ,Right now:

[0113] ;

[0114] in, Represents the modal temporal features of lexical level text, Multiple vocabulary-level text features in a sentence The key vector The bond matrix is ​​composed of Multiple vocabulary-level text features in a sentence The value vector of The value matrix composed of is the key vector Dimension.

[0115] (4) For audio modal data, the contextual relationship of multiple lexical-level audio features is captured through the self-attention mechanism, that is, the contextual relationship between the lexical-corresponding audio is obtained, and the attention vector of the lexical-corresponding audio in the sentence is calculated as the lexical-level audio modal temporal feature of the lexical-corresponding speech. ,Right now:

[0116] ;

[0117] in, Multiple vocabulary-level audio features in a sentence The key vector The bond matrix is ​​composed of Multiple vocabulary-level audio features in a sentence The value vector of The value matrix composed of is the key vector The dimension of and equal.

[0118] (5) Considering the correlation between audio and vocabulary and their context, determine the number of preset vocabulary windows for vocabulary. According to the vocabulary order of sentences in text modal data, smoothly move the windows to select vocabulary-level text features. As the center, determine the contextual vocabulary level text features corresponding to the vocabulary context and contextual word-level audio features The preset number of vocabulary windows is a preset number of vocabulary windows that need to be considered when performing smooth window movement. In this embodiment, the preset number of vocabulary windows may be 4, and the specific value of the preset number of vocabulary windows is not limited.

[0119] (6) Based on multiple vocabulary-level text features , multiple vocabulary level audio features , multiple contextual vocabulary level text features and multiple contextual vocabulary-level audio features , calculate the audio collaborative features of multiple text modalities Text-to-text features with multiple audio modalities .

[0120] (7) By analyzing the temporal features of text modalities at multiple lexical levels , multiple lexical level audio modal temporal features , audio collaborative features of multiple text modalities Text-to-text features with multiple audio modalities Add and average to obtain multiple vocabulary-level audio-text fusion features ,Right now:

[0121] .

[0122] It should be noted that when obtaining multiple vocabulary-level audio-text fusion features Afterwards, multiple vocabulary-level audio-text fusion features are concatenated with multiple vocabulary-level expression features to obtain multiple initial vocabulary-level multimodal fusion features, wherein multiple vocabulary-level audio-text fusion features are concatenated with multiple vocabulary-level expression features according to the following formula:

[0123] ;

[0124] in, Represents the initial vocabulary-level multimodal fusion features of a certain vocabulary, Represents the vocabulary-level audio-text fusion features of the corresponding vocabulary, Represents the lexical level expression features of the corresponding vocabulary.

[0125] It can be understood that this embodiment can make full use of the complementary information between the text modality and the audio modality to improve the accuracy of emotion recognition.

[0126] In a possible implementation, audio collaborative features of multiple text modalities and text collaborative features of multiple audio modalities are calculated based on multiple vocabulary-level text features, multiple vocabulary-level audio features, multiple context vocabulary-level text features, and multiple context vocabulary-level audio features, including:

[0127] Determining a first similarity between the plurality of vocabulary-level text features and the plurality of contextual vocabulary-level audio features according to the plurality of vocabulary-level text features, the plurality of contextual vocabulary-level audio features, and a preset similarity formula;

[0128] Calculate audio collaborative features of multiple text modalities based on multiple contextual vocabulary-level audio features and the first similarity;

[0129] Determining a second similarity between the plurality of vocabulary-level audio features and the plurality of contextual vocabulary-level text features according to the plurality of vocabulary-level audio features, the plurality of contextual vocabulary-level text features, and a preset similarity formula;

[0130] According to multiple contextual vocabulary-level text features and the second similarity, text collaborative features of multiple audio modalities are calculated.

[0131] It should be noted that the preset similarity formula is a formula for calculating the similarity between two features.

[0132] Continuing with the above example, based on multiple vocabulary level text features , multiple vocabulary level audio features , multiple contextual vocabulary level text features and multiple contextual vocabulary-level audio features , calculate the audio collaborative features of multiple text modalities Text-to-text features with multiple audio modalities The specific steps are:

[0133] (1) Calculating audio co-features of text modality :

[0134] First, the vocabulary level text features are calculated according to the preset similarity formula and contextual word-level audio features The first similarity sim( , ),Right now:

[0135] ;

[0136] in, is the vocabulary level text feature The query vector is is the contextual word-level audio feature The key vector of .

[0137] Then, the audio co-features of the text modality are calculated ,Right now:

[0138] ;

[0139] in, is the key vector The dimension of is the contextual word level audio feature The value vector of Where j is the position order of the contextual vocabulary level audio features in the sentence fragment selected by the vocabulary window, j=1,2,…,n, and n represents the number of preset vocabulary windows.

[0140] (2) Calculating textual collaborative features of audio modality :

[0141] First, the vocabulary-level audio features are calculated according to the preset similarity formula Word-level text features with context The second similarity sim( , ),Right now:

[0142] ;

[0143] in, is the vocabulary level audio feature The query vector is is the contextual vocabulary level text feature The key vector of .

[0144] Then, the text co-features of the audio modality are calculated ,Right now:

[0145] ;

[0146] in, is the key vector The dimension of is the contextual vocabulary level text feature The value vector of The j in is the position order of the contextual vocabulary level text features in the sentence fragment selected by the vocabulary window, j=1,2,…,n, and n represents the number of preset vocabulary windows.

[0147] In a possible implementation, multiple initial vocabulary-level multimodal fusion features are subjected to temporal feature processing to obtain multiple vocabulary-level multimodal fusion features, including:

[0148] Through the bidirectional long short-term memory network Bi-LSTM model, the context feature extraction is performed on the initial word-level multimodal fusion feature at each moment, and the forward hidden layer feature and backward hidden layer feature corresponding to the initial word-level multimodal fusion feature at that moment are obtained;

[0149] The forward hidden layer features and the backward hidden layer features are concatenated to obtain the lexical level multimodal fusion features corresponding to the initial lexical level multimodal fusion features at that moment.

[0150] It should be noted that in the video data to be tested, there is a strong correlation between the voice, language, and expression of the speaker at the previous and next moments. Therefore, when performing emotion recognition on the video data to be tested, it is necessary to fully consider the changes in the voice, language, and expression at the previous and next moments. The bidirectional long short-term memory network Bi-LSTM model is a recurrent neural network that contains two independent LSTM layers, one LSTM layer for processing the forward input sequence, and one LSTM layer for processing the reverse input sequence. The Bi-LSTM model can capture the previous and next information in the input sequence, thereby more accurately capturing the features of each moment. The forward hidden layer features reflect the information in the video from the start moment to the current moment; the backward hidden layer features reflect the information in the video from the current moment to the end moment.

[0151] Specifically, following the above example, the Bi-LSTM model is used to process forward and backward bidirectional information at the same time to obtain the initial vocabulary-level multimodal fusion features at each moment. The forward hidden layer features and the backward hidden layer features . The forward hidden layer features and the backward hidden layer features Splice and get ,Will As the vocabulary-level multimodal fusion features of the video at time t.

[0152] It can be understood that this embodiment can more accurately understand the comprehensive information of each word in text and audio modalities and generate word-level multimodal fusion features by using the Bi-LSTM model to capture the context information of the initial word-level multimodal fusion features.

[0153] In a possible implementation, obtaining initial modal features of multiple modal data contained in the video data to be tested includes:

[0154] Decomposing the video data to be tested to obtain text modality data, audio modality data and image modality data in the video data to be tested;

[0155] Feature extraction is performed on text modal data, audio modal data and image modal data respectively to obtain multiple lexical level text features of text modal data, multiple lexical level audio features of audio modal data and multiple lexical level expression features of image modal data.

[0156] It should be noted that the video data to be tested is the video data that needs to be emotion recognized. The video data to be tested includes information of multiple modalities, such as audio, image, and text. By analyzing the information of multiple modalities of audio, image, and text in the video data to be tested, the emotional content in the video data to be tested can be identified. Modal data is media data in different forms in video data, namely audio modal data (sound, background music, etc.), image modal data (picture, face, etc.) and text modal data (subtitles, bullet screen, etc.).

[0157] Text modal data is text information extracted from the video data to be tested, such as subtitles, dialogues, etc.; text modal data is used to analyze the language content in the video. Audio modal data is audio information extracted from the video data to be tested, such as voice, background music, etc.; audio modal data is used to analyze the voice content, emotional tone, etc. in the video. Image modal data is image information extracted from the video data to be tested, such as character expressions and movements; image modal data is used to analyze the non-language content in the video, such as the emotions expressed by facial expressions and movements.

[0158] The initial modal features include multiple lexical-level text features of text modal data, multiple lexical-level audio features of audio modal data, and multiple lexical-level expression features of image modal data. Among them, lexical-level text features are features related to a single word extracted from text modal data. These lexical-level text features may include the semantic meaning, emotional tendency, position in the sentence, etc. of the word, and are used to analyze the language structure and content in the text. Lexical-level audio features are relevant speech features corresponding to the words in the text extracted from audio modal data. These lexical-level audio features usually correspond to the voice changes when the speaker says a certain word in the text, and may include the pitch, volume, speech speed, etc. of the voice corresponding to the word, and are used to analyze the emotional expression in the voice. Lexical-level expression features are relevant expression features corresponding to the words in the text or audio extracted from image modal data. These lexical-level expression features usually correspond to the expression changes when the speaker says a certain word, and may include eye movements, mouth shapes, eyebrow positions, etc., and are used to analyze the non-verbal emotional expression of the character.

[0159] It should be noted that the word-level speech feature refers to the feature of the audio segment aligned with the text modality data on the time axis. The word-level expression feature refers to the feature of the image frame aligned with the text modality data or the audio modality data on the time axis.

[0160] In a possible implementation, feature extraction is performed on the text modality data to obtain multiple vocabulary-level text features in the text modality data, including:

[0161] Performing text segmentation processing on the text modality data through the word vector dictionary to obtain multiple words in the text modality data;

[0162] The word embedding feature processing of multiple words in the text modal data is performed through the pre-trained language representation model to obtain multiple word-level text features.

[0163] It should be noted that the word vector dictionary is a data structure used to map words in a text to a fixed-dimensional vector space. Text segmentation is the process of splitting text modal data into individual words. The pre-trained language representation model (Bidirectional Encoder Representations from Transformers, BERT) is a model that can learn the potential representation of text data and has been trained on large-scale text data. Word embedding feature processing is the process of converting words in text modal data into fixed-dimensional feature vector representations. The semantic relationship between words can be captured through word embedding feature processing.

[0164] Specifically, first, the word vector dictionary is used to perform word segmentation on the text modal data, splitting the text modal data into multiple words to facilitate the subsequent extraction of text features. Then, the BERT model is used to perform word embedding feature processing on the words obtained after word segmentation to obtain word-level text features. These word-level text features reflect the semantic relationship between words and can be used for subsequent tasks such as sentiment analysis.

[0165] In a possible implementation, feature extraction is performed on the audio modality data to obtain multiple vocabulary-level audio features in the audio modality data, including:

[0166] Marking the start time and the end time of multiple words in the text modal data in the audio modal data to obtain first time marking information of the multiple words in the audio modal data;

[0167] Segmenting the audio modality data according to the first time tag information to obtain a plurality of audio segment data corresponding to the plurality of words;

[0168] The automatic speech recognition model is used to extract features from multiple audio clip data to obtain multiple vocabulary-level audio features.

[0169] It should be noted that the start time and the end time are the specific time when each word in the audio modal data starts and ends. The first time marking information is the information of the start time and the end time of the audio segment that aligns the words in the text modal data with the audio modal data on the time axis. The boundary of the audio segment in the audio modal data can be determined by the start time and the end time; the text content and the audio content can be associated by these first time marking information. The audio segment data is the audio segment corresponding to the words in the text modal data on the time axis, which is segmented from the audio modal data according to the time marking information.

[0170] The automatic speech recognition model is a deep learning model for speech recognition. The automatic speech recognition model extracts features from multiple audio clip data to obtain multiple vocabulary-level audio features. In this embodiment, the automatic speech recognition model can be a trained Wav2vec 2.0 model.

[0171] Specifically, first, determine the start time and end time of each word in the audio modal data summarized by the text modal data, that is, the first time tag information; then, segment the audio segment data corresponding to each word in the text modal data from the audio modal data according to the first time tag information; finally, use the automatic speech recognition model to extract features from the audio segment data, so as to obtain the vocabulary-level audio features corresponding to each word in the text modal data. And by adjusting the parameters of the automatic speech recognition model, the dimension of the vocabulary-level audio features is made the same as the dimension of the vocabulary-level text features. These vocabulary-level audio features are used to provide information on speech content, emotional intonation, etc., and are used to analyze emotional expression in speech.

[0172] In a possible implementation, feature extraction is performed on the image modality data to obtain multiple vocabulary-level expression features in the image modality data, including:

[0173] Marking the start time and the end time of the plurality of words in the text modality data in the image modality data to obtain second time marking information of the plurality of words in the image modality data;

[0174] Segmenting the image modality data according to the second time tag information and the preset acquisition frequency to obtain a plurality of segmented images corresponding to the plurality of words;

[0175] Performing face recognition on multiple segmented images through a neural network model to obtain multiple face regions in the multiple segmented images;

[0176] Extracting image features of multiple face regions to obtain multiple image features in multiple segmented images;

[0177] The multiple image features corresponding to the multiple words are summed up and averaged to obtain multiple word-level expression features.

[0178] It should be noted that the start time and the end time are the specific times when each word in the image modal data starts and ends. The second time mark information is information that aligns the words in the text modal data with the images in the image modal data on the time axis. The text content and the image content can be associated through these second time mark information. The preset acquisition frequency is the number of image frames collected per second set when collecting the segmented image. Among them, the preset acquisition frequency can be set to one frame.

[0179] Specifically, first, determine the start time and end time of each word in the image modal data summarized by the text modal data, that is, the second time tag information; secondly, segment the image modal data corresponding to each word in the text modal data according to the second time tag information and the preset acquisition frequency; then perform face recognition on the image modal data through a neural network model (such as the Inception-ResnetV2 model) to identify the face area in the image; then, extract image features from the identified face area to obtain image features, which can describe the local and global features of the face; finally, sum and average multiple image features to obtain multiple vocabulary-level expression features. Vocabulary-level expression features are relevant expression features corresponding to words in text or audio extracted from image modal data. These vocabulary-level expression features can be used to analyze the non-verbal emotional expressions of characters.

[0180] In a possible implementation, a self-attention weighted calculation is performed on multiple vocabulary-level multimodal fusion features to obtain a video-level multimodal fusion feature of the video data to be tested, including:

[0181] The self-attention mechanism is used to calculate the multimodal fusion features at the multiple word levels in multiple sentences, and the word weight information corresponding to multiple words in multiple sentences is obtained;

[0182] According to the vocabulary weight information corresponding to the multiple words in the multiple sentences, a weighted average of the multiple vocabulary-level multimodal fusion features in the multiple sentences is performed to obtain a plurality of sentence-level multimodal fusion features corresponding to the multiple sentences;

[0183] Multiple sentence-level multimodal fusion features corresponding to multiple sentences in the video data to be tested are added and averaged to obtain the video-level multimodal fusion features of the video data to be tested.

[0184] It should be noted that the vocabulary weight information is the weight value representing the importance of each vocabulary in the sentence calculated by the self-attention mechanism. These weight values ​​can reflect the contribution of the vocabulary to the entire sentiment analysis. The sentence-level multimodal fusion feature is a feature obtained by weighted averaging multiple vocabulary-level multimodal fusion features. The sentence-level multimodal fusion feature integrates the multimodal information of all vocabulary in the sentence, and can more comprehensively describe the performance of the sentence in the video data. The video-level multimodal fusion feature is a feature obtained by adding and averaging multiple sentence-level multimodal fusion features. The video-level multimodal fusion feature integrates the multimodal information of all sentences in the video, and can more comprehensively describe the overall content and emotional expression of the video data.

[0185] Specifically, following the above example, we use the self-attention mechanism to integrate multiple vocabulary-level multimodal features in the video. As input, in a single sentence, the weight matrix , , The vector matrix of the multimodal fusion features at the vocabulary level Multiply them together to generate the corresponding matrix Q, matrix K, and matrix V. In this embodiment, the weight matrix , , No specific limitation is given. Using the self-attention formula, the self-attention distribution matrix A is calculated, where the elements in the matrix A are The expression is as follows:

[0186] ;

[0187] in, Indicates the importance of the jth word in the sentence to the ith word; p indicates the number of words in the sentence; is the i-th column of the matrix Q, representing the query vector of the i-th word in the sentence; is the j-th column of the matrix K, representing the key vector of the j-th word in the sentence; Represents the dimension of the vocabulary key vector in the sentence, that is, the number of rows of the matrix K.

[0188] Sum the matrix A by column to get the vocabulary weight information ,Right now: ,in, It indicates the contribution of the jth word in the sentence, that is, the weight value.

[0189] According to the vocabulary weight information corresponding to multiple words in the sentence, multiple vocabulary-level multimodal fusion features in the sentence are weighted averaged to obtain the sentence-level multimodal fusion feature ,Right now:

[0190] ;

[0191] in, Represents vocabulary-level multimodal fusion features.

[0192] The sentence-level multimodal fusion features of each sentence in the video data to be tested By adding and averaging, we can get the video-level multimodal fusion features of the video data to be tested. .

[0193] It can be understood that by performing weighted calculation on the vocabulary-level multimodal fusion features to obtain the corresponding sentence-level multimodal fusion features, and then adding and averaging the sentence-level multimodal fusion features to obtain the video-level multimodal fusion features, the features can more comprehensively describe the overall content and emotional expression of the video data to be tested, so as to make subsequent emotion recognition more accurate.

[0194] In a possible implementation, emotion recognition processing is performed on the video data to be tested according to the video-level multimodal fusion features to obtain the emotion prediction result corresponding to the video data to be tested, including:

[0195] Performing a first linear processing on the video-level multimodal fusion feature through a first linear layer of the multimodal sentiment analysis model to obtain a first linear feature corresponding to the video-level multimodal fusion feature;

[0196] Performing nonlinear correction on the first linear feature by using a first activation function of the multimodal sentiment analysis model to obtain a first corrected feature;

[0197] Performing a second linear processing on the first corrected feature through a second linear layer of the multimodal sentiment analysis model to obtain an initial sentiment prediction result corresponding to the video data to be tested;

[0198] The initial emotion prediction result is normalized by the second activation function of the multimodal emotion analysis model to obtain the emotion prediction result corresponding to the video data to be tested.

[0199] It should be noted that the multimodal sentiment analysis model is a deep learning model that can process multiple modal data and perform sentiment analysis. The multimodal sentiment analysis model can learn the association and complementarity between different modal data, so as to more accurately identify the emotions in the video data.

[0200] In this embodiment, the multimodal sentiment analysis model is a fully connected network including a first linear layer, a second linear layer, a first activation function (ReLU activation function), and a second activation function (Softmax activation function). Among them, the first linear layer is a fully connected layer for linearly transforming the input features. The first linear feature is obtained by performing the first linear processing on the multimodal fusion features at the video level by the first linear layer. The first activation function is a function for nonlinearly correcting the output features of the first linear layer; the activation function can introduce nonlinear factors so that the model can fit more complex functional relationships. The first corrected feature can be obtained by performing nonlinear correction on the first linear feature by the first activation function. The second linear layer is a fully connected layer for further linear transformation of the first corrected feature. The initial sentiment prediction result corresponding to the video data to be tested is obtained by performing the second linear processing on the first corrected feature by the second linear layer. The second activation function is a function for normalizing the initial sentiment prediction result. Normalization can ensure that the value range of the output result is within a reasonable range, which is convenient for subsequent processing and analysis. The sentiment prediction result is obtained by normalizing the initial sentiment prediction result by the second activation function. The emotion prediction result is the output result obtained after emotion recognition is performed on the video data to be tested, and represents the emotion category expressed by the video data to be tested.

[0201] Specifically, following the above example, the input of the first linear layer is the video-level multimodal fusion feature of the video Then, the first activation function ReLU is used for nonlinear correction as the input of the second linear layer. The second linear layer performs the second linear processing, and its output is the initial emotion prediction result. Then, the output of the initial sentiment prediction result Normalization is performed through the second activation function Softmax, namely:

[0202] ;

[0203] in, is the emotion prediction result; N is the number of predicted emotion categories, that is, the initial emotion prediction result Dimensions; is the predicted value of the i-th sentiment category.

[0204] in, is the emotion prediction result, that is, the probability of the predicted emotion category. The position where the maximum probability is located is the predicted emotion number.

[0205] like Figure 2 As shown, Figure 2 FIG. 1 is a flow chart of a multimodal emotion recognition method provided by another embodiment of the present application. Figure 2 In the method, the multimodal emotion recognition method is specifically as follows: (1) multiple vocabulary-level text features and multiple vocabulary-level audio features are fused to obtain multiple vocabulary-level audio-text fusion features; (2) multiple vocabulary-level audio-text fusion features are concatenated with multiple vocabulary-level expression features to obtain multiple initial vocabulary-level multimodal fusion features; (3) multiple initial vocabulary-level multimodal fusion features are processed with time series features to obtain multiple vocabulary-level multimodal fusion features; (4) multiple vocabulary-level multimodal fusion features are weighted with self-attention to obtain video-level multimodal fusion features of the video data to be tested; (5) nonlinear correction processing is performed through the first activation function ReLU of the multimodal emotion analysis model, and normalization processing is performed through the second activation function Softmax of the multimodal emotion analysis model, and finally the emotion prediction result corresponding to the video data to be tested is output.

[0206] It can be understood that, in this embodiment, the multimodal emotion recognition method not only fully considers the synergistic effect brought about by the contextual timing information of the speech modality and the text modality and the correlation between speech and text; it also captures the key information in the sentence through the method of weight calculation and feature construction from the vocabulary level to the sentence level, and enhances the influence of the key information on the sentence as a whole by reasonably allocating weights, thereby improving the accuracy of emotion recognition.

[0207] It should be noted that before using the multimodal sentiment analysis model for sentiment recognition, the multimodal sentiment analysis model needs to be trained. That is, a training set is obtained, wherein the training set includes a plurality of sample video data and sentiment prediction results of the plurality of sample video data; the plurality of sample video data and the sentiment prediction results of the plurality of sample video data are used to train the initial multimodal sentiment analysis model, and a trained multimodal sentiment analysis model is obtained.

[0208] Specifically, the sample video data is divided into 4:1 ratios, where 4 / 5 of the sample video data is used as a training set and 1 / 5 of the sample video data is used as a test set. The loss function uses the cross entropy loss function, that is:

[0209] ;

[0210] in, represents the loss function, N is the number of predicted emotion categories, is the emotion numbered i in the emotion prediction probability vector, represents the true probability of the emotion numbered i in the emotion prediction probability vector, Represents the predicted probability of the emotion numbered i in the emotion prediction probability vector.

[0211] When training the initial modal sentiment analysis model, the weight matrix , , , , , , , , Perform random initialization, set the number of vocabulary windows n, the learning rate lr, and the Bi-LSTM hidden layer dimension, and then use the gradient descent method to iterate and find the optimal parameters. Set the preset threshold and the preset number of iterations. When the loss value is less than the preset threshold or reaches the preset number of iterations, stop model training and save the model, and finally obtain a trained modal sentiment analysis model.

[0212] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0213] Corresponding to a multimodal emotion recognition method in the above embodiment, Figure 3 A structural schematic diagram of a multimodal emotion recognition device provided by an embodiment of the present application is shown. For the sake of ease of explanation, only the parts related to the embodiment of the present application are shown.

[0214] Reference Figure 3 , the multimodal emotion recognition device 2 of this embodiment includes:

[0215] The feature acquisition module 21 is used to acquire initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data and text modal data;

[0216] A first fusion module 22 is used to fuse and splice the initial modal features of multiple modal data and process the temporal features to obtain multiple vocabulary-level multimodal fusion features;

[0217] The second fusion module 23 is used to perform self-attention weighted calculation on multiple vocabulary-level multimodal fusion features to obtain video-level multimodal fusion features of the video data to be tested;

[0218] The emotion recognition module 24 is used to perform emotion recognition processing on the video data to be tested according to the video level multimodal fusion features to obtain the emotion prediction result corresponding to the video data to be tested.

[0219] It can be understood that in this embodiment, the multimodal emotion recognition device 2 of this embodiment obtains the initial modal features of multiple modal data contained in the video data to be tested through the feature acquisition module 21, wherein the modal data includes: audio modal data, image modal data and text modal data; secondly, the initial modal features of the multiple modal data are fused and spliced ​​and the time series feature is processed through the first fusion module 22 to obtain multiple vocabulary-level multimodal fusion features; then, the multiple vocabulary-level multimodal fusion features are self-attention weighted calculated through the second fusion module 23 to obtain the video-level multimodal fusion features of the video data to be tested; finally, the emotion recognition module 24 performs emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain the emotion prediction result corresponding to the video data to be tested. The accuracy of multimodal emotion recognition is improved by fusing the initial modal features of the audio modal data, image modal data and text modal data, and allocating more weight information to the key information of the vocabulary-level multimodal fusion features through self-attention weighted calculation.

[0220] Further, the initial modal features include: a plurality of vocabulary-level text features of text modal data, a plurality of vocabulary-level audio features of audio modal data, and a plurality of vocabulary-level expression features of image modal data. The first fusion module 22 specifically includes:

[0221] A fusion processing submodule, used for fusing a plurality of vocabulary-level text features and a plurality of vocabulary-level audio features to obtain a plurality of vocabulary-level audio-text fusion features;

[0222] A splicing processing submodule is used to splice multiple lexical level audio text fusion features with multiple lexical level expression features to obtain multiple initial lexical level multimodal fusion features;

[0223] The time series processing submodule is used to perform time series feature processing on multiple initial vocabulary level multimodal fusion features to obtain multiple vocabulary level multimodal fusion features.

[0224] Furthermore, the fusion processing submodule specifically includes:

[0225] A first capturing unit is used to capture the contextual relationship of multiple vocabulary-level text features through a self-attention mechanism to obtain multiple vocabulary-level text modal temporal features corresponding to the multiple vocabulary-level text features;

[0226] A second capturing unit is used to capture the contextual relationship of multiple lexical level audio features through a self-attention mechanism to obtain multiple lexical level audio modal time series features corresponding to the multiple lexical level audio features;

[0227] A window translation unit, used to smoothly move the window according to a preset number of vocabulary windows and a vocabulary order in the text modal data, to obtain contextual vocabulary level text features corresponding to a plurality of vocabulary level text features and contextual vocabulary level audio features corresponding to a plurality of vocabulary level audio features;

[0228] A collaborative feature determination unit, configured to calculate audio collaborative features of multiple text modalities and text collaborative features of multiple audio modalities based on multiple vocabulary-level text features, multiple vocabulary-level audio features, multiple contextual vocabulary-level text features, and multiple contextual vocabulary-level audio features;

[0229] The word-level feature determination unit is used to add and average multiple vocabulary-level text modal temporal features, multiple vocabulary-level audio modal temporal features, multiple text modal audio collaborative features, and multiple audio modal text collaborative features to obtain multiple vocabulary-level audio-text fusion features.

[0230] Furthermore, the collaborative feature determination unit is specifically used for:

[0231] Determining a first similarity between the plurality of vocabulary-level text features and the plurality of contextual vocabulary-level audio features according to the plurality of vocabulary-level text features, the plurality of contextual vocabulary-level audio features, and a preset similarity formula;

[0232] Calculate audio collaborative features of multiple text modalities based on multiple contextual vocabulary-level audio features and the first similarity;

[0233] Determining a second similarity between the plurality of vocabulary-level audio features and the plurality of contextual vocabulary-level text features according to the plurality of vocabulary-level audio features, the plurality of contextual vocabulary-level text features, and a preset similarity formula;

[0234] According to multiple contextual vocabulary-level text features and the second similarity, text collaborative features of multiple audio modalities are calculated.

[0235] Furthermore, the timing processing submodule is specifically used for:

[0236] Through the bidirectional long short-term memory network Bi-LSTM model, the context feature extraction is performed on the initial word-level multimodal fusion feature at each moment, and the forward hidden layer feature and backward hidden layer feature corresponding to the initial word-level multimodal fusion feature at that moment are obtained;

[0237] The forward hidden layer features and the backward hidden layer features are concatenated to obtain the lexical level multimodal fusion features corresponding to the initial lexical level multimodal fusion features at that moment.

[0238] Furthermore, the feature acquisition module 21 specifically includes:

[0239] The data decomposition submodule is used to decompose the video data to be tested to obtain text modality data, audio modality data and image modality data in the video data to be tested;

[0240] The feature extraction submodule is used to extract features from text modal data, audio modal data and image modal data respectively, to obtain multiple lexical-level text features of text modal data, multiple lexical-level audio features of audio modal data and multiple lexical-level expression features of image modal data.

[0241] Furthermore, the feature extraction submodule is specifically used for:

[0242] Performing text segmentation processing on the text modality data through the word vector dictionary to obtain multiple words in the text modality data;

[0243] The word embedding feature processing of multiple words in the text modal data is performed through the pre-trained language representation model to obtain multiple word-level text features.

[0244] Furthermore, the feature extraction submodule is also specifically used for:

[0245] Marking the start time and the end time of multiple words in the text modal data in the audio modal data to obtain first time marking information of the multiple words in the audio modal data;

[0246] Segmenting the audio modality data according to the first time tag information to obtain a plurality of audio segment data corresponding to the plurality of words;

[0247] The automatic speech recognition model is used to extract features from multiple audio clip data to obtain multiple vocabulary-level audio features.

[0248] Furthermore, the feature extraction submodule is also specifically used for:

[0249] Marking the start time and the end time of the plurality of words in the text modality data in the image modality data to obtain second time marking information of the plurality of words in the image modality data;

[0250] Segmenting the image modality data according to the second time tag information and the preset acquisition frequency to obtain a plurality of segmented images corresponding to the plurality of words;

[0251] Performing face recognition on multiple segmented images through a neural network model to obtain multiple face regions in the multiple segmented images;

[0252] Extracting image features of multiple face regions to obtain multiple image features in multiple segmented images;

[0253] The multiple image features corresponding to the multiple words are summed up and averaged to obtain multiple word-level expression features.

[0254] Furthermore, the second fusion module 23 specifically includes:

[0255] The vocabulary weight calculation submodule is used to perform self-attention calculation on multiple vocabulary-level multimodal fusion features in multiple sentences through a self-attention mechanism to obtain vocabulary weight information corresponding to multiple words in multiple sentences;

[0256] A sentence-level fusion feature determination submodule is used to perform weighted averaging of multiple lexical-level multimodal fusion features in multiple sentences according to lexical weight information corresponding to multiple words in multiple sentences, so as to obtain multiple sentence-level multimodal fusion features corresponding to the multiple sentences;

[0257] The video-level fusion feature determination submodule is used to add and average multiple sentence-level multimodal fusion features corresponding to multiple sentences in the video data to be tested, so as to obtain the video-level multimodal fusion features of the video data to be tested.

[0258] Furthermore, the emotion recognition module 24 specifically includes:

[0259] A first linear processing submodule, used for performing a first linear processing on the video-level multimodal fusion feature through a first linear layer of the multimodal sentiment analysis model to obtain a first linear feature corresponding to the video-level multimodal fusion feature;

[0260] A correction processing submodule, used for performing nonlinear correction on the first linear feature through a first activation function of the multimodal sentiment analysis model to obtain a first corrected feature;

[0261] A second linear processing submodule, used for performing a second linear processing on the first corrected feature through a second linear layer of the multimodal sentiment analysis model to obtain an initial sentiment prediction result corresponding to the video data to be tested;

[0262] The normalization processing submodule is used to normalize the initial emotion prediction result through the second activation function of the multimodal emotion analysis model to obtain the emotion prediction result corresponding to the video data to be tested.

[0263] It should be noted that the information interaction, execution process and other contents between the modules in the above-mentioned multimodal emotion recognition device 2 are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0264] The present application also provides a terminal device, such as Figure 4 As shown, Figure 4A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. Figure 4 The terminal device 3 of this embodiment includes: a memory 31, a processor 32, and a computer program stored in the memory 31 and executable on the processor 32. When the processor 32 executes the computer program, the steps of any one of the above-mentioned multimodal emotion recognition method embodiments are implemented.

[0265] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0266] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0267] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0268] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0269] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0270] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0271] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0272] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A multimodal emotion recognition method, characterized in that: include: Acquire initial modal features of multiple modal data contained in the video data to be tested, wherein the modal data includes: audio modal data, image modal data, and text modal data; Performing fusion and splicing and temporal feature processing on the initial modal features of the plurality of modal data to obtain a plurality of vocabulary-level multimodal fusion features; Performing self-attention weighted calculation on the plurality of vocabulary-level multimodal fusion features to obtain video-level multimodal fusion features of the video data to be tested; Performing emotion recognition processing on the video data to be tested according to the video-level multimodal fusion features to obtain an emotion prediction result corresponding to the video data to be tested; The initial modality features include: a plurality of vocabulary-level text features of the text modality data, a plurality of vocabulary-level audio features of the audio modality data, and a plurality of vocabulary-level expression features of the image modality data. The fusing, splicing and temporal feature processing of the initial modal features of the plurality of modal data to obtain a plurality of vocabulary-level multimodal fusion features includes: Capturing the contextual relationship of the plurality of lexical-level text features through a self-attention mechanism to obtain a plurality of lexical-level text modal temporal features corresponding to the plurality of lexical-level text features; Capturing the contextual relationship of the plurality of lexical-level audio features through the self-attention mechanism to obtain a plurality of lexical-level audio modal time series features corresponding to the plurality of lexical-level audio features; Performing smooth window movement according to a preset number of vocabulary windows and a vocabulary order in the text modality data to obtain contextual vocabulary-level text features corresponding to a plurality of the vocabulary-level text features and contextual vocabulary-level audio features corresponding to a plurality of the vocabulary-level audio features; Calculating audio collaboration features of multiple text modalities and text collaboration features of multiple audio modalities based on the multiple word-level text features, the multiple word-level audio features, the multiple context word-level text features, and the multiple context word-level audio features; A plurality of the vocabulary-level text modality time series features, a plurality of the vocabulary-level audio modality time series features, a plurality of the audio collaborative features of the text modalities, and a plurality of the text collaborative features of the audio modalities are added and averaged to obtain a plurality of vocabulary-level audio-text fusion features.

2. The multimodal emotion recognition method according to claim 1, wherein: The fusing, splicing and temporal feature processing of the initial modal features of the plurality of modal data to obtain a plurality of vocabulary-level multimodal fusion features further includes: splicing the plurality of the lexical-level audio-text fusion features with the plurality of the lexical-level expression features to obtain a plurality of initial lexical-level multimodal fusion features; Temporal feature processing is performed on the multiple initial vocabulary-level multimodal fusion features to obtain the multiple vocabulary-level multimodal fusion features.

3. The multimodal emotion recognition method according to claim 1, wherein: The step of calculating the audio collaborative features of multiple text modalities and the text collaborative features of multiple audio modalities based on the multiple word-level text features, the multiple word-level audio features, the multiple context word-level text features, and the multiple context word-level audio features includes: Determining a first similarity between the plurality of word-level text features and the plurality of context word-level audio features based on the plurality of word-level text features, the plurality of context word-level audio features, and a preset similarity formula; Calculating audio collaborative features of a plurality of the text modalities based on the plurality of the contextual vocabulary-level audio features and the first similarity; Determining a second similarity between the plurality of word-level audio features and the plurality of context word-level text features based on the plurality of word-level audio features, the plurality of context word-level text features, and a preset similarity formula; Based on the multiple contextual vocabulary-level text features and the second similarity, multiple text collaborative features of the audio modalities are calculated.

4. The multimodal emotion recognition method according to claim 2, wherein: The performing time series feature processing on the plurality of the initial vocabulary-level multimodal fusion features to obtain the plurality of the vocabulary-level multimodal fusion features includes: Using a bidirectional long short-term memory network (Bi-LSTM) model, context feature extraction is performed on the initial word-level multimodal fusion feature at each moment to obtain forward hidden layer features and backward hidden layer features corresponding to the initial word-level multimodal fusion feature at that moment; The forward hidden layer features and the backward hidden layer features are concatenated to obtain the vocabulary-level multimodal fusion features corresponding to the initial vocabulary-level multimodal fusion features at that moment.

5. The multimodal emotion recognition method according to claim 2, wherein: The obtaining of initial modal features of multiple modal data contained in the video data to be tested includes: Decomposing the video data to be tested to obtain the text modality data, the audio modality data, and the image modality data in the video data to be tested; Feature extraction is performed on the text modality data, the audio modality data, and the image modality data, respectively, to obtain a plurality of the vocabulary-level text features of the text modality data, a plurality of the vocabulary-level audio features of the audio modality data, and a plurality of the vocabulary-level expression features of the image modality data.

6. The multimodal emotion recognition method according to claim 5, wherein: The feature extraction of the text modality data to obtain a plurality of vocabulary-level text features in the text modality data includes: Performing text segmentation processing on the text modality data using a word vector dictionary to obtain a plurality of words in the text modality data; The word embedding feature processing is performed on multiple words in the text modality data through a pre-trained language representation model to obtain multiple word-level text features.

7. The multimodal emotion recognition method according to claim 6, wherein: The extracting features from the audio modality data to obtain a plurality of vocabulary-level audio features in the audio modality data includes: Marking the start time and the end time of the plurality of words in the text modal data in the audio modal data to obtain first time marking information of the plurality of words in the audio modal data; Segmenting the audio modality data according to the first time tag information to obtain a plurality of audio segment data corresponding to a plurality of words; Feature extraction is performed on the plurality of audio segment data using an automatic speech recognition model to obtain a plurality of vocabulary-level audio features.

8. The multimodal emotion recognition method according to claim 6, wherein: The feature extraction is performed on the image modality data to obtain a plurality of the vocabulary-level expression features in the image modality data, including: Marking the start time and the end time of the plurality of words in the text modality data in the image modality data to obtain second time marking information of the plurality of words in the image modality data; Segmenting the image modality data according to the second time tag information and a preset acquisition frequency to obtain a plurality of segmented images corresponding to a plurality of words; Performing face recognition on the plurality of segmented images using a neural network model to obtain a plurality of face regions in the plurality of segmented images; Extracting image features of the plurality of face regions to obtain a plurality of image features in the plurality of segmented images; The plurality of image features corresponding to the plurality of words are summed up and averaged to obtain a plurality of word-level expression features.

9. The multimodal emotion recognition method according to claim 1, wherein: The performing self-attention weighted calculation on the plurality of vocabulary-level multimodal fusion features to obtain the video-level multimodal fusion features of the video data to be tested includes: Performing self-attention calculation on the plurality of lexical-level multimodal fusion features in the plurality of sentences through a self-attention mechanism to obtain lexical weight information corresponding to the plurality of lexical words in the plurality of sentences; According to the vocabulary weight information corresponding to the multiple words in the multiple sentences, a weighted average of the multiple vocabulary-level multimodal fusion features in the multiple sentences is performed to obtain a plurality of sentence-level multimodal fusion features corresponding to the multiple sentences; A plurality of the sentence-level multimodal fusion features corresponding to a plurality of sentences in the video data to be tested are added and averaged to obtain the video-level multimodal fusion features of the video data to be tested.

Citation Information

Patent Citations

  • Multi-modal lie detection method and device, and equipment

    CN112329746A

  • Video multi-modal emotion recognition method and device based on cross-modal dynamic convolution and computer equipment

    CN114511906A