Non-contact multi-mode decoupling emotion recognition method and device in dialogue scene

By employing a non-contact, multimodal decoupled emotion recognition method in dialogue scenarios, and utilizing a dedicated encoder, a shared feature projector, and a dynamic gating network, the problem of modal feature differences and inconsistencies in multimodal emotion recognition is solved, achieving higher emotion recognition accuracy and robustness.

CN120910683APending Publication Date: 2025-11-07XIDIAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510960172.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies for multimodal emotion recognition in dialogue scenarios suffer from problems such as inefficient fusion due to differences in modal features, inconsistency in information between modalities, modal imbalance, and missing information, resulting in low accuracy and insufficient robustness in emotion recognition.

Method used

A non-contact, multimodal decoupled emotion recognition method is adopted. The original features are extracted by a dedicated encoder, and weighted fusion is performed using a shared feature projector and a multi-path routing module. Combined with a cross-attention module and a dynamic gating network, the modal feature weights are adaptively adjusted. Cross-modal projectors and reconstruction projectors are introduced for training to achieve effective fusion and robustness of intermodal information.

Benefits of technology

It improves the accuracy and robustness of multimodal emotion recognition, enabling accurate emotion classification even in the case of modality loss or imbalance, reducing intermodal interference, and enhancing the precision and training efficiency of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910683A_ABST
    Figure CN120910683A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact multi-modal decoupling emotion recognition method and device in a dialogue scene. The method comprises the following steps: acquiring original data of multiple modals in the dialogue scene; encoding the original data into original features by using a mode-dedicated encoder; projecting the original features by using a shared feature projector to obtain projection features, and performing weighted fusion to obtain shared features; extracting exclusive features from the original features by using a modal-specific expert network, and carrying out weighted fusion on the exclusive features to obtain private features; fusing the shared features and the private features through a cross attention fusion module to obtain multi-modal fusion features; and classifying the multi-modal fusion features by using a first classifier to obtain an emotion recognition result. According to the method, the key problems of high modal feature heterogeneity, inconsistent modal information, unbalanced modal, missing and the like in the field of multi-modal emotion recognition are solved, and the performance and robustness of emotion recognition in a dialogue scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of information technology, and particularly relates to a non-contact multi-modal disentangled emotion recognition method and device in a dialogue scenario. BACKGROUND

[0002] Emotion recognition in a dialogue scenario is an important topic in the field of artificial intelligence, and is widely used in intelligent customer service, human-computer interaction and psychological health analysis fields. Emotion in a dialogue scenario has multi-modal characteristics: human emotion expression is not only reflected in the text content (such as words and semantics) of the dialogue, but also transmitted through non-verbal signals such as voice tone and visual expression. Therefore, multi-modal emotion recognition usually needs to process text, speech and visual three kinds of heterogeneous data at the same time. Because the data form and distribution of text, speech and visual modalities are different, and different modalities may have inconsistency in expressing the same emotion, for example, the text content may express positive emotion, but the voice tone reveals negative. Therefore, how to fuse the information of heterogeneous modalities to accurately recognize emotion is a challenging technical problem. SUMMARY

[0003] In order to solve the above problems existing in the prior art, the present application provides a non-contact multi-modal disentangled emotion recognition method and device in a dialogue scenario.

[0004] The technical problem to be solved by the present application is realized by the following technical scheme: A non-contact multi-modal disentangled emotion recognition method in a dialogue scenario, comprising: obtaining dialogue scenario data; the dialogue scenario data comprises original data of multiple modalities in a dialogue scenario; encoding the original data of each modality into original features by using an encoder special for each modality; projecting the original features of each modality to a shared feature space by using a shared feature projector to obtain projection features and weighting fusion of the projection features, to obtain shared features; wherein the weight for weighting fusion of the projection features of each modality is generated by a first gating network based on the projection features of each modality; extracting exclusive features from the original features of each modality by using an expert network special for each modality, and weighting fusion of the exclusive features of each modality to obtain private features; wherein the weight for weighting fusion of the exclusive features of each modality is generated by a second gating network based on the exclusive features of each modality; fusing the shared features and the private features by using a cross-attention fusion module to obtain multi-modal fusion features; classifying the multi-modal fusion features by using a first classifier to obtain an emotion recognition result; The encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained through pre-training, and in the training process, a cross-modal projector is introduced to extract cross features between different modalities, a reconstruction projector is introduced to reconstruct original features based on the cross features, a shared feature loss is calculated based on the similarity between the cross features and the projected features, a cross reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, and the gradient during back propagation is jointly calculated based on at least the classification loss of the first classifier, the shared feature loss, and the cross reconstruction loss.

[0005] The application also provides a non-contact multi-modal disentangled emotion recognition device in a dialogue scenario, comprising: A first acquisition module is configured to acquire dialogue scenario data, wherein the dialogue scenario data comprises original data of multiple modalities in a dialogue scenario. An encoding module is configured to encode the original data of each modality into original features using an encoder dedicated to each modality. A shared feature generation module is configured to project the original features of each modality into a shared feature space using a shared feature projector to obtain projected features and to perform weighted fusion on the projected features to obtain shared features, wherein the weights for the weighted fusion of the projected features of each modality are generated by a first gating network based on the projected features of each modality. A private feature generation module is configured to extract exclusive features from the original features of each modality using an expert network dedicated to each modality and to perform weighted fusion on the exclusive features of each modality to obtain private features, wherein the weights for the weighted fusion of the exclusive features of each modality are generated by a second gating network based on the exclusive features of each modality. A fusion module is configured to fuse the shared features and the private features through a cross-attention fusion module to obtain multi-modal fusion features. A classification module is configured to classify the multi-modal fusion features using a first classifier to obtain an emotion recognition result. The encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained through pre-training, and in the training process, a cross-modal projector is introduced to extract cross features between different modalities, a reconstruction projector is introduced to reconstruct original features based on the cross features, a shared feature loss is calculated based on the similarity between the cross features and the projected features, a cross reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, and the gradient during back propagation is jointly calculated based on at least the classification loss of the first classifier, the shared feature loss, and the cross reconstruction loss.

[0006] The non-contact multi-modal decoupling emotion recognition method in a dialogue scene provided by the present application can balance multi-modal information through an adaptive routing strategy, reduce the interference of information conflict between modalities on the recognition result, fully retain unique emotional expressions of each modality, reduce the adverse effects brought by direct interaction of heterogeneous modalities, effectively reduce the interference of irrelevant factors between modalities, and make the fused features more accurately reflect multi-modal emotions, so as to realize more accurate emotion classification based on the fused features, improve the fusion efficiency and the accuracy of emotion recognition. Moreover, the present application dynamically adjusts the fusion weights of the features of each modality according to the features of each modality through a gating network, realizes on-demand adaptive allocation, and compared with the fixed weighting method, this dynamic routing method can highlight the most critical modality for different input scenes, improve the utilization efficiency of multi-modal information and improve the classification accuracy.

[0007] The present application will be further described in detail below with reference to the accompanying drawings and the present application. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 is a flowchart of a non-contact multi-modal decoupling emotion recognition method in a dialogue scene provided by an embodiment of the present application; Figure 2 is a schematic diagram of a multi-path routing module provided by the present application; Figure 3 is a schematic diagram of a multi-stage adaptive knowledge distillation training process provided by the present application; Figure 4 is Figure 3 is a schematic diagram of the training process in which a cross-modality projector and a reconstruction projector are introduced to participate in training; Figure 5 is a schematic diagram of a dynamic modality replacement mechanism provided by the present application; Figure 6 is a schematic diagram of a cache manager provided by the present application. DETAILED DESCRIPTION

[0009] The present application will be further described in detail below with reference to the accompanying drawings and the present application.

[0010] The present application will be further described in detail below with reference to the accompanying drawings and the present application. Inefficient fusion caused by modal feature difference: simple early-stage concatenation fusion ignores the huge difference in characteristics between modalities, often leading to the dominance of information from a certain modality in the fusion result, while useful information from other modalities is submerged or irrelevant noise is introduced. Without proper mechanisms to distinguish between modal private information and shared information, the information from different modalities in the fusion feature space may conflict, reducing the accuracy of emotion recognition.

[0011] Insufficient solution to the problem of inconsistent information between modalities: different modalities may have inconsistencies when expressing the same emotion, for example, the text content may express positive emotion, but the voice tone reveals negative emotion. Although existing attention fusion models can highlight relevant information to some extent, they lack decoupled extraction of modal-specific information and common information, and cannot guarantee that the model can capture both "common" emotional cues and individual modal "individual" cues, so the processing of some conflicting information is still not robust enough.

[0012] Modal imbalance and lack of robustness: in practical applications, the text modality often provides the most direct and explicit emotional cues, while the voice and visual modalities provide auxiliary means. However, existing multi-modal models fail to consider using strong modalities to guide weak modalities to learn, resulting in insufficient feature extraction for voice and visual modalities. When encountering modal missing (e.g., no image or no audio), the performance of most models will decrease significantly.

[0013] Therefore, in view of the key problems of modal feature heterogeneity, modal information inconsistency, modal imbalance and missing in the field of multi-modal emotion recognition, the embodiment of the present application provides a non-contact multi-modal decoupled emotion recognition method in a dialogue scene to improve the performance and robustness of emotion recognition in a dialogue scene, as shown in Figure 1 The method comprises the following steps: S10, obtaining dialogue scene data, the dialogue scene data comprising original data of multiple modalities in a dialogue scene.

[0014] Among them, the multiple modalities at least include a text modality. Preferably, the original data of the text modality, the voice modality and the visual (video) modality in the dialogue scene can be obtained respectively in this step S10.

[0015] S20, respectively encoding the original data of each modality into original features using an encoder dedicated to each modality.

[0016] Here, the present application sets a dedicated encoder for each modality to map its original data to a homogeneous feature space, so that the original features of different modalities can be directly operated and compared.

[0017] Specifically, the text encoder receives an input text sequence (e.g., a sentence or a dialogue), encodes the text sequence (e.g., a user's sentence or dialogue) into a high-dimensional semantic feature vector, and outputs a corresponding text feature vector. As an example, RoBERTa (Robustly Optimized BERT Approach), which is an improved BERT-based pure Transformer encoder, can be used as the text encoder. The speech encoder receives an audio signal (a dialogue speech segment), captures emotion-related signals such as tone and rhythm in the speech, and outputs a speech feature vector. As an example, data2vec based on Transformer, which is a general self-supervised learning framework suitable for multiple modalities, can be used as the speech encoder. The visual encoder receives a video segment or image sequence (a speaker's facial expression video during a dialogue), processes the input facial video frames or image sequence, and outputs a visual feature vector reflecting the dynamics of facial expressions. As an example, VideoMAE, which is a ViT (Vision Transformer) based video self-supervised learning model inspired by the MAE (Masked Autoencoder) of images, can be used as the visual encoder.

[0018] It should be noted that the above-mentioned encoder is only an example and is not limited to this in practice.

[0019] In this step S20, the modal-specific encoder can focus on learning the most discriminative features within each modality, thereby preserving the information unique to the modality (e.g., tone changes in speech, facial micro-expressions in vision, etc.).

[0020] In addition, in the specific process, after the encoder of each modality extracts the original features, if the encoder output is a sequence feature (e.g., a word vector sequence of text, a frame feature of speech), it can be further converted into a fixed-length vector using global pooling or attention pooling method, to realize consistent alignment with other modalities.

[0021] S30, using a shared feature projector to project the original features of each modality to a shared feature space to obtain projected features and to weight and fuse the projected features to obtain shared features; wherein the weight for weighting and fusing the projected features of each modality is generated by the first gating network based on the projected features of each modality.

[0022] In order to extract the common emotional information between different modalities, the application constructs a shared feature projector with consistent structure and completely shared parameters, which is used to project and map the original features of each modality, and extract shared features that can reflect the commonality of multi-modalities. Since the module uses a unified network structure and parameters between modalities, the feature extraction process remains consistent regardless of the input modality data, ensuring that all modalities have consistent semantics when mapped to the shared feature space. This design effectively improves the consistency and comparability of cross-modal features, making different modalities comparable and compatible in numerical range, which helps to mine and utilize common emotional expression clues between different modalities. Preferably, a multi-layer nonlinear fully connected network can be used to implement the shared feature projector. It is worth mentioning that the shared feature projector with the above characteristics needs to be trained through a specific training method. This specification first explains how to use the already trained model to implement the process of emotion recognition, and then details the specific training method used by the application.

[0023] Assuming that the multi-modalities include audio modalities, text modalities and visual modalities, the original features of each modality are projected into the shared feature space using the shared feature projector, which can be expressed by the formula:

[0024] wherein, represents the shared feature projector, represents the original features of the speech modality, the text modality and the visual modality, respectively, , , are the projection features of the speech, text and visual modalities obtained by using the shared feature projector. This mapping preserves the cross-modality common information while capturing the unique expression of each modality. In this way, the model can better understand and utilize the shared information in multi-modal data, thereby improving the accuracy of subsequent emotion recognition.

[0025] Then, the projection features of the multi-modalities are fused, and the shared features for expressing the common emotional information of the cross-modalities are obtained.

[0026] Generally, the mixed expert model can achieve the purpose of feature fusion. However, in the existing mixed expert model, the inputs of the expert network and the gating network must be completely the same, which limits the flexibility of processing multi-modal features. If this method is directly used, all modal features must be spliced together to be uniformly input into the mixed expert model, but this will cause interference between modal information, especially when the characteristics of data between modalities are quite different, the spliced features may lose the original modal features. To solve this problem, the present application designs a multiple channel routing module (MCR) to integrate information from different modalities to obtain overall shared features.

[0027] Specifically, referring to Figure 2 , the multiple channel routing module includes a sparse gating network (indicated as Gate in the figure) and a weighted fusion unit. The sparse gating network is used to predict a set of weights, and the weighted fusion unit uses the set of weights output by the sparse gating network to perform weighted fusion on the features input into the multiple channel routing module.

[0028] The multiple channel routing module provided by the present application can be represented by the following formula:

[0029] , wherein, indicates the sparse gating network, is the input of the sparse gating network, indicates the i-th expert network, indicates the features output by the multiple channel routing module.

[0030] Based on the above MCR, the present application uses a shared feature projector to project the original features of each modality to a shared feature space and perform weighted fusion to obtain shared features, including: inputting the projected features of each modality into a multiple channel routing module, which internally splices the projected features of multiple modalities to form features , and inputs the features into a sparse gating network (herein referred to as the first gating network) to make it output a set of weights based on the features. Then, the weighted fusion unit performs weighted summation on the set of weights and the projected features of multiple modalities , so as to obtain shared features, is the number of modalities, .

[0031] It can be understood that in the process of fusing to form shared features, the i-th expert network is used by the i-th modality in step S30. ​​​a shared feature projector, and correspondingly, the MCR output of the first gating network i.e. the shared feature obtained in step S30.

[0032] Thus, by introducing a multi-path routing mechanism, the present application utilizes the gating network to dynamically allocate the weight of each expert according to the input feature, allowing the gating network to flexibly and adaptively adjust the feature fusion strategy according to the integrated multi-modal information, thereby flexibly weighting and fusing the output of different modal experts, and providing strong support for processing complex multi-modal emotion recognition tasks.

[0033] S40, respectively using an expert network dedicated to each modality to extract exclusive features from the original features of the modality, and weighting and fusing the exclusive features of each modality to obtain private features; wherein the weight when weighting and fusing the exclusive features of each modality is generated by the second gating network based on the exclusive features of each modality.

[0034] Specifically, in order to capture the unique emotional expression of each modality, the present application designs an independent expert network for each modality to extract private features of the modality to capture the unique emotional expression of each modality. The expert network adopts a multi-layer nonlinear fully connected network, and the parameters of the expert networks of each modality are not shared and are independently trained. Through the special expert network, the model can focus on learning the most recognizable features within each modality, thereby retaining the information unique to the modality (e.g. pitch changes in speech, facial micro-expressions in vision, etc.).

[0035] Then, the exclusive features of each modality are weighted and fused by another MCR to output private features, and the sparse gating network in the MCR is the second gating network in step S40. It can be understood that in the process of forming the private features, the MCR formula is i.e. the first expert network in step S40, and correspondingly, the feature output by the current MCR is the private feature.

[0036] S50, the shared feature and the private feature are fused by a cross-attention fusion module to obtain a multi-modal fusion feature.

[0037] ​Specifically, after obtaining the private and shared features, the application further fuses the private and shared features by using a cross-attention module. The cross-attention module calculates the attention score by letting the features of different modalities serve as the query and key value, thereby discovering the deeper cross-modal association between the private and shared features, deeply fusing the shared features of each modality, fully mining the complementary information between modalities, and forming a unified multi-modal fusion feature representation. Then, the private and shared features after cross fusion are connected or further nonlinearly transformed to obtain the final fusion feature representation (i.e., the multi-modal fusion feature) of the model.

[0038] Specifically, the process of fusing the shared and private features by the cross-attention fusion module can be represented by the following formula:

[0039] wherein, is the private feature obtained in step S40, is the shared feature obtained in step S30, represents the calculation of the cross-attention of and is the multi-modal fusion feature obtained in step S50.

[0040] S60, classifying the multi-modal fusion feature by using the first classifier to obtain an emotion recognition result.

[0041] Specifically, the multi-modal fusion feature obtained in step S50 is input into the pre-trained first classifier to make it output the result of classifying the multi-modal fusion feature (the probability of the multi-modal fusion feature belonging to each possible classification), so as to select the classification corresponding to the highest probability from the classification result as the emotion recognition result. In practice, the first classifier can be a one-layer or multi-layer fully connected network for mapping the fused features to the emotion label space, and then outputting the final emotion recognition result by using the Softmax function, but is not limited thereto.

[0042] ​The application provides a non-contact multi-modal disentangled emotion recognition method in a dialogue scene, which fully retains the unique emotional expression of each modality by decomposing the features of each modality into private features and shared features and extracting and fusing them respectively, while extracting cross-modal common information, effectively reducing the interference of irrelevant factors between modalities, and making the fused features more accurately reflect multi-modal emotions, so that more accurate emotion classification can be realized based on the fused features. Moreover, the application dynamically adjusts the fusion weights of the features of each modality according to the features of each modality through a gating network, realizing on-demand adaptive allocation. Compared with the fixed weighting method, this dynamic routing method can highlight the most critical modality for different input scenes, improve the utilization efficiency of multi-modal information and improve the classification accuracy.

[0043] The training process of the emotion recognition model (including an encoder, a shared feature projector, a first gating network, an expert network, a second gating network, a cross-attention fusion module, and a first classifier) used in the non-contact multi-modal disentangled emotion recognition method in a dialogue scene provided by the application is described below.

[0044] Specifically, in the application, the encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained through pre-training, and in the training process, a cross-modal projector is introduced to extract cross features between different modalities, a reconstruction projector is introduced to reconstruct the original features based on the cross features, a shared feature loss is calculated based on the similarity between the cross features and the projected features, a cross-reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, and the gradient during back propagation is jointly calculated based on at least the classification loss of the first classifier, the shared feature loss, and the cross-reconstruction loss.

[0045] The classification loss can be calculated using a weighted cross-entropy loss function or a focal loss function.

[0046] The shared feature loss and the cross-reconstruction loss are calculated as follows: Referring to FIG. 1, Figure 4 In the forward propagation process, the shared projector projects the original features of each modality into a shared feature space, and then fuses the projected features of each modality through a multi-path routing module to obtain shared features, Figure 4 The other modules in FIG. 1 do not need to be used in the forward propagation. In the back propagation, the cross-modal projector is used to extract cross features between different modalities, and then a pair of cross features extracted from different modalities are combined and input into the reconstruction projector for feature reconstruction, and the reconstructed original features are output.

[0047] Specifically, when back propagation is performed, the original features of the text modality and the visual modality are input into a cross-modal projector of the audio modality (referred to as an audio cross-modal projector), so that the audio modality related features are extracted from the original features of the text modality , and the audio modality related features are extracted from the original features of the visual modality Then, the features and the features are combined and input into a reconstruction projector of the audio modality (referred to as an audio reconstruction projector) to perform feature reconstruction. Similarly, the original features of the audio modality and the visual modality are input into a cross-modal projector of the text modality (referred to as a text cross-modal projector), so that the text modality related features are extracted from the original features of the audio modality , and the text modality related features are extracted from the original features of the visual modality Then, the features and the features are combined and input into a reconstruction projector of the text modality (referred to as a text reconstruction projector) to perform feature reconstruction. Similarly, the original features of the text modality and the audio modality are input into a cross-modal projector of the visual modality (referred to as a visual cross-modal projector), so that the visual modality related features are extracted from the original features of the text modality , and the visual modality related features are extracted from the original features of the audio modality Then, the features and the features are combined and input into a reconstruction projector of the visual modality (referred to as a visual reconstruction projector) to perform feature reconstruction.

[0048] It can be understood that the purpose of extracting cross features is to extract shared collaborative information between modalities, and such feature representation reveals complex relationships that cannot be captured by a single modality.

[0049] Then, a shared feature loss is calculated based on the similarity between the cross features and the projection features, and a cross reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector.

[0050] Specifically, the present application uses a smooth distance to calculate the shared feature loss, and the calculation formula is as follows:

[0051] Wherein, represents the shared feature loss, represents the projection feature of the audio modality, represents the projection feature of the text modality, represents the projection feature of the visual modality, a function representing computing a smooth L1 distance.

[0052] In the present application, the shared feature loss ensures that the model can maintain the effectiveness and consistency of the features during the learning process. The shared feature loss is achieved by calculating the distance between the features, the goal is to minimize the distance between similar features, while maximizing the distance between different features, this loss mechanism improves the discriminability and generalization performance of the features.

[0053] The cross-reconstruction loss calculated in the present application aims to verify the effectiveness of the cross features by reconstructing the input modality data. Specifically, the reconstruction quality is measured by calculating the difference between the reconstructed data and the original data, so as to ensure that the cross features have discriminability by minimizing the cross-reconstruction loss, while preserving the key information of the original data, enhancing the robustness and accuracy of the model. The specific calculation formula of the cross-reconstruction loss is as follows:

[0054] wherein, the cross-reconstruction loss, the original feature of the audio modality output by the encoder of the audio modality, the original feature of the text modality output by the encoder of the text modality, the original feature of the visual modality output by the encoder of the visual modality, the reconstructed feature of the reconstructed feature of the reconstructed feature of

[0055] It can be understood that the gradient of the back propagation in the joint calculation of the classification loss of the first classifier, the shared feature loss and the cross-reconstruction loss in the present application is that the classification loss of the first classifier, the shared feature loss and the cross-reconstruction loss are included as components of the total loss, the way of inclusion includes summation or weighted summation, etc., so as to calculate the gradient of the model training back propagation based on the total loss.

[0056] In one implementation, the present application adopts a multi-stage adaptive knowledge distillation training process. Referring to Figure 3 ​​​The pre-training process can include two stages; in the first stage (stage 1), a text modality-specific encoder (i.e., a text encoder) and a second classifier (classifier 2) are trained, and the gradient of the reverse propagation is calculated based on the classification loss of the second classifier during the training process; in the second stage (stage 2), the trained text encoder and the second classifier are frozen as a teacher network, and a student network is trained based on the knowledge distillation idea, and it is noted that in stage 2, for the model parameter variable x that needs to be partially frozen for knowledge distillation, the gradient blocking can be performed on the variable x, that is, the forward propagation is sg(x) = x, but the gradient of the reverse propagation is 0, that is, “no gradient is passed at this point”.

[0057] The student network includes a modality-specific encoder, a shared feature projector, a first gating network, an expert network, a second gating network, a cross-attention fusion module, and a first classifier (classifier 1). Figure 3 In the method, the module marked as C is composed of a multi-path routing module in which the shared feature projector and the first gating network are located, and the module marked as S is composed of a multi-path routing module in which the expert network and the second gating network are located.

[0058] Correspondingly, in the process of training the student network, the gradient of the reverse propagation is calculated based on at least the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, and the distillation loss; that is, the distillation loss can be further included in the total loss. The distillation loss is calculated based on the difference in classification performance between the first classifier during training and the trained second classifier.

[0059] Specifically, the distillation loss Ldistillation is divided into an inter-class loss Linter-class and an intra-class loss Lintra-class, and the calculation formula is as follows:

[0060]

[0061]

[0062] wherein the positive number is a temperature coefficient for controlling the smoothness, is a prediction matrix output by the student network, is a prediction matrix output by the teacher network, in order to calculate the inter-class loss and the intra-class loss, the present application uses a softmax function to convert the prediction matrix output by the student network and the teacher network into a probability distribution Y . is a training batch, and the subscript is a batch number, ​​is the number of emotion categories, is the number of emotion categories, is a Pearson correlation coefficient calculation function.

[0063] It is worth mentioning that the existing multi-modal model rarely specially designs a mechanism to utilize the strong modal to guide the weak modal learning, resulting in insufficient feature extraction of the voice, vision and other modalities, in the present application, through two-stage pre-training and combining the knowledge distillation technology, the knowledge of the strong modal (text) is fully utilized to improve the feature representation ability of the weak modal (voice, vision), and the modal imbalance problem in the training process is alleviated.

[0064] In an implementation manner, in the process of training the student network, the present application can also incorporate a total reconstruction loss in the total loss, that is, the gradient of back propagation can be calculated based on the classification loss of the first classifier, the shared feature loss, the cross reconstruction loss, the distillation loss and the total reconstruction loss; the total reconstruction loss is calculated based on the multi-modal fusion feature using the SimSiam method.

[0065] Specifically, SimSiam constructs a pair of twin networks, the two networks are called projector and predictor , which have the same structure but different parameters. Through the two networks and a similarity function, a total reconstruction loss function of self-supervised learning can be constructed :

[0066] , wherein, is the reconstruction target, that is, the original feature, is the multi-modal fusion feature, is the similarity function, and the cosine similarity function is usually used. In order to avoid the problem of gradient collapse, SimSiam also uses gradient stopping technology, that is, when calculating the similarity, the gradient of the second parameter of the function will be truncated, so the function in the above total reconstruction loss function represents the function that does not perform gradient propagation on the second parameter of the cosine similarity.

[0067] Therefore, in the training process of the second stage, the total loss used can be calculated as follows:

[0068] , wherein λ1~λ3 are balance factors between losses, represents the total reconstruction loss, represents the distillation loss, represents the classification loss, .

[0069] Therefore, by calculating the overall loss, the gradient when adjusting the model parameters during back propagation is adjusted, so as to obtain a trained overall model by optimizing the model parameters, and thus the pre-trained model can be used for emotion recognition.

[0070] To verify the effectiveness of the multi-stage adaptive knowledge distillation training process, the inventors conducted comparative experiments on two public datasets commonly used in sentiment recognition research, MELD and IEMOCAP. During the experiment, a multi-modal base model without a distillation mechanism was trained, and a multi-modal model with a cross-modal knowledge distillation mechanism was introduced. The training process of the multi-modal model with the cross-modal knowledge distillation mechanism is described in Figure 3 , and the training process of the multi-modal base model without the distillation mechanism is directly on Figure 3 The trainable branch in stage2 is trained. Then, the accuracy (Accuracy) and weighted F1 score (Weighted F1) are used as model performance evaluation indicators, and the experimental results show that after introducing distillation, the average F1 of the model on the MELD dataset is improved by 3-5 percentage points, and on the IEMOCAP dataset, it is improved by 2-4 percentage points. In addition, cross-modal distillation can also speed up the convergence speed of model training. For example, on the MELD dataset, the model without distillation needs about N iterations to converge, and after introducing distillation, it can converge in advance. The experimental results fully show that cross-modal knowledge distillation effectively promotes the complementary fusion of multi-modal information and improves the accuracy and training efficiency of emotion recognition.

[0071] In practical applications, there are often cases of modal imbalance or even modal loss: for example, some dialogues lack video information or audio quality is poor, which leads to insufficient single-modal information to support accurate emotion judgment. These backgrounds make it particularly necessary to design an emotion recognition system that can fully utilize multi-modal information, take into account the differences between each modality, and be robust to incomplete modal input.

[0072] To solve this problem, the present application also proposes a dynamic modal replacement mechanism. Specifically, when obtaining dialogue scene data in step S10, it is determined whether the dialogue scene data contains original data in the full modal; the full modal at least covers the text modal, the speech modal and the visual modal; when the determination result is no, the original data of each modal is encoded into original features by using an encoder dedicated to each modal, and the pre-trained placeholder is used to replace the original features in the missing modal; wherein the pre-trained placeholder is obtained by introducing a trainable placeholder parameter during the training of the student network and training the student network together.

[0073] As Figure 5A trainable parameter vector is created for each modality, and the missing modality is automatically detected during forward propagation and replaced with a trained placeholder, which is optimized together with the network during training. Through end-to-end training, the optimal replacement representation for the missing modality can be learned, maintaining semantic correlation between modalities. Compared with the traditional simple method of zero padding or mean padding, the dynamic modality replacement mechanism used in the application can better capture the semantic correlation between modalities, and can automatically adapt when modalities are missing, avoiding additional training and inference costs for missing modalities. The traditional invasive method needs to modify the structure and parameters of the existing network, resulting in an increase in model complexity and difficulty in integrating with existing networks. The method of the application enhances the existing fusion layer in a non-invasive manner, and can directly adapt to existing networks that meet the conditions without modifying the original network, seamlessly integrating with existing multi-modal fusion networks, improving their adaptability to missing modalities, and significantly improving their robustness under incomplete input.

[0074] To verify the above effects, the inventors designed a comprehensive robustness evaluation experiment for the scenario of single modality missing in multi-modal input. The specific process is as follows: based on the MELD and IEMOCAP datasets, the input situations of missing text, speech and vision modalities are simulated respectively. Then the emotion classification accuracy of the model of the application under different modality missing is evaluated, and compared with the same model without using the modality missing compensation mechanism. The comparison results show that when the speech modality is missing, the accuracy of the model of the application only decreases by about 5% compared with the full modality, and the comparison model decreases by more than 15%; when the text modality is missing, the F1 score of the model of the application can still reach about 80% of the full modality, while the comparison model is less than 60%; when the visual modality is missing, the performance of the model of the application is almost lossless, maintaining the same emotion recognition result as the full modality input. These data fully show that the modality missing self-adaptive mechanism proposed by the application greatly improves the robustness and practicality of the system to incomplete and diversified input scenarios.

[0075] In one embodiment, the method of the present application can further comprise: before encoding the original data of each modality into original features by the encoder dedicated to each modality, calculating a model hash value according to the parameters of the encoder dedicated to each modality, and calculating a data hash value according to the original data of the modality; constructing an identity of the original data according to the model hash value and the data hash value, and determining whether the original features of the same original data have been stored in the video memory according to the identity; if yes, directly obtaining the original features of the original data from the video memory without encoding the original data again; if no, further determining whether the original features of the same original data have been stored in the memory according to the identity; if yes, directly obtaining the original features of the original data from the memory without encoding the original data again; if no, further determining whether the original features of the same original data have been stored in the disk according to the identity; if yes, directly obtaining the original features of the original data from the memory without encoding the original data again; if no, continuing to encode the original data. Then, after encoding the original data of each modality into original features by the encoder dedicated to each modality, storing the original features and the identity of the original data into the video memory, the memory and the disk in sequence.

[0076] In practice, the above functions can be realized by designing a cache manager. By intercepting the forward propagation process and calling the cache manager, the calculation redundancy in the training of the deep learning model is significantly reduced, and the training efficiency is improved.

[0077] Specifically, in the cache manager, the GPU memory (graphics processing unit memory) is used as a first-level cache to store the recently used feature tensors and provides the fastest access speed; the memory cache is used as a second-level cache and has a larger storage capacity than the first-level cache but slower access speed, and requires additional CPU-to-GPU data transmission; the disk cache provides the largest storage capacity but the slowest access speed. When the device cache is not hit, the cache manager checks the memory cache, and if the memory cache is also not hit, attempts to load the cache data from the disk. When calculating the hash value, the model parameters of the encoder can be serialized by using SafeTensors to convert them into binary data, and then spliced with the original data, and then processed by an algorithm to generate a hash value. SafeTensors is a high-efficiency and safe tensor serialization file format, which has the advantages of fast reading and writing and avoiding memory copying, can effectively improve the access efficiency of large-scale feature data, and reduce the security risks caused by serialization.

[0078] When writing to the cache, the cache manager needs to record three parameters: the modal corresponding to the encoder, the original data, and the encoded original features. A unique cache hash value can be generated by calculating the hash value of the model parameters of the encoder and the original data respectively and splicing them together, and used as the cache key. The encoded original features are stored as cache values in the cache. The cache manager will store them in the order of video memory, memory, and disk. The cache values of the disk need to be serialized before storage. If the cache space is insufficient, the cache manager will clear the least recently used cache item according to the least recently used strategy to free up space to store new cache data.

[0079] When reading the cache, the cache manager needs to receive two parameters: the modal corresponding to the encoder and the original data. Similarly, a unique cache hash value can be obtained by calculating the hash value of the model parameters of the encoder and the original data respectively and splicing them together. Then, according to this hash value, the cache key is found, and the video memory, memory, and disk cache are checked in hierarchical order until the matching cache key is found. If the cache hits, the cache value is returned to the caller, and the forward propagation process of the encoder is skipped. If the cache misses, the original data is passed to the encoder for forward propagation, and the calculated feature vector result is written to the cache. To further improve cache efficiency, the cache manager can also use a lazy loading mechanism, that is, only when needed will the encoded original features be loaded from the disk to the memory or device memory, to avoid unnecessary data transmission.

[0080] It can be understood that traditional models also have bottlenecks in training efficiency and deployment convenience, such as model parameter redundancy, repeated calculation, and other problems, making it difficult to meet the needs of large-scale practical applications. In comparison, the multi-level feature cache component of the present application can significantly reduce the computational redundancy in deep learning model training, improve training efficiency, especially in scenarios involving large frozen networks. This technology not only speeds up the model training process, but also optimizes the utilization of computing resources, providing a new solution for efficient training and inference of deep learning models.

[0081] In summary, the present application is directed to text, speech and vision three modalities, respectively setting special feature encoder, mapping the original multi-modal data to the homogeneous feature space, fully mining the semantic information of each. Secondly, through the "private-shared feature separation" mechanism, the unique emotional features of each modality are extracted, and the common emotional information between different modalities is extracted by using the projection network with parameter sharing, realizing the collaborative modeling of modality specificity and modality commonality. Thirdly, the multi-path routing and cross attention mechanism is introduced, realizing the adaptive dynamic weighting and deep interaction of each modality information in the feature fusion stage, and enhancing the comprehensive discrimination ability of the model to complex emotional signal. Finally, the two-stage distillation training strategy is adopted, the knowledge of single modality strong classifier and the advantages of multi-modal joint modeling are combined in the training process, and the generalization performance and robustness of the system in the actual scene of modal missing and incomplete information are effectively improved. Through the above steps, the present application solves the problems of reduced accuracy and insufficient generalization ability caused by modal expression difference, feature redundancy and modal missing in traditional multi-modal emotion recognition, significantly improves the practicability and application value of the multi-modal emotion recognition system, and the specific beneficial effects mainly reflect in the following aspects: (1) Layered decoupled feature extraction improves fusion effectiveness The present application decomposes the features of each modality into private features and shared features, and extracts and fuses them respectively, fully retains the unique emotional expression of each modality, extracts cross-modal common information, effectively reduces the interference of irrelevant factors between modalities, and makes the fused features more accurately reflect multi-modal emotion. Experimental results show that the model with private and shared modules simultaneously outperforms the single structure, and both types of features are indispensable for improving the emotion recognition task.

[0082] (2) Multi-path routing dynamic fusion, adaptive allocation of modal weight The multi-path routing module (MCR) is introduced, which dynamically adjusts the fusion weight of each modality feature through a gating network, realizing on-demand adaptive allocation. Compared with the fixed weighting method, dynamic routing can highlight the most critical modality for different input scenarios, improve the utilization efficiency of multi-modal information, and play a key role in improving the overall model performance.

[0083] (3) Cross-modal knowledge distillation, enhance the weak modality representation The present application adopts a two-stage training strategy to transfer the rich knowledge contained in the text and other dominant modalities to the multi-modal model, improves the discrimination ability of the speech and visual branches through distillation loss, effectively alleviates the performance bottleneck caused by modal imbalance. Even in the text missing scene, the distilled speech / visual model can also achieve high emotion recognition accuracy, and the overall performance is more balanced and robust.

[0084] (4) Feature reconstruction regularization and placeholder replacement mechanism, improve the robustness of the model The introduction of the cross-reconstruction loss ensures high-fidelity preservation of key information of each modality by the fused features. The dynamic modality missing simulation and trainable placeholder mechanism enable the model to automatically adapt to any missing combination of modalities, supporting complete and incomplete inputs without the need for additional training of specialized models, greatly simplifying the actual deployment process. In testing, the method exhibits graceful performance degradation when a single modality is missing, maintaining high accuracy and demonstrating strong fault tolerance and generalization ability.

[0085] (5) Efficient engineering, supporting large-scale applications The pre-training model, efficient projection network and gating routing mechanism are fused, and each module can run in parallel on a GPU. For the repeated calculation problem in the distillation training process, a multi-level feature caching strategy is introduced, and the output of the frozen teacher model is cached to the video memory, memory or disk, significantly reducing unnecessary calculations and improving the overall training and inference speed, thereby meeting the efficiency requirements in actual scenarios.

[0086] In summary, the multi-modal emotion recognition method proposed in the present application not only significantly improves the effectiveness and robustness of multi-modal information fusion, but also greatly optimizes the engineering implementation efficiency of the system. This has important popularization and application value for emotion recognition application scenarios such as dialogue systems, intelligent customer service, human-computer interaction, which require high accuracy and high robustness.

[0087] The method provided by the embodiment of the present application can be applied to an electronic device. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc. Herein, it is not limited, and any electronic device that can implement the present application belongs to the protection scope of the present application.

[0088] Corresponding to the above-mentioned non-contact multi-modal disentangled emotion recognition method in a dialogue scenario, the present application also provides a non-contact multi-modal disentangled emotion recognition device in a dialogue scenario, comprising: A first acquisition module is configured to acquire dialogue scenario data, wherein the dialogue scenario data comprises original data of multiple modalities in a dialogue scenario; An encoding module is configured to encode the original data of each modality into original features using a modality-specific encoder; A shared feature generation module is configured to project the original features of each modality into a shared feature space using a shared feature projector to obtain projection features and perform weighted fusion on the projection features to obtain shared features. The weights for the weighted fusion of the projection features of each modality are generated by a first gating network based on the projection features of each modality. The private feature generation module is configured to extract exclusive features from the original features of each modality by using an expert network dedicated to each modality, and to fuse the exclusive features of each modality by weighting to obtain private features; wherein the weight for fusing the exclusive features of each modality is generated by the second gating network based on the exclusive features of each modality; The fusion module is configured to fuse the shared features and the private features by using a cross-attention fusion module to obtain multi-modal fusion features. The classification module is configured to classify the multi-modal fusion features by using a first classifier to obtain an emotion recognition result. The encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained by pre-training, and in the training process, a cross-modal projector is introduced to extract cross-features between different modalities, a reconstruction projector is introduced to reconstruct the original features based on the cross-features, a shared feature loss is calculated based on the similarity between the cross-features and the projected features, and a cross-reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, so as to jointly calculate the gradient during back propagation based on at least the classification loss of the first classifier, the shared feature loss, and the cross-reconstruction loss.

[0089] Optionally, the plurality of modalities at least includes a text modality. The pre-training process includes two stages; wherein the first stage trains the encoder dedicated to the text modality and a second classifier, and the gradient of back propagation is calculated based on the classification loss of the second classifier during the training process; the second stage uses the trained encoder dedicated to the text modality and the second classifier as a teacher network, and guides the training of a student network based on the knowledge distillation idea. The student network includes an encoder dedicated to each modality, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier; during the training of the student network, the gradient of back propagation is jointly calculated based on the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, and a distillation loss; wherein the distillation loss is calculated based on the difference in classification performance between the first classifier and the trained second classifier.

[0090] Optionally, during the training of the student network, the gradient of back propagation is jointly calculated based on the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, the distillation loss, and a total reconstruction loss; the total reconstruction loss is calculated by using the SimSiam method based on the multi-modal fusion features.

[0091] Optionally, the apparatus further comprises: a judging module, configured to judge whether the dialogue scene data contains original data in a full modality when the dialogue scene data is acquired, the full modality covering at least a text modality, a voice modality and a visual modality; a replacing module, configured to, when the judging result is no, replace original features in a missing modality with a pre-trained placeholder while encoding original data in each modality into original features with an encoder dedicated to the modality; wherein the pre-trained placeholder is obtained by introducing a trainable placeholder parameter in the process of training the student network and training the student network together.

[0092] Optionally, the apparatus further comprises: a second obtaining module, configured to, before encoding original data in each modality into original features with an encoder dedicated to the modality, calculate a model hash value according to the parameter of the encoder dedicated to the modality, calculate a data hash value according to the original data in the modality, construct an identifier of the original data according to the model hash value and the data hash value, and judge whether original features of the same original data already exist in the video memory according to the identifier; if yes, directly obtain the original features of the original data from the video memory without encoding the original data into features again; if no, further judge whether original features of the same original data already exist in the internal memory according to the identifier; if yes, directly obtain the original features of the original data from the internal memory without encoding the original data into features again; if no, further judge whether original features of the same original data already exist in the disk according to the identifier; if yes, directly obtain the original features of the original data from the internal memory without encoding the original data into features again; if no, continue to encode the original data into features; a storing module, configured to, after the encoding module encodes original data in each modality into original features with an encoder dedicated to the modality, store the original features of the original data and the identifier into the video memory, the internal memory and the disk in sequence.

[0093] It should be noted that, for the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment, and the same or similar beneficial effects can be achieved.

[0094] It should be noted that the terms "first", "second", and the like, are used to distinguish like objects, and are not necessarily used to describe a particular sequential or chronological order. It is to be understood that such terms as used herein are interchangeable under appropriate circumstances such that the embodiments of the application described herein are capable of operation in other sequences than described or otherwise illustrated herein. The embodiments described in this example implementation are not meant to represent all implementations consistent with the application. On the contrary, they are merely examples of apparatus and methods consistent with some aspects of the application.

[0095] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification.

[0096] Although the present application is described herein in conjunction with various embodiments, those skilled in the art, with reference to the attached drawings and the disclosure content, can understand and implement other variations of the disclosed embodiments in the implementation of the claimed application. In the description of the present application, the word "comprising" does not exclude other components or steps, "one" or "an" does not exclude a plurality, and "plurality" means two or more, unless otherwise expressly specified. In addition, some measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0097] The above is a further detailed description of the present application in conjunction with specific preferred embodiments, and cannot be considered as limiting the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or substitutions can be made, which should be considered as falling within the scope of protection of the present application.

Claims

1. A method for non-contact multi-modal disentangled emotion recognition in a dialogue scenario, characterized in that, The method comprises: obtaining dialogue scene data; the dialogue scene data comprises original data of multiple modalities in a dialogue scene; encoding the original data of each modality into original features by using an encoder dedicated to each modality; projecting the original features of each modality into a shared feature space by using a shared feature projector to obtain projected features and performing weighted fusion on the projected features to obtain shared features; wherein the weights for performing weighted fusion on the projected features of each modality are generated by a first gating network based on the projected features of each modality; extracting exclusive features from the original features of each modality by using an expert network dedicated to each modality, and performing weighted fusion on the exclusive features of each modality to obtain private features; wherein the weights for performing weighted fusion on the exclusive features of each modality are generated by a second gating network based on the exclusive features of each modality; fusing the shared features and the private features by using a cross-attention fusion module to obtain multi-modal fusion features; classifying the multi-modal fusion features by using a first classifier to obtain an emotion recognition result; wherein the encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained by pre-training, and in the training process, a cross-modal projector is introduced to extract cross-features between different modalities, a reconstruction projector is introduced to reconstruct the original features based on the cross-features, a shared feature loss is calculated based on the similarity between the cross-features and the projected features, a cross-reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, and the gradient during back propagation is calculated based on at least the classification loss of the first classifier, the shared feature loss, and the cross-reconstruction loss.

2. The non-contact multi-modal disentangled emotion recognition method in a dialogue scenario according to claim 1, characterized in that, The multiple modalities at least include a text modality; The pre-training process includes two stages; wherein the first stage trains the encoder dedicated to the text modality and a second classifier, and the gradient during back propagation is calculated based on the classification loss of the second classifier; the second stage uses the trained encoder dedicated to the text modality and the second classifier as a teacher network, and trains a student network based on the knowledge distillation idea; The student network includes an encoder dedicated to each modality, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier; during the training of the student network, the gradient during back propagation is calculated based on at least the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, and a distillation loss; wherein the distillation loss is calculated based on the difference in classification performance between the first classifier and the trained second classifier.

3. The non-contact multi-modal disentangled emotion recognition method under dialogue scenario according to claim 2, characterized in that, During the training of the student network, the gradient during back propagation is calculated based on at least the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, the distillation loss, and a total reconstruction loss; the total reconstruction loss is calculated by using the SimSiam method based on the multi-modal fusion features.

4. The method of claim 2, wherein the non-contact multi-modal de-coupled emotion recognition in dialog scene is characterized by, The method further comprises: In the process of obtaining the dialogue scene data, it is determined whether the dialogue scene data contains original data in a full modality; the full modality at least covers a text modality, a speech modality and a visual modality; When the determination result is no, while encoding the original data in each modality into original features by using an encoder special for the modality, a pre-trained placeholder is used to replace the original feature in the missing modality; wherein the pre-trained placeholder is obtained by introducing a trainable placeholder parameter in the process of training the student network and training the student network together.

5. The method of claim 1, wherein, The method further comprises: Before encoding the original data in each modality into original features by using an encoder special for the modality, a model hash value is calculated according to the parameters of the encoder special for the modality, and a data hash value is calculated according to the original data in the modality; an identifier of the original data is constructed according to the model hash value and the data hash value, and it is determined whether the original feature of the same original data already exists in the video memory according to the identifier; if yes, the original feature of the original data is directly obtained from the video memory, and the original data does not need to be encoded again; if no, it is further determined whether the original feature of the same original data already exists in the memory according to the identifier; if yes, the original feature of the original data is directly obtained from the memory, and the original data does not need to be encoded again; if no, it is further determined whether the original feature of the same original data already exists in the disk according to the identifier; if yes, the original feature of the original data is directly obtained from the memory, and the original data does not need to be encoded again; if no, the original data is continuously encoded; After encoding the original data in each modality into original features by using an encoder special for the modality, the original feature of the original data and the identifier are sequentially stored in the video memory, the memory and the disk.

6. A non-contact multi-modal disentangled emotion recognition apparatus in a dialogue scenario, characterized in that, Comprise: A first obtaining module is configured to obtain dialogue scene data; The dialogue scene data comprises original data of multiple modalities in a dialogue scene; An encoding module is configured to encode the original data in each modality into original features by using an encoder special for the modality; A shared feature generation module is configured to project the original features of the modalities to a shared feature space by using a shared feature projector to obtain projection features and perform weighted fusion on the projection features to obtain shared features; wherein the weights for performing weighted fusion on the projection features of the modalities are generated by a first gating network based on the projection features of the modalities; A private feature generation module is configured to extract exclusive features from the original features of each modality by using an expert network special for the modality, and perform weighted fusion on the exclusive features of the modalities to obtain private features; wherein the weights for performing weighted fusion on the exclusive features of the modalities are generated by a second gating network based on the exclusive features of the modalities; A fusion module is configured to fuse the shared features and the private features by using a cross-attention fusion module to obtain multi-modal fusion features; A classification module is configured to classify the multi-modal fusion features by using a first classifier to obtain an emotion recognition result. The encoder, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier are obtained through pre-training, and in the training process, a cross-modal projector is introduced to extract cross-features between different modalities, a reconstruction projector is introduced to reconstruct original features based on the cross-features, a shared feature loss is calculated based on the similarity between the cross-features and the projected features, and a cross-reconstruction loss is calculated based on the similarity between the original features output by the encoder and the original features reconstructed by the reconstruction projector, so that the gradient during back propagation is jointly calculated based on at least the classification loss of the first classifier, the shared feature loss, and the cross-reconstruction loss.

7. The non-contact multi-modal de-coupling emotion recognition apparatus under conversational scenario according to claim 6, characterized in that, The plurality of modalities at least includes a text modality; The pre-training process includes two stages; in the first stage, the encoder and the second classifier dedicated to the text modality are trained, and the gradient during back propagation is calculated based on the classification loss of the second classifier; in the second stage, the trained encoder and the second classifier dedicated to the text modality are used as a teacher network, and a student network is trained based on the knowledge distillation idea; The student network includes an encoder dedicated to each modality, the shared feature projector, the first gating network, the expert network, the second gating network, the cross-attention fusion module, and the first classifier; in the process of training the student network, the gradient during back propagation is jointly calculated based on the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, and a distillation loss; the distillation loss is calculated based on the difference in classification performance between the first classifier and the trained second classifier.

8. The non-contact multi-modal de-coupling emotion recognition apparatus under conversational scenario according to claim 7, characterized in that, In the process of training the student network, the gradient during back propagation is jointly calculated based on the classification loss of the first classifier, the shared feature loss, the cross-reconstruction loss, the distillation loss, and a total reconstruction loss; the total reconstruction loss is calculated based on the multi-modal fusion features using the SimSiam method.

9. The non-contact multi-modal de-coupling emotion recognition apparatus under conversational scenario according to claim 7, characterized in that, The device further comprises: A judgment module is configured to determine whether the dialogue scene data contains original data in a full modality when the dialogue scene data is obtained; the full modality at least covers a text modality, a speech modality, and a visual modality; An alternative module is configured to, when the determination result is negative, replace the original features in the missing modality with a pre-trained placeholder while encoding the original data in each modality into original features using an encoder dedicated to the modality; the pre-trained placeholder is obtained by introducing a trainable placeholder parameter in the process of training the student network and training the student network together.

10. The non-contact multi-modal de-coupling emotion recognition apparatus under conversational scenario as claimed in claim 6, wherein, The device further comprises: The second obtaining module is configured to, before encoding the original data of each modality into original features by using the encoder dedicated to each modality, calculate a model hash value according to the parameters of the encoder dedicated to each modality, calculate a data hash value according to the original data of the modality, construct an identifier of the original data according to the model hash value and the data hash value, and determine whether the original features of the same original data have been stored in the video memory according to the identifier; if yes, directly obtain the original features of the original data from the video memory without further encoding the original data into features; if no, further determine whether the original features of the same original data have been stored in the memory according to the identifier; if yes, directly obtain the original features of the original data from the memory without further encoding the original data into features; if no, further determine whether the original features of the same original data have been stored in the disk according to the identifier; if yes, directly obtain the original features of the original data from the memory without further encoding the original data into features; if no, continue to encode the original data into features; The storage module is configured to, after the encoding module encodes the original data of each modality into original features by using the encoder dedicated to each modality, store the original features and the identifier of the original data into the video memory, the memory and the disk in sequence.

Citation Information

Cited By

  • Multi-modal sentiment analysis system and method based on dynamic routing and feature decoupling

    CN121615016A

  • Multi-modal sentiment analysis system and method based on dynamic routing and feature decoupling

    CN121615016B

  • A multi-modal emotion recognition method based on feature decoupling

    CN122388702A