Self-adaptive multi-expert cooperative multi-modal emotion recognition method and related equipment

Through the adaptive multi-expert collaborative multi-modal emotion recognition method, the high-dimensional Gaussian distribution and self-attention mechanism are used to dynamically adjust the number and weight of experts, solving the problems of uneven modal data quality and noise interference, and improving the accuracy and generalization ability of multi-modal emotion recognition.

CN120524167APending Publication Date: 2025-08-22SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510425728.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing multimodal emotion recognition methods are difficult to effectively deal with when facing uneven modal data quality and noise interference, resulting in poor feature loss and fusion effects.

Method used

Adaptive multi-expert collaborative multi-modal emotion recognition method is adopted to quantify modal uncertainty through high-dimensional Gaussian distribution, dynamically adjust the number and weight of experts, and combine the self-attention mechanism and cascaded cross-attention fusion structure to optimize the modal fusion strategy.

Benefits of technology

It improves the accuracy and generalization ability of multimodal emotion recognition, can better handle data interference and noise in real scenarios, and improves the effect of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524167A_ABST
    Figure CN120524167A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive multi-expert cooperative multi-mode emotion recognition method and related equipment, and the system comprises a feature extraction module which is used for extracting the primary feature representation of three modes of images, voices and characters in the emotion; the uncertainty perception self-adaptive expert distribution mechanism module is used for quantifying the uncertainty level of each modal and obtaining the stable representation of each modal feature; the reliability level emotion decoding module is used for realizing self-adaptive modal fusion based on uncertainty and fusing multi-modal features in sequence according to a reliability priority; and the emotion intensity prediction module is used for performing intensity prediction on the emotion by using a multi-layer perceptron. According to the method, uncertainty in an emotion modal sample is mined and quantified, so that self-adaptive multi-expert cooperative multi-modal emotion recognition based on uncertainty driving is realized; and more stable emotion feature representation is extracted in combination with uncertainty, and a more efficient multi-modal fusion strategy is proposed, so that the accuracy of multi-modal emotion recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal learning and emotion recognition, and in particular to an adaptive multi-expert collaborative multimodal emotion recognition method and related equipment. Background Art

[0002] Emotion plays a crucial role in human communication. Emotion recognition involves analyzing and interpreting multiple modalities, including natural language, facial expressions, and verbal expressions, to infer a person's emotional state, intentions, and experiences. Emotion recognition can help us better understand others' emotions, intentions, and attitudes, fostering effective communication. It can also help us build closer interpersonal relationships and enhance mutual understanding and empathy. Furthermore, emotion recognition can timely identify and predict an individual's emotional state and mental health, enabling timely identification and intervention of potential psychological issues. This provides objective evidence for personalized psychological counseling and psychotherapy, benefiting both physical and mental health. Emotion recognition technology plays a significant role not only at the individual level but also at the organizational and societal levels. With the development of the internet, user-generated online content (such as short videos, product reviews, and social media user sentiment) has generated a vast amount of voice, image, and video data. Emotion recognition of this data can provide insights into user experience and product evaluation. This is particularly valuable for businesses, governments, market research institutions, and other organizations, enabling them to better understand user needs and optimize products. The goal of emotion recognition is to enable computers to automatically identify human emotional responses, helping people better understand, make decisions, communicate, and provide personalized services. Multimodal emotion recognition plays a vital role in various fields, including human-computer interaction, intelligent healthcare, robotics, and intelligent customer service. Through emotion recognition, we can improve the emotional intelligence of machines.

[0003] Most existing multimodal emotion recognition methods aim to investigate multimodal fusion, attempting to integrate multimodal information through various multimodal fusion approaches, such as early fusion, feature fusion, and late fusion. Most work assumes that all modalities in a dataset are of reliable quality. However, in real-world scenarios, modal data often varies. Factors such as background interference, device interference, or partial data loss can affect modal quality. When one item in multimodal data exhibits sample ambiguity, directly participating in multimodal interaction without processing will undermine the contribution of other reliable modalities. Some current work directly discards or reconstructs features from noisy or missing data. However, such approaches result in the loss of original feature data and the presence of feature discrepancies. Therefore, effectively assessing and addressing noise in datasets presents a challenge in multimodal emotion recognition. Summary of the Invention

[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide an uncertainty-driven adaptive multi-expert collaborative multimodal emotion recognition method and related equipment.

[0005] The first technical solution adopted by the present invention is:

[0006] An adaptive multi-expert collaborative multimodal emotion recognition method includes the following steps:

[0007] Acquire multimodal data, extract primary feature representations for each modality, and perform modal alignment by filling in features from different modalities to obtain the correct primary feature representations for each modality. Multimodal data includes three modalities: image, speech, and text.

[0008] Use high-dimensional Gaussian distribution to model the primary feature representation and obtain the mean and variance of the current modal feature; use the variance to quantify the uncertainty level of the current mode, and obtain a stable representation of the primary feature through Gaussian distribution reparameterization;

[0009] The hybrid multi-expert model further extracts reparameterized features, where uncertainty and reparameterized features dynamically determine the number of experts in the multi-expert model and the weight of each expert. The higher the modal uncertainty, the more multi-experts are activated. Finally, the modal features are weighted summed by the selected experts with their weights to obtain a more advanced and reliable latent feature representation.

[0010] The three modal features are vector-concatenated to form a joint representation, and the internal correlation of the joint representation is modeled through the self-attention mechanism to enhance the complementary information between the modalities;

[0011] Based on the uncertainty information, the three modes are sorted from high to low in terms of reliability, and a modality priority queue is constructed;

[0012] According to the modality priority queue, a three-level cascade cross-attention fusion structure is designed to fuse multimodal feature representations and obtain multimodal joint feature representation;

[0013] The intensity of emotion is predicted based on the multimodal joint feature representation to obtain the prediction result.

[0014] Furthermore, the step of acquiring multimodal data and extracting primary feature representations of each modality includes:

[0015] We obtain images, speech, and text files from the sentiment dataset and extract primary features of each modality using the three feature extractors FACET, COVAREP, and BERT. We then design a one-dimensional convolutional network for each modality to compress the features. The feature extraction formula is as follows:

[0016] X v =1DCNN(FACET(x v ))

[0017] X a =1DCNN(COVAREP(x a ))

[0018] X l =1DCNN(BERT(x l ))

[0019] Where x v , x a , x l are the original modal data, X l ∈R b×D ,X a ∈R b×D ,X v ∈R b×D is the extracted modal primary feature, where b represents the batch_size and D represents the feature dimension.

[0020] Furthermore, the method of using a high-dimensional Gaussian distribution to model the primary feature representation to obtain the mean and variance of the current modal feature; using the variance to quantify the uncertainty level of the current modality, and obtaining a stable representation of the primary feature through Gaussian distribution reparameterization includes:

[0021] Mapping the high-dimensional Gaussian distribution of each mode: Each mode is mapped to a high-dimensional Gaussian distribution through two fully connected layers to obtain the mean and variance of each mode. At the same time, the Gaussian distribution is sampled using the reparameterization technique to enable the variables to be back-propagated. The relevant formula is as follows:

[0022] p(z m |x m )~N(μ m ,σ m2 )

[0023] μ m =W1x m ,σ m =W2x m

[0024] z m =μ m +∈σ m ,∈~N(0,I)

[0025] Where x m Represents the primary modal characteristics of mode m, μ m and σ mRespectively represent the mean and variance after high-dimensional Gaussian distribution mapping, W1 represents the mean weight matrix, W2 represents the variance weight matrix, z m Represents the reparameterized eigenvector, ∈σ m Indicates random sampling of noise, so that the modal data is no longer a fixed value, but is sampled from the mapped distribution;

[0026] During model training, the mapped high-dimensional Gaussian distribution needs to be compared with the standard normal distribution to calculate the Kullback-Leibler divergence loss. The formula is as follows:

[0027]

[0028] Where, KL() represents the KL divergence calculation formula, N(μ m ,σ m2 ) represents the high-dimensional Gaussian distribution of the mapping, N(0,I) represents the standard normal distribution, log is the logarithmic function with e as the base, and the loss is to ensure that N(μ m ,σ m2 ) can correctly map the primary features of the modal to the high-dimensional Gaussian distribution, and can better ensure σ m2 Can correctly quantify modal uncertainty.

[0029] Furthermore, the hybrid multi-expert model extracts the re-parameterized features, and the final modal features are weighted summed by the selected experts with weights to obtain the hidden layer feature representation, including:

[0030] The hybrid multi-expert model extracts features from the re-parameterized features. In order to achieve the goal of activating more multi-experts when the modal uncertainty is higher, a dynamic top-k gating unit is designed to select the number of experts and initialize the weights.

[0031] Dynamic top-k gating unit utilizes U m =[z m ,σ m ] as input, and use two fully connected layers to output the expert activation ratio and expert weight distribution respectively. Finally, the selected experts will retain their weights, and the unselected experts will be blocked. Finally, the outputs of multiple experts will be weighted fused. The specific formula is as follows:

[0032]

[0033] G top-k ,I top-k =TopK(softmax(W g z),k)

[0034]

[0035] Where Wk For U m The input weight matrix represents the fully connected layer calculated by k value, N is the total number of set experts; Sigmoid is the activation function, the value range is (0,1), the topK function means extracting the first k weights and experts according to the k value; M means based on I top-k The mask set, G′ is the normalized expert assignment probability, ∈ is a numerical stability constant; the final G′ will be assigned to each expert, and the product summation with the expert's output will obtain the dynamic multi-expert decision feature representation.

[0036] Furthermore, during training, the number of dynamic multi-experts needs to be constrained. First, the multimodal data of the same batch needs to be normalized. The calculated value of the target k needs to be constrained by taking its expectation using the normalized variance. Then, its mean square error loss is calculated as its self-supervisory loss. The specific formula is as follows:

[0037]

[0038] k target =1+E[σ norm ]·(N-1)

[0039]

[0040] Where σ norm is the normalized variance value, E[] represents the expected variance, min(σ) and max(σ) represent the minimum variance and maximum variance in the same batch respectively, k target It represents the expected value of k. Using MSELoss calculation can make the model self-supervised and converge to an appropriate level, and make the number of dynamic multi-experts more reasonable and reliable.

[0041] Furthermore, the three modal features are vector-concatenated to form a joint representation, and the internal correlation of the joint representation is modeled through a self-attention mechanism to enhance the complementary information between the modalities, including:

[0042] The latent feature representations of the three modalities (language, audio, and vision) are vector-concatenated to form a joint feature representation, achieving preliminary integration of modal information. The self-attention mechanism is used to model the internal correlation of the joint representation, enhancing the complementary information between the modalities, and this is used as the subsequent query input. The specific formula is as follows:

[0043]

[0044] Where Q, K, and V are the query, key, and value input into the self-attention mechanism, respectively. Here, Q = K = V = [z v , z a , z l], where z v , z a , z l They represent the visual, speech, and text features obtained by multi-expert learning output by the uncertainty-aware adaptive expert allocation mechanism module, and H represents the obtained joint feature representation.

[0045] Furthermore, the three modalities are sorted from high to low according to reliability based on the uncertainty information to construct a modality priority queue, including:

[0046] Uncertainty-driven modality sorting and feature recombination: Acquire features of the three modalities (language, audio, and vision) and calculate the uncertainty scalar for each modality; sort based on uncertainty to ensure that the modalities with high reliability are prioritized for fusion;

[0047] The formula for calculating the mean uncertainty of each mode is as follows:

[0048]

[0049] Where, It is calculated as the uncertainty scalar, and π represents the uncertainty size operator. π is used to rearrange the modal features so that they are organized in descending order of reliability to improve the stability of modal fusion.

[0050] Furthermore, the three-level cascaded cross-attention fusion structure uses the self-attention output as the initial query vector, and sequentially inputs the modal features as key-value pairs according to the reliability priority. Each level uses the output of the previous level as the query vector of the next level, thus achieving progressive feature fusion.

[0051] According to the modality priority queue, a three-level cascade cross-attention fusion structure is designed to fuse multimodal feature representations to obtain a multimodal joint feature representation, including:

[0052] Cascaded cross-attention fusion multimodal feature representation: Based on uncertainty-driven modal sorting and feature reorganization, feature representations are obtained with priority sorted from high to low, ensuring that the most reliable modality is prioritized for cascaded cross-attention fusion. The specific formula is as follows:

[0053]

[0054] Where, d k is the dimension size of Query and Key in the attention mechanism; H is the joint feature representation output by the self-attention mechanism as Query, guiding the first priority mode Perform fusion and obtain h1, which is used as the next layer input Query to guide the second priority mode Perform fusion and obtain h2, which is used as the next layer input Query to guide the third priority mode After fusion, h3 is finally obtained as the fusion feature of the final multimodal output for subsequent sentiment intensity prediction.

[0055] Furthermore, the step of predicting the intensity of emotion based on the multimodal joint feature representation to obtain a prediction result includes:

[0056] A three-layer multilayer perceptron is used to evaluate the sentiment intensity of multimodal fusion features. The first layer is the input layer, the size of which is 72, the second layer is a hidden layer with a size of 128, and the third layer is the output layer to predict the sentiment intensity.

[0057] The second technical solution adopted by the present invention is:

[0058] An adaptive multi-expert collaborative multimodal emotion recognition system, comprising:

[0059] The feature extraction module is used to acquire multimodal data and extract the primary feature representation of each modality; the multimodal data includes three modalities: image, speech, and text;

[0060] The uncertainty-aware adaptive expert allocation mechanism module is used to model the primary feature representation using a high-dimensional Gaussian distribution to obtain the mean and variance of the current modal feature. The variance is used to quantify the uncertainty level of the current modality, and a stable representation of the primary feature is obtained through Gaussian distribution reparameterization. A hybrid multi-expert model is used to extract the reparameterized features. The final modal feature is weighted summed by the selected experts with weights to obtain the hidden feature representation.

[0061] The reliability-level sentiment decoding module is used to concatenate the three modal features into a joint representation through vector concatenation. It also uses a self-attention mechanism to model the internal correlation of the joint representation and enhance the complementary information between the modalities. Based on uncertainty information, the three modalities are sorted from high to low in terms of reliability to construct a modality priority queue. Based on the modality priority queue, a three-level cascaded cross-attention fusion structure is designed to fuse the multimodal feature representations to obtain a multimodal joint feature representation.

[0062] The emotion intensity prediction module is used to predict the intensity of emotion based on the multimodal joint feature representation and obtain the prediction result.

[0063] The third technical solution adopted by the present invention is:

[0064] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the adaptive multi-expert collaborative multimodal emotion recognition method as described above.

[0065] The fourth technical solution adopted by the present invention is:

[0066] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the adaptive multi-expert collaborative multimodal emotion recognition method as described above.

[0067] The fifth technical solution adopted by the present invention is:

[0068] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the above-mentioned adaptive multi-expert collaborative multimodal emotion recognition method.

[0069] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0070] (1) Compared with existing technologies, this invention does not directly use modal data containing noise for feature extraction and feature fusion. Instead, it innovatively introduces a high-dimensional Gaussian distribution into the model to quantify uncertainty. This method effectively quantifies and utilizes the noise in the modal data in the form of uncertainty.

[0071] (2) Compared to existing hybrid multi-expert model approaches, this paper proposes an uncertainty-aware adaptive expert allocation mechanism that dynamically selects the number of experts and assigns weights based on the uncertainty of the sample modalities, and then performs multi-expert feature extraction. This approach effectively ensures that features present in uncertain data can be effectively extracted and utilized.

[0072] (3) A reliability-level emotion decoding module is designed. The present invention can sort the modalities based on uncertainty-driven and perform feature reorganization to obtain modalities sorted from high to low reliability and fuse multimodal features through cascaded cross-attention at one time, so that the fusion of each modality is more complete, the generalization of the model is further improved, and the emotion recognition effect can be better improved in situations such as data interference that exist in reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.

[0074] Figure 1 Schematic diagram of the overall module relationship of the adaptive multi-expert collaborative multimodal emotion recognition system in an embodiment of the present invention.

[0075] Figure 2 Schematic diagram of a feature extraction module in an embodiment of the present invention.

[0076] Figure 3 This is a workflow diagram of the uncertainty-aware adaptive expert allocation mechanism module in an embodiment of the present invention.

[0077] Figure 4 Schematic diagram of the reliability level emotion decoding module and the emotion intensity assessment module in an embodiment of the present invention.

[0078] Figure 5 This is a flowchart of the steps of an adaptive multi-expert collaborative multimodal emotion recognition method in an embodiment of the present invention. DETAILED DESCRIPTION

[0079] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. For the step numbers in the following embodiments, they are provided only for the convenience of explanation and are not intended to limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0080] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as setting, installing, and connecting should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0081] In the description of this application, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.

[0082] In the description of this application, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The terms "first" and "second" are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or as implicitly specifying the number or order of the technical features indicated.

[0083] In the description of this application, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0084] In response to the shortcomings and deficiencies of the existing technology, the present invention proposes an uncertainty-driven adaptive multi-expert collaborative multimodal emotion recognition scheme, which mines the uncertainty of each modality in different samples and quantifies the noise in the modality. Secondly, in order to make more full use of the implicit features of the modality, a hybrid multi-expert model structure is adopted. The uncertainty of different samples dynamically determines the number of experts to extract more stable feature representations, and optimizes the emotion fusion strategy according to the uncertainty ranking, thereby improving the accuracy of multimodal emotion recognition and providing a new solution for multimodal emotion recognition.

[0085] Example 1

[0086] like Figures 1 to 4 As shown, this embodiment provides an uncertainty-driven adaptive multi-expert collaborative multimodal emotion recognition system, which is developed and implemented based on the PyTorch framework in a Windows environment. The specific implementation includes:

[0087] The feature extraction module is used to extract the primary feature representation of each emotion modality. The modal data mainly includes three modalities: image, speech, and text. The features of different modalities are filled to achieve modal alignment to obtain the correct primary feature representation of the modality.

[0088] The uncertainty-aware adaptive expert allocation mechanism module is used to quantify the uncertainty information of each modality to represent the reliability of the modality. First, the primary feature representation is modeled using a high-dimensional Gaussian distribution to obtain the mean and variance of the current modal feature. The variance is used to quantify the uncertainty level of the current modality, and the stable representation of the primary feature is obtained through Gaussian distribution reparameterization. The hybrid multi-expert model further extracts the reparameterized features, where the uncertainty and reparameterized features will dynamically determine the number of experts and the weight of each expert in the multi-expert model, so that the higher the modal uncertainty, the more multi-experts are activated. The final modal features are weighted summed by the screened experts with the weights to obtain a more advanced and reliable hidden modal feature representation;

[0089] Reliability-level sentiment decoding module, used to achieve adaptive modal fusion based on uncertainty. This module receives the latent feature representations of the three modalities and their corresponding uncertainty information. First, the three modal features are vector-concatenated to form a joint representation, and the internal correlation of the joint representation is modeled through the self-attention mechanism to enhance the complementary information between the modalities. Then, based on the uncertainty information, the three modalities are sorted from high to low according to reliability to construct a modal priority queue. A three-level cascaded cross-attention fusion structure is designed, with the self-attention output as the initial query vector, and the modal features are input as key-value pairs in sequence according to the reliability priority. Each level uses the output of the previous level as the query vector of the next level to achieve progressive feature fusion. This fusion method ensures that high-reliability modal information participates in the fusion process first, reducing the influence of noise and uncertainty. At the same time, the fusion strategy is dynamically adjusted according to the specific uncertainty distribution of each sample to achieve adaptive modal fusion, and finally outputs high-quality multimodal joint feature representation for subsequent sentiment classification tasks;

[0090] The sentiment intensity prediction module takes the high-quality multimodal joint feature representation output by the reliability-level sentiment decoding module as input and uses a multi-layer perceptron to predict the intensity of sentiment.

[0091] As an optional implementation, Figure 2 As shown, the feature extraction module performs the following operations:

[0092] First, we load the images, speech, and text files in the sentiment dataset separately, and extract the primary features of the modality using the three feature extractors FACET, COVAREP, and BERT. Then, we design a one-dimensional convolutional network for each modality to compress the features. The feature extraction formula is as follows:

[0093] X v =1DCNN(FACET(x v ))

[0094] X a =1DCNN(COVAREP(x a ))

[0095] X l =1DCNN(BERT(x l ))

[0096] Where x v , x a , x l are the original modal data, X l ∈R b×D ,X a ∈N b×D ,X v ∈N b×D is the extracted modal primary feature, where b represents the batch_size, which is set to 16, and D represents the feature dimension, which is set to 72.

[0097] As an optional implementation, Figure 3 As shown, the uncertainty-aware adaptive expert allocation mechanism module performs the following operations:

[0098] 1) Map the high-dimensional Gaussian distribution of each mode. In this step, each mode is mapped to a high-dimensional Gaussian distribution through two fully connected layers to obtain the mean and variance of each mode. At the same time, the Gaussian distribution is sampled using the reparameterization technique so that the variables can be back-propagated. The relevant formula is as follows:

[0099] p(z m |x m )~N(μ m ,σ m2 )

[0100] μ m =W1x m ,σ m =W2x m

[0101] z m =μ m +∈σ m ,∈~N(0,I)

[0102] Where x m Represents the primary modal characteristics of mode m, μ m and σ m Respectively represent the mean and variance after high-dimensional Gaussian distribution mapping, W1 represents the mean weight matrix, W2 represents the variance weight matrix, z m Represents the reparameterized eigenvector, ∈σ mIndicates random sampling of noise, so that the modal data is no longer a fixed value, but is sampled from the mapped distribution; during model training, the mapped high-dimensional Gaussian distribution needs to be compared with the standard normal distribution to calculate the Kullback-Leibler divergence loss, the formula is as follows:

[0103]

[0104] Where, KL() represents the KL divergence calculation formula, N(μ m ,σ m2 ) represents the high-dimensional Gaussian distribution of the mapping, N(0,I) represents the standard normal distribution, log is the logarithmic function with e as the base, and the loss is to ensure that N(μ m ,σ m2 ) can correctly map the primary features of the modal to the high-dimensional Gaussian distribution, and can better ensure σ m2 Can correctly quantify modal uncertainty.

[0105] 2) The hybrid multi-expert model extracts the heavy parameterized features. In order to achieve the goal of activating more multi-experts with higher modal uncertainty, a dynamic top-k gating unit is designed to select the number of experts and initialize the weights. In this step, the dynamic top-k gating unit uses U m =[z m ,σ m ] as input, and use two fully connected layers to output the expert activation ratio and expert weight distribution respectively. Finally, the selected experts will retain their weights, and the unselected experts will be blocked. Finally, the outputs of multiple experts will be weighted fused. The specific formula is as follows:

[0106]

[0107] G top-k ,I top-k =TopK(softmax(W g z),k)

[0108]

[0109] Where W k For U m The input weight matrix represents the fully connected layer calculated by k value, N is the total number of set experts; Sigmoid is the activation function, the value range is (0,1), the topK function means extracting the first k weights and experts according to the k value; M means based on I top-kThe mask is set, G′ is the normalized expert assignment probability, and ∈ is a numerical stability constant. The final G′ will be assigned to each expert and multiplied and summed with the expert's output to obtain the feature representation after dynamic multi-expert decision. During training, it is necessary to constrain the number of dynamic multi-experts. First, the multimodal data of the same batch needs to be normalized. The calculated value of the target k needs to be constrained by taking its expected variance after normalization. Then, its mean square error loss is calculated as its self-supervised loss. The specific formula is as follows:

[0110]

[0111] k target =1+E[σ norm ]·(N-1)

[0112]

[0113] Where σ norm is the normalized variance value, E[] represents the expected variance, min(σ) and max(σ) represent the minimum variance and maximum variance in the same batch respectively, k target It represents the expected value of k. Using MSELoss calculation can make the model self-supervised and converge to an appropriate level, and make the number of dynamic multi-experts more reasonable and reliable.

[0114] As an optional implementation, Figure 4 As shown, the reliability level emotion decoding module performs the following steps:

[0115] 1) Obtaining a joint multimodal feature representation. In this step, the features of the three modalities are obtained from the uncertainty-aware adaptive expert allocation mechanism module. The latent feature representations of the three modalities (language, audio, and vision) are vector-concatenated to form a joint feature representation, achieving preliminary integration of modal information. The self-attention mechanism is used to model the internal correlation of the joint representation, enhancing the complementary information between the modalities, and this is used as the subsequent query input. The specific formula is as follows:

[0116]

[0117] Where Q, K, and V are the query, key, and value input into the self-attention mechanism, respectively. Here, Q = K = V = [z v , z a , z l ], where z v , z a , z l They represent the visual, speech, and text features obtained by multi-expert learning output by the uncertainty-aware adaptive expert allocation mechanism module, and H represents the obtained joint feature representation.

[0118] 2) Uncertainty-driven modality sorting and feature recombination. In this step, the features of the three modalities (language, audio, and vision) are obtained from the uncertainty-aware adaptive expert allocation mechanism module, and the uncertainty scalar of each modality is calculated. Based on the uncertainty sorting, the modal with high reliability is prioritized for fusion. First, the uncertainty mean of each modality is calculated:

[0119]

[0120] Where, It is calculated as the uncertainty scalar, and π represents the uncertainty size operator. π is used to rearrange the modal features so that they are organized in order of reliability from high to low to improve the stability of modal fusion so as to input the subsequent cascaded cross-attention module.

[0121] 3) Cascaded cross-attention fusion of multimodal feature representations. In this step, feature representations are obtained based on uncertainty-driven modal sorting and feature reorganization, sorted from high to low priority, ensuring that the most reliable modality is prioritized for cascaded cross-attention fusion. The specific formula is as follows:

[0122]

[0123] Where, represents the three-fold cascade cross-attention mechanism module. First, H is the joint feature representation output by the self-attention mechanism as the query, guiding the first priority mode Perform fusion and obtain h1, which is used as the next layer input Query to guide the second priority mode The fusion is performed and h2 is obtained as the output of the next layer, and so on. Finally, h3 is obtained as the fusion feature of the final multimodal output for subsequent sentiment intensity prediction.

[0124] As an optional implementation, the emotion intensity prediction module is implemented as follows:

[0125] A three-layer multilayer perceptron is used to evaluate the sentiment intensity of multimodal fusion features. The first layer is the input layer, the size of which is 72, the second layer is a hidden layer with a size of 128, and the third layer is the output layer to predict the sentiment intensity.

[0126] Example 2

[0127] like Figure 5 As shown, this embodiment provides an adaptive multi-expert collaborative multimodal emotion recognition method, comprising the following steps:

[0128] S1. Acquire multimodal data and extract primary feature representations of each modality; multimodal data includes three modalities: image, speech, and text.

[0129] In some embodiments, first, the images, voice, and text files in the emotion dataset are loaded separately, and the primary features of the modality are extracted using the three feature extractors FACET, COVAREP, and BERT, respectively. Then, a one-dimensional convolutional network is designed for each modality to compress the features. The feature extraction formula is as follows:

[0130] X v =1DCNN(FACET(x v ))

[0131] X a =1DCNN(COVAREP(x a ))

[0132] X l =1CDNN(BERT(x l ))

[0133] Where x v , x a , x l are the original modal data, X l ∈R b×D ,X a ∈R b×D ,X v ∈R b×D is the extracted modal primary feature, where b represents the batch_size, which is set to 16, and D represents the feature dimension, which is set to 72.

[0134] S2. Use high-dimensional Gaussian distribution to model the primary feature representation and obtain the mean and variance of the current modal feature; use the variance to quantify the uncertainty level of the current mode and obtain a stable representation of the primary feature through Gaussian distribution reparameterization.

[0135] For example, the high-dimensional Gaussian distribution of each mode is mapped. In this step, each mode is mapped to a high-dimensional Gaussian distribution through two fully connected layers to obtain the mean and variance of each mode. At the same time, the Gaussian distribution is sampled using the reparameterization technique so that the variables can be back-propagated. The relevant formula is as follows:

[0136] p(z m |x m )~N(μ m ,σ m2 )

[0137] μ m =W1x m ,σm =W2x m

[0138] z m =μ m +∈σ m ,∈~N(0,I)

[0139] Where x m Represents the primary modal characteristics of mode m, μ m and σ m Respectively represent the mean and variance after high-dimensional Gaussian distribution mapping, W1 represents the mean weight matrix, W2 represents the variance weight matrix, z m Represents the reparameterized eigenvector, ∈σ m Indicates random sampling of noise, so that the modal data is no longer a fixed value, but is sampled from the mapped distribution; during model training, the mapped high-dimensional Gaussian distribution needs to be compared with the standard normal distribution to calculate the Kullback-Leibler divergence loss, the formula is as follows:

[0140]

[0141] Where, KL() represents the KL divergence calculation formula, N(μ m ,σ m2 ) represents the high-dimensional Gaussian distribution of the mapping, N(0,I) represents the standard normal distribution, log is the logarithmic function with e as the base, and the loss is to ensure that N(μ m ,σ m2 ) can correctly map the primary features of the modal to the high-dimensional Gaussian distribution, and can better ensure σ m2 Can correctly quantify modal uncertainty.

[0142] S3. The hybrid multi-expert model extracts the heavy parameterized features, and the final modal features are weighted summed by the selected experts with weights to obtain the hidden layer feature representation.

[0143] The hybrid multi-expert model extracts features from the re-parameterized features. In order to achieve the goal of activating more multi-experts with higher modal uncertainty, a dynamic top-k gating unit is designed to select the number of experts and initialize the weights. In this step, the dynamic top-k gating unit uses U m =[z m ,σ m ] as input, and use two fully connected layers to output the expert activation ratio and expert weight distribution respectively. Finally, the selected experts will retain their weights, and the unselected experts will be blocked. Finally, the outputs of multiple experts will be weighted fused. The specific formula is as follows:

[0144]

[0145] G top-k ,I top-k =TopK(softmax(W g z),k)

[0146]

[0147] Where W k For U m The input weight matrix represents the fully connected layer calculated by k value, N is the total number of set experts; Sigmoid is the activation function, the value range is (0,1), the topK function means extracting the first k weights and experts according to the k value; M means based on I top-k The mask is set, G′ is the normalized expert assignment probability, and ∈ is a numerical stability constant. The final G′ will be assigned to each expert and multiplied and summed with the expert's output to obtain the feature representation after dynamic multi-expert decision. During training, it is necessary to constrain the number of dynamic multi-experts. First, the multimodal data of the same batch needs to be normalized. The calculated value of the target k needs to be constrained by taking its expected variance after normalization. Then, its mean square error loss is calculated as its self-supervised loss. The specific formula is as follows:

[0148]

[0149] k target =1+E[σ norm ]·(N-1)

[0150]

[0151] Where σ norm is the normalized variance value, E[] represents the expected variance, min(σ) and max(σ) represent the minimum variance and maximum variance in the same batch respectively, k target It represents the expected value of k. Using MSELoss calculation can make the model self-supervised and converge to an appropriate level, and make the number of dynamic multi-experts more reasonable and reliable.

[0152] S4. Concatenate the three modal features into vectors to form a joint representation, and use the self-attention mechanism to model the internal correlation of the joint representation to enhance the complementary information between the modalities.

[0153] Acquire a joint representation of multimodal features. In this step, the features of the three modalities are obtained from the uncertainty-aware adaptive expert allocation mechanism module. The latent feature representations of the three modalities (language, audio, and vision) are vector-concatenated to form a joint feature representation, achieving preliminary integration of modal information. A self-attention mechanism is used to model the internal correlation of the joint representation, enhancing the complementary information between the modalities, and this is used as the subsequent query input. The specific formula is as follows:

[0154]

[0155] Where Q, K, and V are the query, key, and value input into the self-attention mechanism, respectively. Here, Q = K = V = [z v , z a , z l ], where z v , z a , z l They represent the visual, speech, and text features obtained by multi-expert learning output by the uncertainty-aware adaptive expert allocation mechanism module, and H represents the obtained joint feature representation.

[0156] S5. Sort the three modalities from high to low according to their reliability based on the uncertainty information and build a modality priority queue.

[0157] Uncertainty-driven modality sorting and feature recombination: In this step, the features of the three modalities (language, audio, and vision) are obtained from the uncertainty-aware adaptive expert allocation mechanism module, and the uncertainty scalar of each modality is calculated. Based on the uncertainty sorting, the modalities with high reliability are prioritized for fusion. First, the uncertainty mean of each modality is calculated:

[0158]

[0159] Where, It is calculated as the uncertainty scalar, and π represents the uncertainty size operator. π is used to rearrange the modal features so that they are organized in order of reliability from high to low to improve the stability of modal fusion so as to input the subsequent cascaded cross-attention module.

[0160] S6. Based on the modality priority queue, a three-level cascaded cross-attention fusion structure is designed to fuse the multimodal feature representations and obtain a multimodal joint feature representation.

[0161] Cascaded cross-attention fusion of multimodal feature representations. In this step, feature representations are obtained based on uncertainty-driven modal sorting and feature reorganization, sorted from high to low priority, to ensure that the most reliable modality is prioritized for cascaded cross-attention fusion. The specific formula is as follows:

[0162]

[0163]

[0164] Where, represents the three-fold cascade cross-attention mechanism module. First, H is the joint feature representation output by the self-attention mechanism as the query, guiding the first priority mode Perform fusion and obtain h1, which is used as the next layer input Query to guide the second priority mode The fusion is performed and h2 is obtained as the output of the next layer, and so on. Finally, h3 is obtained as the fusion feature of the final multimodal output for subsequent sentiment intensity prediction.

[0165] S7. Predict the intensity of emotion based on the multimodal joint feature representation to obtain the prediction result.

[0166] Specifically, a three-layer multi-layer perceptron is used to evaluate the sentiment intensity of multimodal fusion features. The first layer is the input layer, the size of which is 72, the second layer is a hidden layer with a size of 128, and the third layer is the output layer to predict the sentiment intensity.

[0167] Example 3

[0168] An embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the following Figure 5 An adaptive multi-expert collaborative multimodal emotion recognition method is shown.

[0169] It is understood that the memory may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the server, etc.

[0170] The processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the server. It executes various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, as well as accessing data stored in memory. Optionally, the processor may be implemented using at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate one or a combination of a central processing unit (CPU) and a modem. The CPU primarily processes the operating system and application programs, while the modem handles wireless communications. It is understood that the modem may not be integrated into the processor and may be implemented separately via a single chip.

[0171] Since the electronic device is an electronic device corresponding to an adaptive multi-expert collaborative multimodal emotion recognition method in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0172] Example 4

[0173] An embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 5 An adaptive multi-expert collaborative multimodal emotion recognition method is shown.

[0174] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0175] Since the storage medium is a storage medium corresponding to an adaptive multi-expert collaborative multimodal emotion recognition method in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0176] Example 5

[0177] In some possible implementations, various aspects of the method of the embodiments of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to execute the steps of an adaptive multi-expert collaborative multimodal emotion recognition method according to various exemplary embodiments of the present application described above in this specification. The executable computer program code or "code" for executing each embodiment may be written in a high-level programming language such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0178] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0179] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0180] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. An adaptive multi-expert collaborative multimodal emotion recognition method, characterized in that: The following steps are involved: Acquire multimodal data and extract primary feature representations for each modality; multimodal data includes three modalities: image, speech, and text; Use high-dimensional Gaussian distribution to model the primary feature representation and obtain the mean and variance of the current modal feature; use the variance to quantify the uncertainty level of the current mode, and obtain a stable representation of the primary feature through Gaussian distribution reparameterization; The hybrid multi-expert model extracts the heavily parameterized features, and the final modal features are weighted summed by the selected experts with weights to obtain the hidden layer feature representation; The three modal features are vector-concatenated to form a joint representation, and the internal correlation of the joint representation is modeled through the self-attention mechanism to enhance the complementary information between the modalities; Based on the uncertainty information, the three modes are sorted from high to low in terms of reliability, and a modality priority queue is constructed; According to the modality priority queue, a three-level cascade cross-attention fusion structure is designed to fuse multimodal feature representations and obtain multimodal joint feature representation; The intensity of emotion is predicted based on the multimodal joint feature representation to obtain the prediction result.

2. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The step of acquiring multimodal data and extracting primary feature representations of each modality includes: We obtain images, speech, and text files from the sentiment dataset and extract primary features of each modality using the three feature extractors FACET, COVAREP, and BERT. We then design a one-dimensional convolutional network for each modality to compress the features. The feature extraction formula is as follows: X v =1DCNN(FACET(x v )) X a =1DCNN(COVAREP(x a )) X l =1DCNN(BERT(x l )) Where x v , x a , x l are the original modal data, X l ∈R b×D ,X a ∈R b×D ,X v ∈R b×D is the extracted modal primary feature, where b represents the batch_size and D represents the feature dimension.

3. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The method uses a high-dimensional Gaussian distribution to model the primary feature representation, obtains the mean and variance of the current modal feature, uses the variance to quantify the uncertainty level of the current modality, and obtains a stable representation of the primary feature through Gaussian distribution reparameterization, including: Mapping the high-dimensional Gaussian distribution of each mode: Each mode is mapped to a high-dimensional Gaussian distribution through two fully connected layers to obtain the mean and variance of each mode. At the same time, the Gaussian distribution is sampled using the reparameterization technique to enable the variables to be back-propagated. The relevant formula is as follows: p(z m |x m )~N(μ m ,s m2 ) m m =W1x m ,s m =W2x m With m =μ m +∈σ m ,∈~N(0,I) Where x m Represents the primary modal characteristics of mode m, μ m and σ m Respectively represent the mean and variance after high-dimensional Gaussian distribution mapping, W1 represents the mean weight matrix, W2 represents the variance weight matrix, z m Represents the reparameterized eigenvector, ∈σ m Indicates random sampling of noise, so that the modal data is no longer a fixed value, but is sampled from the mapped distribution; during model training, the mapped high-dimensional Gaussian distribution needs to be compared with the standard normal distribution to calculate the Kullback-Leibler divergence loss, the formula is as follows: Where, KL() represents the KL divergence calculation formula, N(μ m ,σ m2 ) represents the high-dimensional Gaussian distribution of the mapping, N(0,I) represents the standard normal distribution, log is the logarithmic function with e as the base, and the loss is to ensure that N(μ m ,σ m2 ) can correctly map the primary features of the modal to the high-dimensional Gaussian distribution, and can better ensure σ m2 Can correctly quantify modal uncertainty.

4. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The hybrid multi-expert model extracts the re-parameterized features, and the final modal features are weighted summed by the selected experts with weights to obtain the hidden layer feature representation, including: The hybrid multi-expert model extracts features from the re-parameterized features. In order to achieve the goal of activating more multi-experts when the modal uncertainty is higher, a dynamic top-k gating unit is designed to select the number of experts and initialize the weights. Dynamic top-k gating unit utilizes U m =[z m ,σ m ] as input, and use two fully connected layers to output the expert activation ratio and expert weight distribution respectively. Finally, the selected experts will retain their weights, and the unselected experts will be blocked. Finally, the outputs of multiple experts will be weighted fused. The specific formula is as follows: G top-k ,I top-k =TopK(softmax(W g z), k) Where W k For U m The input weight matrix represents the fully connected layer calculated by k value, N is the total number of set experts; Sigmoid is the activation function, the value range is (0,1), the topK function means extracting the first k weights and experts according to the k value; M means based on I top-k The mask set, G′ is the normalized expert assignment probability, ∈ is a numerical stability constant; the final G′ will be assigned to each expert, and the product summation with the expert's output will obtain the dynamic multi-expert decision feature representation.

5. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 4, characterized in that: During training, the number of dynamic multi-experts needs to be constrained. First, the multimodal data of the same batch needs to be normalized. The calculated value of the target k needs to be constrained by taking its expectation using the normalized variance. Then, the mean square error loss is calculated as its self-supervisory loss. The specific formula is as follows: k target =1+E[σ norm ]·(N-1) Where σ norm is the normalized variance value, E[] represents the expected variance, min(σ) and max(σ) represent the minimum variance and maximum variance in the same batch respectively, k target It represents the expected value of k. Using MSELoss calculation can make the model self-supervised and converge to an appropriate level, and make the number of dynamic multi-experts more reasonable and reliable.

6. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The three modal features are vector-concatenated to form a joint representation, and the internal correlation of the joint representation is modeled through the self-attention mechanism to enhance the complementary information between the modalities, including: The latent feature representations of the three modalities (language, audio, and vision) are vector-concatenated to form a joint feature representation, achieving preliminary integration of modal information. The self-attention mechanism is used to model the internal correlation of the joint representation, enhancing the complementary information between the modalities, and this is used as the subsequent query input. The specific formula is as follows: Where Q, K, and V are the query, key, and value input into the self-attention mechanism, respectively. Here, Q = K = V = [z v , z a , z l ], where z v , z a , z l They represent the visual, speech and text features learned by multiple experts respectively, and H represents the joint feature representation obtained.

7. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The method of sorting the three modalities from high to low according to reliability based on uncertainty information and constructing a modality priority queue includes: Uncertainty-driven modal sorting and feature recombination: Obtain the features of the three modes and calculate the uncertainty scalar for each mode; sort based on uncertainty to ensure that the modes with high reliability are prioritized in fusion; The formula for calculating the mean uncertainty of each mode is as follows: Where, It is calculated as the uncertainty scalar, and π represents the uncertainty size operator. π is used to rearrange the modal features so that they are organized in descending order of reliability to improve the stability of modal fusion.

8. The adaptive multi-expert collaborative multimodal emotion recognition method according to claim 1, characterized in that: The three-level cascaded cross-attention fusion structure uses the self-attention output as the initial query vector, and sequentially inputs the modal features as key-value pairs according to the reliability priority. Each level uses the output of the previous level as the query vector for the next level, achieving progressive feature fusion. According to the modality priority queue, a three-level cascade cross-attention fusion structure is designed to fuse multimodal feature representations to obtain a multimodal joint feature representation, including: Cascaded cross-attention fusion multimodal feature representation: Based on uncertainty-driven modal sorting and feature reorganization, feature representations are obtained with priority sorted from high to low, ensuring that the most reliable modality is prioritized for cascaded cross-attention fusion. The specific formula is as follows: Where, d k is the dimension size of Query and Key in the attention mechanism; H is the joint feature representation output by the self-attention mechanism as Query, guiding the first priority mode Perform fusion and obtain h1, which is used as the next layer input Query to guide the second priority mode Perform fusion and obtain h2, which is used as the next layer input Query to guide the third priority mode After fusion, h3 is finally obtained as the fusion feature of the final multimodal output for subsequent sentiment intensity prediction.

9. An adaptive multi-expert collaborative multimodal emotion recognition system, characterized by: include: Feature extraction module, used to acquire multimodal data and extract primary feature representations of each modality; The multimodal data includes three modalities: image, voice and text; The uncertainty-aware adaptive expert allocation mechanism module is used to model the primary feature representation using a high-dimensional Gaussian distribution to obtain the mean and variance of the current modal feature. The variance is used to quantify the uncertainty level of the current modality, and a stable representation of the primary feature is obtained through Gaussian distribution reparameterization. The hybrid multi-expert model extracts the heavily parameterized features, and the final modal features are weighted summed by the selected experts with weights to obtain the hidden layer feature representation; The reliability-level sentiment decoding module is used to concatenate the three modal features into a joint representation through vector concatenation. It also uses a self-attention mechanism to model the internal correlation of the joint representation and enhance the complementary information between the modalities. Based on uncertainty information, the three modalities are sorted from high to low in terms of reliability to construct a modality priority queue. Based on the modality priority queue, a three-level cascaded cross-attention fusion structure is designed to fuse the multimodal feature representations to obtain a multimodal joint feature representation. The emotion intensity prediction module is used to predict the intensity of emotion based on the multimodal joint feature representation and obtain the prediction result.

10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Emotion recognition method and system based on visual and auditory collaboration

    CN120852890A

  • Multi-modal sentiment analysis method and system based on credibility driving

    CN121389026A

  • Method and system for multi-modal sentiment analysis based on credibility driving

    CN121389026B

  • Method and system for predicting remaining service life of battery based on dynamic expert hybrid model

    CN121679363A