Emotion feature visualization method and system based on modal mixing and dynamic evidence, terminal and storage medium
Patent Information
- Application Number
- CN202610806818.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-05
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于模态混合和动态证据的情感特征可视化方法、系统、终端及计算机可读存储介质,旨在解决现有技术中对缺失模态进行恢复时难以控制,导致无法准确表达细粒度情绪线索,进而限制最终识别准确性的问题
[0016]In this invention, an observed modality set and a missing modality set are obtained from an emotional sample. Multiple encoders are used to extract features from the observed modality set and the missing modality set, and all extracted features are mapped to a latent representation space to obtain an observed modality feature set and a missing modality feature set. A conditional distribution of the observed modality set and the missing modality set is constructed. A diffusion module is trained using the conditional distribution, the observed modality feature set, and the missing modality feature set, and the diffusion module is used to recover the missing modality set to obtain a complete modality set. Corresponding modality weights are assigned to different types of modalities in the complete modality set, and multiple different loss terms are constructed using all the modality weights. When optimizing the diffusion module using all the loss terms, the constructed error transformation function is used to adaptively modulate the reverse update process of each single modality to output the classification result of each modality feature. The classification result is then visualized to obtain all the visualization results of the emotional sample. This invention can dynamically allocate modal contributions for different samples, thereby achieving more robust and stable results under missing modal conditions and reducing the misleading effect of recovered modal noise on downstream tasks.
Smart Images

Figure CN122332930B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of feature visualization technology, and in particular to a method, system, terminal, and computer-readable storage medium for visualizing emotional features based on modality mixing and dynamic evidence. Background Technology
[0002] Multimodal Emotion Recognition (MER) can more comprehensively depict a person's emotional state by jointly modeling heterogeneous signals such as language content, facial expressions, and speech prosody. It has important application value in scenarios such as human-computer interaction, intelligent customer service, video understanding, and medical assistance.
[0003] However, in real-world deployment scenarios, multimodal data is often not fully available. Due to factors such as sensor malfunction, occlusion, changes in the acquisition environment, packet loss during transmission, and privacy restrictions, some samples may only contain text and vision, lacking acoustic modalities; or only the text modality may be retained, while other modalities are missing. This lack of modality directly disrupts the complementarity and consistency between modalities, causing traditional emotion recognition models that rely on complete input to significantly degrade in real-world environments.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] The main objective of this invention is to provide a method, system, terminal, and computer-readable storage medium for visualizing emotional features based on modality mixing and dynamic evidence. This aims to solve the problem in the prior art where it is difficult to control the recovery of missing modalities, resulting in the inability to accurately express fine-grained emotional cues and thus limiting the accuracy of the final recognition.
[0006] To achieve the above objectives, this invention provides a method for visualizing sentiment features based on modality mixture and dynamic evidence. The method includes the following steps: The observed modality set and the missing modality set are obtained from the emotion sample. Multiple encoders are used to extract features from the observed modality set and the missing modality set. All extracted features are mapped to the latent representation space to obtain the observed modality feature set and the missing modality feature set. A conditional distribution of the observed modality set and the missing modality set is constructed. A diffusion module is trained using the conditional distribution, the observed modality feature set, and the missing modality feature set. The diffusion module is then used to recover the missing modality set to obtain a complete modality set. Assign corresponding modal weights to different types of modes in the complete modality set, and construct various different loss terms using all the modal weights; When optimizing the diffusion module using all the aforementioned loss terms, the backward update process of each single modality is adaptively modulated using the constructed error transformation function to output the classification result of each modality feature, and the classification result is visualized to finally obtain all the visualization results of the sentiment sample.
[0007] Optionally, the emotional feature visualization method based on modality mixture and dynamic evidence, wherein obtaining the observed modality set and the missing modality set from the emotional sample, extracting features from the observed modality set and the missing modality set using multiple encoders, and mapping all extracted features to a latent representation space to obtain the observed modality feature set and the missing modality feature set, specifically includes: Obtain the emotional samples input by the user, and judge the emotional samples to determine all observed modalities and all missing modalities in the emotional samples, so as to obtain the set of observed modalities and the set of missing modalities; The text encoder, visual encoder, and acoustic encoder are used to extract features from the text modality, visual modality, and acoustic modality in the observed modality set and the missing modality set, respectively, to obtain the corresponding sequence semantic representation, the corresponding visual temporal features, and the corresponding acoustic features. By using a multilayer perceptron, all the sequence semantic representations, all the visual temporal features, and all the acoustic features are mapped to a latent representation space of the same dimension, the observed text features, observed visual features, observed acoustic features, missing text features, missing visual features, and missing acoustic features are obtained. An observation modality feature set is constructed using the observed text features, the observed visual features, and the observed acoustic features; a missing modality feature set is constructed using the missing text features, the missing visual features, and the missing acoustic features.
[0008] Optionally, the sentiment feature visualization method based on modality mixture and dynamic evidence, wherein constructing the conditional distributions of the observed modality set and the missing modality set, and training a diffusion module using the conditional distributions, the observed modality feature set, and the missing modality feature set, specifically includes: A conditional distribution is constructed using the observed mode set and the missing mode set, and a scoring network is then built. The observed modality feature set is input into the scoring network to learn the conditional distribution and obtain a conditional score. The missing modality feature set is input into the scoring network to learn the conditional distribution and obtain an unconditional score. Based on the conditional and unconditional scores, a guiding coefficient score is constructed, and the constructed initial diffusion module is trained using the guiding coefficient score to obtain a diffusion module for reconstructing the missing modality set: ; in, This represents the guiding coefficient score. Indicates the guiding strength parameter. Indicates conditional score. This indicates an unconditional score. This represents the missing modal feature set. Indicates diffusion time, Represents the set of observed modal features. This indicates a missing feature.
[0009] Optionally, the emotional feature visualization method based on modality mixture and dynamic evidence, wherein the step of recovering the missing modality set using the diffusion module to obtain the complete modality set specifically includes: The emotional samples are input into the missing modality recovery model of the diffusion module. During the forward diffusion stage, the missing modality recovery model gradually adds Gaussian noise to each missing modality feature in the missing modality set, converting each missing modality feature into corresponding pure noise. In the backdiffusion phase, for each missing modal feature, the pure noise, the current diffusion time step, and the observed modal feature set are input into the scoring network. The scoring network predicts the denoising direction of the missing modal feature to obtain the recovered modal feature of the missing modal feature. By integrating all the recovered modal features and the observed modal feature set, a complete modal set of the emotion sample is obtained.
[0010] Optionally, the sentiment feature visualization method based on modality mixture and dynamic evidence, wherein assigning corresponding modality weights to different types of modalities in the complete modality set, constructing multiple different loss terms using all the modality weights, and constructing an error transformation function using the modality features in the complete modality set, specifically includes: The complete modality set is input into the routing network, which analyzes the missing status of the sentiment samples based on the complete modality set and outputs the modality weight of each single modality. Based on all the modal weights, the sentiment sample is predicted as a whole and as a single modal, respectively, to obtain the final prediction result and multiple single modal prediction results. The final prediction result and all the single modal prediction results are then used to construct the multimodal prediction loss and the single modal prediction loss, respectively. Based on all the single-modal prediction results, the importance of each single mode is constructed, and alignment constraints are applied to all the importance and the corresponding modal weights to obtain the expert balance loss. Calculate the difference between the multimodal prediction loss and the single-modal prediction loss, and construct the single-modal distillation loss based on the difference; Calculate the representation distance between each missing modality feature and its corresponding recovered modality feature, and construct the reconstruction loss and score matching loss using all the representation distances.
[0011] Optionally, the sentiment feature visualization method based on modality mixture and dynamic evidence, wherein the step of performing overall prediction and single-modal prediction on the sentiment sample according to all the modality weights to obtain a final prediction result and multiple single-modal prediction results, and constructing a multimodal prediction loss and a single-modal prediction loss using the final prediction result and all the single-modal prediction results, specifically includes: The emotion sample is input into the emotion recognition fusion model, and a multimodal fusion representation is predicted and output based on all the modal weights. The multimodal fusion representation is then input into the prediction head for prediction, and the final prediction result is output. The sentiment samples are input into each unimodal prediction branch, and the unimodal latent representation of each missing modality is predicted and output according to all the modality weights. Each unimodal latent representation is input into the prediction head for prediction, and the unimodal prediction result of each missing modality is output. Based on the difference between the final prediction result and the multimodal true result, a multimodal prediction loss is constructed; Based on the difference between each single-mode prediction result and the corresponding single-mode true result, a corresponding single-mode loss function is constructed. All single-mode loss functions are then integrated to obtain the single-mode prediction loss.
[0012] Optionally, the sentiment feature visualization method based on modality mixture and dynamic evidence, wherein when optimizing the diffusion module using all the loss terms, adaptively modulates the reverse update process of each single modality using the constructed error transformation function to output the classification result of each modality feature, specifically includes: By combining the reconstruction loss, the score matching loss, the multimodal prediction loss, the unimodal prediction loss, the expert balancing loss, and the unimodal distillation loss, a joint loss function is obtained; The missing modality recovery model and the emotion recognition fusion model in the diffusion module are trained using the joint loss function, respectively. For each modality feature, a corresponding error transformation function is constructed using the modality feature. During the backpropagation of the joint loss function, the error transformation function is used to adaptively adjust various parameters of the missing modality recovery model and the emotion recognition fusion model to output the classification result corresponding to the modality feature.
[0013] Furthermore, to achieve the above objectives, the present invention also provides an emotional feature visualization system based on modality mixture and dynamic evidence, wherein the emotional feature visualization system based on modality mixture and dynamic evidence includes: The modality recognition module is used to acquire the observed modality set and the missing modality set in the emotion sample, use multiple encoders to extract features from the observed modality set and the missing modality set, and map all the extracted features to the latent representation space to obtain the observed modality feature set and the missing modality feature set; The modality reconstruction module is used to construct the conditional distribution of the observed modality set and the missing modality set, train a diffusion module using the conditional distribution, the observed modality feature set, and the missing modality feature set, and use the diffusion module to recover the missing modality set to obtain a complete modality set; The loss construction module is used to assign corresponding modal weights to different types of modes in the complete modality set, and to construct multiple different loss terms using all the modal weights; The model optimization and result visualization module is used to adaptively modulate the reverse update process of each single modality using the constructed error transformation function when optimizing the diffusion module with all the loss terms, so as to output the classification result of each modality feature, and to visualize the classification result, so as to finally obtain all the visualization results of the sentiment sample.
[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an emotional feature visualization program based on modal mixture and dynamic evidence stored in the memory and executable on the processor, wherein when the emotional feature visualization program based on modal mixture and dynamic evidence is executed by the processor, it implements the steps of the emotional feature visualization method based on modal mixture and dynamic evidence as described above.
[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a sentiment feature visualization program based on modal mixture and dynamic evidence, and the sentiment feature visualization program based on modal mixture and dynamic evidence, when executed by a processor, implements the steps of the sentiment feature visualization method based on modal mixture and dynamic evidence as described above.
[0016] In this invention, an observed modality set and a missing modality set are obtained from an emotional sample. Multiple encoders are used to extract features from the observed modality set and the missing modality set, and all extracted features are mapped to a latent representation space to obtain an observed modality feature set and a missing modality feature set. A conditional distribution of the observed modality set and the missing modality set is constructed. A diffusion module is trained using the conditional distribution, the observed modality feature set, and the missing modality feature set, and the diffusion module is used to recover the missing modality set to obtain a complete modality set. Corresponding modality weights are assigned to different types of modalities in the complete modality set, and multiple different loss terms are constructed using all the modality weights. When optimizing the diffusion module using all the loss terms, the constructed error transformation function is used to adaptively modulate the reverse update process of each single modality to output the classification result of each modality feature. The classification result is then visualized to obtain all the visualization results of the emotional sample. This invention can dynamically allocate modal contributions for different samples, thereby achieving more robust and stable results under missing modal conditions and reducing the misleading effect of recovered modal noise on downstream tasks. Attached Figure Description
[0017] Figure 1 This is a flowchart of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 2 This is a schematic diagram of the overall framework of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 3 This is a diffusion diagram of a preferred embodiment of the emotional feature visualization method based on modal mixing and dynamic evidence of the present invention; Figure 4 This is a schematic diagram of the routing network of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 5 This is a schematic diagram of single-modal distillation of a preferred embodiment of the emotional feature visualization method based on modal mixing and dynamic evidence of the present invention; Figure 6 This is a schematic diagram of the prediction process of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 7 This is a flowchart of the self-attention mechanism of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 8 This is a comparison chart of the visualization results of a preferred embodiment of the emotional feature visualization method based on modality mixing and dynamic evidence of the present invention; Figure 9This is a structural diagram of a preferred embodiment of the emotional feature visualization system based on modal mixing and dynamic evidence of the present invention; Figure 10 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] The preferred embodiment of the present invention describes a method for visualizing sentiment features based on modality mixture and dynamic evidence, such as... Figure 1 As shown, the sentiment feature visualization method based on modality mixture and dynamic evidence includes the following steps: Step S10: Obtain the observation modality set and the missing modality set from the emotion sample, use multiple encoders to extract features from the observation modality set and the missing modality set, and map all extracted features to the latent representation space to obtain the observation modality feature set and the missing modality feature set.
[0020] Specifically, the user inputs an emotional sample, and the emotional sample is judged to determine all observed modalities and all missing modalities in the emotional sample, thus obtaining the observed modality set and the missing modality set; The text encoder, visual encoder, and acoustic encoder are used to extract features from the text modality, visual modality, and acoustic modality in the observed modality set and the missing modality set, respectively, to obtain the corresponding sequence semantic representation, the corresponding visual temporal features, and the corresponding acoustic features. By using a multilayer perceptron, all the sequence semantic representations, all the visual temporal features, and all the acoustic features are mapped to a latent representation space of the same dimension, the observed text features, observed visual features, observed acoustic features, missing text features, missing visual features, and missing acoustic features are obtained. An observation modality feature set is constructed using the observed text features, the observed visual features, and the observed acoustic features; a missing modality feature set is constructed using the missing text features, the missing visual features, and the missing acoustic features.
[0021] In the embodiments disclosed in this invention, the input emotion sample is identified to determine the available and missing modalities. Then, all available modalities are input into their respective encoders, such as... Figure 2 As shown, the language modality is then input into the language encoder. (in, (For the language modality), the visual modality is input into the visual encoder. (in, (For the visual modality), the missing modality corresponding to the audio modality (if the audio modality is a usable modality, then input to the audio encoder). (among them, For audio modalities, generate features for each modality, and for missing modalities, generate corresponding latent representations.
[0022] In the embodiments disclosed in this invention, the text modality can be obtained from word vectors or a pre-trained language model to obtain sequential semantic representations; the visual modality can be obtained from facial regions, expression frames, or video frame sequences to obtain visual temporal features; and the acoustic modality can be obtained from speech waveforms, spectrograms, prosody, or acoustic descriptors to obtain acoustic features. Since the original dimensions, sampling frequencies, and statistical distributions of the three modalities differ, the system further maps the three modalities to a latent representation space of the same dimension through a linear mapping layer, a multilayer perceptron, or a temporal projection layer, obtaining text features, visual features, and acoustic features. During the training phase, if a modality is set as a missing modality, its true encoded features can still be used as a supervisory signal to calculate the diffusion reconstruction loss; during the inference phase, the true missing modality is not visible, and the system only uses the observed modality features as conditions to generate the reconstructed representation of the missing modality in the shared latent space. Thus, the encoding module not only completes the conversion of heterogeneous original data to a unified feature space but also provides a unified input for subsequent missing modality diffusion reconstruction and dynamic fusion.
[0023] In this invention, the original, heterogeneous multi-view data is transformed into unified numerical features that can be processed by the subsequent diffusion process, providing high-quality input for subsequent cluster analysis.
[0024] Step S20: Construct the conditional distribution of the observed mode set and the missing mode set, train a diffusion module using the conditional distribution, the observed mode feature set and the missing mode feature set, and use the diffusion module to recover the missing mode set to obtain the complete mode set.
[0025] In this invention, based on the aforementioned set of observed modes and set of missing modes, the missing modes are further recovered using a diffusion recovery mechanism based on stochastic differential equations, thereby obtaining a complete set of modes. The forward process generates noise by adding noise. Since it does not involve network training, the reverse process is directly denoised step by step under the constraints of observable modes in actual execution to obtain the reconstruction result.
[0026] Specifically, a conditional distribution is constructed using the observed mode set and the missing mode set, and a scoring network is built. The observed modality feature set is input into the scoring network to learn the conditional distribution and obtain a conditional score. The missing modality feature set is input into the scoring network to learn the conditional distribution and obtain an unconditional score. Based on the conditional and unconditional scores, a guiding coefficient score is constructed, and the constructed initial diffusion module is trained using the guiding coefficient score to obtain a diffusion module for reconstructing the missing modality set: ; in, This represents the guiding coefficient score. This represents the guiding strength parameter, used to control the dependence of the recovery results on the observed modal conditions. Indicates conditional score. This indicates an unconditional score. This represents the missing modal feature set. Indicates diffusion time, Represents the set of observed modal features. This indicates a missing feature.
[0027] Furthermore, the emotional samples are input into the missing modality recovery model of the diffusion module. During the forward diffusion stage, the missing modality recovery model progressively adds Gaussian noise to each missing modality feature in the missing modality set, converting each missing modality feature into corresponding pure noise. In the backdiffusion phase, for each missing modal feature, the pure noise, the current diffusion time step, and the observed modal feature set are input into the scoring network. The scoring network predicts the denoising direction of the missing modal feature to obtain the recovered modal feature of the missing modal feature. By integrating all the recovered modal features and the observed modal feature set, a complete modal set of the emotion sample is obtained.
[0028] This invention does not directly recover missing modalities from the original image, speech, or text, but rather recovers the missing modal representations from the shared latent space obtained by the encoder. During the training phase, this is done to learn the conditional distribution of the missing modalities relative to the observed modalities. First, missing modalities are randomly constructed from the complete training samples. One or more modalities are used as targets to be recovered, and their true encoded features are preserved as supervision. Figure 3 As shown, during the forward diffusion process, Gaussian noise is gradually added to the missing modal feature set, transforming it into a pure noise state. During the reverse denoising process, the scoring network uses the current noise state, the diffusion time step, and the observed modal features as input to gradually predict the denoising direction, thereby recovering the missing modal features. In the embodiments disclosed in this invention, the recovered audio modality is obtained. .
[0029] To achieve classifier-free guidance, this invention covers two input forms simultaneously when training the same scoring network: one is a conditional form, which inputs the observed modal conditions to learn conditional scores, and the other is an unconditional form, which sets the conditions to empty or randomly discards them to learn unconditional scores. During inference, the conditional scores and unconditional scores are combined according to the guidance coefficient to obtain the guided score.
[0030] This invention uses the true missing modal features in the training samples as the target and the observed modal features as conditions to learn the conditional generation process from the observed modal feature set to the missing modal feature set. During inference, it utilizes this conditional distribution to recover the invisible modality. Smaller reconstruction results place greater emphasis on semantic consistency with the observed modality; When the size is larger, the reconstruction process retains more smoothness and diversity of the generated distribution.
[0031] Furthermore, such as Figure 3 As shown, the score matching loss can be comprehensively considered during model training. and reconstruction losses This ensures that the recovered modality conforms to the true distribution and is semantically consistent with the existing modality.
[0032] ; in, This represents the set of modalities that are set as missing in the current sample. Indicates the first One missing mode, Indicates the first Individual modeling state; ; in, Represents the time weighting function. Indicates the diffusion time step. Indicates Gaussian noise. Indicates to , and Find the expectation of the joint distribution.
[0033] Step S30: Assign corresponding modal weights to different types of modes in the complete modal set, and construct various different loss terms using all the modal weights.
[0034] Among them, such as Figure 4 As shown, after obtaining the complete modality set, the present invention uses a routing network. (in, Represents a mapping function for routing networks. Represents the features of the input. The parameters representing the routing network are used to assign weights to different modalities (where the language modality weight is...). The visual modality weights are The audio modal weights are , l , v and a These represent the language module, visual modality, and audio modality, respectively, satisfying that each weight is non-negative and sums to 1. The temperature parameter t is used to adjust the routing sparsity: when t is small, the system tends to select a few high-confidence modalities; when t is large, the system adopts a smoother fusion strategy.
[0035] Specifically, the complete modality set is input into the routing network, which analyzes the missing status of the sentiment samples based on the complete modality set and outputs the modality weight of each single modality; Based on all the modal weights, the sentiment sample is predicted as a whole and as a single modal, respectively, to obtain the final prediction result and multiple single modal prediction results. The final prediction result and all the single modal prediction results are then used to construct the multimodal prediction loss and the single modal prediction loss, respectively. Based on all the single-modal prediction results, the importance of each single mode is constructed, and alignment constraints are applied to all the importance and the corresponding modal weights to obtain the expert balance loss. Calculate the difference between the multimodal prediction loss and the single-modal prediction loss, and construct the single-modal distillation loss based on the difference; Calculate the representation distance between each missing modality feature and its corresponding recovered modality feature, and construct the reconstruction loss and score matching loss using all the representation distances.
[0036] First, it should be noted that the complete modality set mentioned here refers to "the integrated result of observed modal features and reconstructed modal features." Specifically, if a sample originally contains text and visual modalities but lacks an acoustic modality, the system first reconstructs the acoustic latent representation using text and visual features through a diffusion module; then, the originally observed text features, visual features, and reconstructed acoustic features together form the complete modality set. If multiple modalities are missing, the corresponding modalities are reconstructed separately and then combined with the existing observed modalities to form the complete modality set.
[0037] For each input sample, the system calculates the modal weight of the sample individually based on the quality of its textual, visual, and acoustic features and the status of missing or reconstructed features. This weight is not a globally fixed parameter, but is dynamically output by the routing network based on the complete modal set of the current sample.
[0038] After obtaining the reconstructed mode, the system combines the observed mode and the reconstructed mode to form a complete mode set. If a mode is originally observable, the encoder outputs the true observed features; if a mode is missing, the encoder outputs the reconstructed features. Subsequently, the routing network takes the complete mode set of the sample as input, outputs the corresponding mode weights, and uses Softmax to ensure that each weight is non-negative and sums to 1, ultimately obtaining the fused representation. This invention, through a dynamic modality expert hybrid mechanism, can adaptively allocate mode importance based on the actual available information of different samples, reducing the risk of overuse of low-quality recovered modes.
[0039] In this process, the sparsity of the weight distribution can also be adjusted using the temperature parameter. When the temperature parameter is small, the routing network tends to select a few high-confidence modes; when the temperature parameter is large, the weight distribution is smoother, allowing multiple modes to participate in the fusion. Since this fusion is a weighted convex combination, the fusion result will not amplify noise when the modal features are bounded, which helps to suppress the negative impact of low-quality remodeled modes.
[0040] Furthermore, to mitigate potential biases caused by the recovery modes, this invention introduces an expert balancing mechanism. The system calculates the modal importance based on the individual prediction errors of each mode. (i.e., the first) The importance score of each modality is determined by similarity constraints, with modalities having smaller prediction errors receiving higher importance scores. To align the routing weights with this importance distribution, an entropy regularization term is added. To prevent weights from degenerating into single-modal monopolization, this invention designs a single-modal distillation strategy: preserving the discriminative capabilities of textual, visual, and acoustic single-modal branches respectively, and applying weighted distillation loss. The fusion representation incorporates knowledge from single-modal discrimination. The distillation contribution of unreliable modes is naturally suppressed due to their lower weight.
[0041] Regarding the execution process of the expert balancing mechanism, firstly, the system obtains the corresponding single-modal prediction results using textual, visual, and acoustic single-modal branches respectively. Then, each single-modal prediction result is compared with the true label, and the individual prediction error of that modality is calculated. The smaller the prediction error, the more reliable the modality is for the current sample; the larger the prediction error, the more likely the modality has noise, missing reconstruction bias, or insufficient discriminative information. The system further converts the prediction errors of each modality into a modality importance distribution, for example, by using the inverse of the error and exponential normalization to obtain the importance. Subsequently, the system aligns the modality weights output by the routing network with the importance, so that the routing weights are not only learned freely by the network but also supervised by the reliability of the single modality. That is, when a certain modality is predicted more accurately, its routing weight should be increased accordingly; when a certain modality has a large prediction error, its routing weight should be suppressed. At the same time, the system adds an entropy regularization term to prevent the routing weight from degenerating into single-modality monopoly in the long term, thereby maintaining the effective participation of multiple modality experts. Ultimately, the expert balance loss consists of the weight-importance alignment loss and the entropy regularization term.
[0042] Furthermore, such as Figure 5 As shown, unimodal distillation preserves the discriminative power of each modality and prevents the fusion representation from over-relying on a single modality or being interfered with by low-quality reconstructed modalities. Specifically, the system sets up unimodal prediction branches for text, vision, and acoustic respectively, outputting unimodal prediction results and unimodal latent representations; simultaneously, the fusion network outputs a multimodal fusion representation and the final prediction result. During training, on the one hand, the prediction loss of each unimodal is calculated, so that each modal branch has independent sentiment discrimination ability; on the other hand, the latent representations of each unimodal are weighted and combined according to routing weights to obtain a reliability-aware unimodal knowledge representation. Subsequently, the consistency between the fusion representation and the weighted unimodal knowledge representation is constrained by the distillation loss, allowing the fusion network to absorb effective discriminative information from each unimodal. Since the unimodal knowledge is weighted by routing weights before distillation, unreliable modalities or modalities with large reconstruction errors will have reduced distillation contributions due to their lower weights, while reliable modalities will have a greater impact on the fusion representation. Therefore, unimodal distillation is not a simple averaging of knowledge from each modality, but a selective knowledge transfer based on modal reliability. This mechanism can reduce noise propagation and improve the discriminativeness and robustness of the fused representation.
[0043] Among them, the The single-modal implicit representation or prediction logic vector of each modality is: The implicit representation or prediction logistic vector output by the multimodal fusion branch is The single-mode distillation loss is obtained by weighting the single modes according to the routing weights: ; in, It can represent and The mean square error or cosine distance between them (i.e., the measure used to calculate the error).
[0044] In this invention, the expert balancing mechanism enables dynamic routing to not only "allocate" but also "allocate reasonably," reducing the impact of modes with large recovery errors on the final output; and the use of single-mode distillation ensures that the fusion result maintains overall performance without losing the discrimination information within each mode.
[0045] Furthermore, the emotion sample is input into the emotion recognition fusion model, and a multimodal fusion representation is predicted and output based on all the modal weights. The multimodal fusion representation is then input into the prediction head for prediction, and the final prediction result is output. The sentiment samples are input into each unimodal prediction branch, and the unimodal latent representation of each missing modality is predicted and output according to all the modality weights. Each unimodal latent representation is input into the prediction head for prediction, and the unimodal prediction result of each missing modality is output. Based on the difference between the final prediction result and the multimodal true result, a multimodal prediction loss is constructed; Based on the difference between each single-mode prediction result and the corresponding single-mode true result, a corresponding single-mode loss function is constructed. All single-mode loss functions are then integrated to obtain the single-mode prediction loss.
[0046] Among them, such as Figure 6 As shown, after the diffusion module recovers the missing modes, the system combines the observed modes and the reconstructed modes to form a complete mode set, and obtains the fused representation through a dynamic routing network. The prediction head outputs multimodal prediction results based on the fusion representation. (No. (These are the prediction results), and we can then construct a multimodal prediction loss based on them: ; in, This represents the multimodal prediction loss. Indicates the number of missing modes. Indicates the first A real modality.
[0047] To maintain the independent discriminative power of each single-modal branch, this invention sets up single-modal prediction heads for text, visual, and acoustic modalities respectively. The single-mode prediction results for each mode are as follows: Then the first The single-mode loss of each mode is: ; in, Indicates the first The single-modal loss is applied to each modality. This loss allows text, visual, and acoustic modalities to maintain sentiment discrimination even when inputting individually, preventing the fusion network from relying entirely on a single modality or ignoring effective information within the single modality.
[0048] Furthermore, to ensure that the mode weights output by the routing network are consistent with the mode reliability, this invention estimates mode importance based on single-mode prediction error. For the first... The first sample and the first Each mode has a prediction error that can Represented as: The smaller the error, the more reliable the mode is for the sample.
[0049] Modal importance is constructed based on this error: ; in, This represents a very small constant to prevent the denominator from being zero. Represents mode, Indicates the first The first sample and the first Importance level.
[0050] Among them, the output of the routing network is the first The first sample and the first The modal weights are denoted as The similarity constraint in expert balancing can be expressed as: ; in, Represents similarity constraints, These represent the language modality, visual modality, and audio modality, respectively. express and The mean square error or cosine distance (i.e., the measure used to calculate the error).
[0051] in, Combined with the entropy regularization term weight ( This can be combined with similarity constraints to construct an expert equilibrium loss: .
[0052] Step S40: When optimizing the diffusion module using all the loss terms, the reverse update process of each single modality is adaptively modulated using the constructed error transformation function to output the classification result of each modality feature, and the classification result is visualized to finally obtain all the visualization results of the sentiment sample.
[0053] Among them, the score-based matching loss and reconstruction losses The diffusion mode loss can be constructed, and the fusion and prediction loss can be built based on expert balance loss, multimodal prediction loss, single-modal prediction loss, and single-modal distillation loss. ; in, Indicates fusion and prediction loss; This represents the overall single-modal loss, enabling text, visual, and acoustic modalities to maintain sentiment discrimination capabilities even when inputting individually, thus preventing the fusion network from relying entirely on a single modality or ignoring effective information within the single modality; These represent the weights of the multimodal prediction loss, single-modal prediction loss, expert balance loss, and single-modal distillation loss, respectively.
[0054] Specifically, the joint loss function is obtained by combining the reconstruction loss, the score matching loss, the multimodal prediction loss, the single-modal prediction loss, the expert balance loss, and the single-modal distillation loss; The missing modality recovery model and the emotion recognition fusion model in the diffusion module are trained using the joint loss function, respectively. For each modality feature, a corresponding error transformation function is constructed using the modality feature. During the backpropagation of the joint loss function, the error transformation function is used to adaptively adjust various parameters of the missing modality recovery model and the emotion recognition fusion model to output the classification result corresponding to the modality feature.
[0055] Among them, such as Figure 7 As shown, this invention utilizes a Derf (error function) feature modulation layer to replace the traditional LayerNorm (layer normalization) or BatchNorm (batch normalization) followed by a fixed activation function structure; for the input features Derf employs a learnable error transformation function: ; in, All of these represent learnable parameters. This is the error function.
[0056] During training, the error transformation function is used to jointly update the encoder, diffusion network, routing network, and prediction head through backpropagation. Specifically, when the fusion prediction loss, expert balance loss, single-modal loss, distillation loss, and diffusion loss are backpropagated, the Derf layer will automatically adjust the scaling, translation, and shape parameters according to the gradient, so that the different modal features are adaptively modulated before entering the subsequent transformer or fusion network.
[0057] Unlike normalization methods that rely on batch or sample statistics, Derf is a learnable nonlinear transformation that acts on features point-by-point, without requiring estimation of the mean and variance. Therefore, when different samples have different missing modes, large differences in mode distribution, or noise in the reconstructed mode, Derf can suppress extreme activations and maintain gradient continuity through a smooth, bounded function form, thereby improving training stability and the robustness of the fused representation.
[0058] Furthermore, in another embodiment of this invention, validation was performed on two publicly available benchmark datasets: CMU-MOSI (a widely used benchmark dataset in the field of multimodal sentiment analysis) and CMU-MOSEI (a widely used benchmark dataset in the field of multimodal sentiment analysis). CMU-MOSI contains 2199 samples, and CMU-MOSEI contains 22856 samples. The evaluation metrics used were ACC2 (Binary Accuracy), F1 (F1 Score), and ACC7 (7-class Accuracy). Random missing data ratio settings and fixed missing data protocol settings were examined to verify the stability of the model under different completeness conditions. Specific results are shown in Tables 1, 2, 3, and 4. Table 1: Results under the CMU-MOSI random missing mode
[0059] Table 2: Results of CMU-MOSEI Random Missing Mode
[0060] Table 3: Results under CMU-MOSI fixed missing mode
[0061] Table 4: Results of CMU-MOSEI with Fixed Missing Mode
[0062] In this table, MMIN stands for Missing Modality Imagination Network; GCNet stands for Graph Completion Network; DiCMoR stands for Distribution-consistent Modality Recovery; IMDer stands for Incomplete Modality Diffusion Recovery; and GSDNet stands for Graph Spectral Diffusion Network. As shown in Tables 1, 2, 3, and 4, this invention significantly outperforms existing methods in both CMU-MOSI and CMU-MOSEI. In scenarios with random missing rates, the binary classification accuracy of CMU-MOSI and CMU-MOSEI is generally higher than other models, while their seven-class classification accuracy is comprehensively higher. In scenarios with fixed missing modalities, the seven-class classification accuracy of CMU-MOSI is higher than other models, and the seven-class classification accuracy of CMU-MOSEI is also higher than other models in most cases. These results indicate that this invention is highly competitive in both overall recognition accuracy and fine-grained seven-class classification metrics. As the missing data worsens, the model's prediction accuracy decreases at a significantly slower rate than other models, fully demonstrating the model's excellent robustness.
[0063] Based on the experiments in Tables 3 and 4, this invention further visualizes feature images randomly sampled from the CMU-MOSEI dataset, such as... Figure 8 As shown (in this experiment, Figure 8 In this diagram, L, V, and A represent the language module, visual modality, and audio modality, respectively. Visualization results demonstrate the modality distribution recovered by this invention. Figure 8 The recovered value in the data and the true distribution ( Figure 8 The actual value is closer to the true value in the data, while the comparison method generally suffers from distribution shift and higher degree of dispersion.
[0064] This invention can dynamically allocate modal contributions for different samples, thereby achieving more robust and stable results under missing modal conditions and reducing the misleading effect of recovered modal noise on downstream tasks.
[0065] Furthermore, such as Figure 9 As shown, based on the above-mentioned emotional feature visualization method based on modality mixture and dynamic evidence, the present invention also provides a corresponding emotional feature visualization system based on modality mixture and dynamic evidence, wherein the emotional feature visualization system based on modality mixture and dynamic evidence includes: Modality recognition module 51 is used to acquire the observed modality set and the missing modality set in the emotion sample, use multiple encoders to extract features from the observed modality set and the missing modality set, and map all extracted features to the latent representation space to obtain the observed modality feature set and the missing modality feature set; The modality reconstruction module 52 is used to construct the conditional distribution of the observed modality set and the missing modality set, train a diffusion module using the conditional distribution, the observed modality feature set and the missing modality feature set, and use the diffusion module to recover the missing modality set to obtain a complete modality set; The loss construction module 53 is used to assign corresponding modal weights to different types of modes in the complete modality set, and to construct multiple different loss terms using all the modal weights; The model optimization and result visualization module 54 is used to adaptively modulate the reverse update process of each single modality using the constructed error transformation function when optimizing the diffusion module with all the loss terms, so as to output the classification result of each modality feature and visualize the classification result, and finally obtain all the visualization results of the sentiment sample.
[0066] Furthermore, such as Figure 10 As shown, based on the above-mentioned emotional feature visualization method and system based on modality mixing and dynamic evidence, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 10 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0067] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a sentiment feature visualization program 40 based on modality mixture and dynamic evidence, which can be executed by the processor 10 to implement the sentiment feature visualization method based on modality mixture and dynamic evidence in this application.
[0068] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the emotional feature visualization method based on modality mixture and dynamic evidence.
[0069] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0070] In one embodiment, when the processor 10 executes the sentiment feature visualization program 40 based on modality mixture and dynamic evidence in the memory 20, it implements the steps of the sentiment feature visualization method based on modality mixture and dynamic evidence as described above.
[0071] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a sentiment feature visualization program based on modal mixture and dynamic evidence, and the sentiment feature visualization program based on modal mixture and dynamic evidence, when executed by a processor, implements the steps of the sentiment feature visualization method based on modal mixture and dynamic evidence as described above.
[0072] In summary, this invention provides a method and related equipment for visualizing sentiment features based on modality mixture and dynamic evidence. The method includes: acquiring an observed modality set and a missing modality set from an sentiment sample; extracting features from the observed modality set and the missing modality set using multiple encoders, and mapping all extracted features to a latent representation space to obtain an observed modality feature set and a missing modality feature set; constructing a conditional distribution of the observed modality set and the missing modality set; training a diffusion module using the conditional distribution, the observed modality feature set, and the missing modality feature set; recovering the missing modality set using the diffusion module to obtain a complete modality set; assigning corresponding modality weights to different types of modalities in the complete modality set; constructing multiple different loss terms using all the modality weights; when optimizing the diffusion module using all the loss terms, adaptively modulating the reverse update process of each single modality using a constructed error transformation function to output the classification result of each modality feature; visualizing the classification result to finally obtain all visualization results of the sentiment sample. This invention can dynamically allocate modal contributions for different samples, thereby achieving more robust and stable results under missing modal conditions and reducing the misleading effect of recovered modal noise on downstream tasks.
[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0074] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0075] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for visualizing sentiment features based on modality mixture and dynamic evidence, characterized in that, The sentiment feature visualization method based on modality mixture and dynamic evidence includes: The process involves obtaining the observed modality set and the missing modality set from the sentiment samples, extracting features from the observed modality set and the missing modality set using multiple encoders, and mapping all extracted features to a latent representation space to obtain the observed modality feature set and the missing modality feature set. Specifically, this includes: Obtain the emotional samples input by the user, and judge the emotional samples to determine all observed modalities and all missing modalities in the emotional samples, so as to obtain the set of observed modalities and the set of missing modalities; The text encoder, visual encoder, and acoustic encoder are used to extract features from the text modality, visual modality, and acoustic modality in the observed modality set and the missing modality set, respectively, to obtain the corresponding sequence semantic representation, the corresponding visual temporal features, and the corresponding acoustic features. By using a multilayer perceptron, all the sequence semantic representations, all the visual temporal features, and all the acoustic features are mapped to a latent representation space of the same dimension, the observed text features, observed visual features, observed acoustic features, missing text features, missing visual features, and missing acoustic features are obtained. An observation modality feature set is constructed using the observed text features, the observed visual features, and the observed acoustic features; a missing modality feature set is constructed using the missing text features, the missing visual features, and the missing acoustic features. Constructing conditional distributions for the observed mode set and the missing mode set, and training a diffusion module using the conditional distributions, the observed mode feature set, and the missing mode feature set, specifically includes: A conditional distribution is constructed using the observed mode set and the missing mode set, and a scoring network is then built. The observed modality feature set is input into the scoring network to learn the conditional distribution and obtain a conditional score. The missing modality feature set is input into the scoring network to learn the conditional distribution and obtain an unconditional score. Based on the conditional and unconditional scores, a guiding coefficient score is constructed, and the constructed initial diffusion module is trained using the guiding coefficient score to obtain a diffusion module for reconstructing the missing modality set: ; in, This represents the guiding coefficient score. Indicates the guiding strength parameter. Indicates conditional score. This indicates an unconditional score. This represents the missing modal feature set. Indicates diffusion time. Represents the set of observed modal features. Indicates missing features; The missing mode set is recovered using the diffusion module to obtain the complete mode set; Assign corresponding modal weights to different types of modes in the complete modality set, and construct various different loss terms using all the modal weights; When optimizing the diffusion module using all the aforementioned loss terms, the backward update process of each single modality is adaptively modulated using the constructed error transformation function to output the classification result of each modality feature, and the classification result is visualized to finally obtain all the visualization results of the sentiment sample.
2. The emotional feature visualization method based on modality mixture and dynamic evidence according to claim 1, characterized in that, The process of recovering the missing mode set using the diffusion module to obtain the complete mode set specifically includes: The emotional samples are input into the missing modality recovery model of the diffusion module. During the forward diffusion stage, the missing modality recovery model gradually adds Gaussian noise to each missing modality feature in the missing modality set, converting each missing modality feature into corresponding pure noise. In the backdiffusion phase, for each missing modal feature, the pure noise, the current diffusion time step, and the observed modal feature set are input into the scoring network. The scoring network predicts the denoising direction of the missing modal feature to obtain the recovered modal feature of the missing modal feature. By integrating all the recovered modal features and the observed modal feature set, a complete modality set of the emotion sample is obtained.
3. The emotional feature visualization method based on modality mixture and dynamic evidence according to claim 1, characterized in that, The process of assigning corresponding mode weights to different types of modes in the complete mode set, constructing multiple different loss terms using all the mode weights, and constructing an error transformation function using the mode features in the complete mode set specifically includes: The complete modality set is input into the routing network, which analyzes the missing status of the sentiment samples based on the complete modality set and outputs the modality weight of each single modality. Based on all the modal weights, the sentiment sample is predicted as a whole and as a single modal, respectively, to obtain the final prediction result and multiple single modal prediction results. The final prediction result and all the single modal prediction results are then used to construct the multimodal prediction loss and the single modal prediction loss, respectively. Based on all the single-modal prediction results, the importance of each single mode is constructed, and alignment constraints are applied to all the importance and the corresponding modal weights to obtain the expert balance loss. Calculate the difference between the multimodal prediction loss and the single-modal prediction loss, and construct the single-modal distillation loss based on the difference; Calculate the representation distance between each missing modality feature and its corresponding recovered modality feature, and construct the reconstruction loss and score matching loss using all the representation distances.
4. The emotional feature visualization method based on modality mixture and dynamic evidence according to claim 3, characterized in that, The step involves performing overall prediction and single-modal prediction on the sentiment sample based on all the modal weights, obtaining a final prediction result and multiple single-modal prediction results, and constructing a multimodal prediction loss and a single-modal prediction loss using the final prediction result and all the single-modal prediction results, specifically including: The emotion sample is input into the emotion recognition fusion model, and a multimodal fusion representation is predicted and output based on all the modal weights. The multimodal fusion representation is then input into the prediction head for prediction, and the final prediction result is output. The sentiment samples are input into each unimodal prediction branch, and the unimodal latent representation of each missing modality is predicted and output according to all the modality weights. Each unimodal latent representation is input into the prediction head for prediction, and the unimodal prediction result of each missing modality is output. Based on the difference between the final prediction result and the multimodal true result, a multimodal prediction loss is constructed; Based on the difference between each single-mode prediction result and the corresponding single-mode true result, a corresponding single-mode loss function is constructed. All single-mode loss functions are then integrated to obtain the single-mode prediction loss.
5. The emotional feature visualization method based on modality mixing and dynamic evidence according to claim 3 or 4, characterized in that, When optimizing the diffusion module using all the aforementioned loss terms, the back-update process for each single mode is adaptively modulated using the constructed error transformation function to output the classification result for each mode feature, specifically including: By combining the reconstruction loss, the score matching loss, the multimodal prediction loss, the unimodal prediction loss, the expert balancing loss, and the unimodal distillation loss, a joint loss function is obtained; The missing modality recovery model and the emotion recognition fusion model in the diffusion module are trained using the joint loss function, respectively. For each modality feature, a corresponding error transformation function is constructed using the modality feature. During the backpropagation of the joint loss function, the error transformation function is used to adaptively adjust various parameters of the missing modality recovery model and the emotion recognition fusion model to output the classification result corresponding to the modality feature.
6. A sentiment feature visualization system based on modality mixture and dynamic evidence, characterized in that, The sentiment feature visualization system based on modal mixture and dynamic evidence is used to implement the sentiment feature visualization method based on modal mixture and dynamic evidence as described in any one of claims 1-5, wherein the sentiment feature visualization system based on modal mixture and dynamic evidence comprises: The modality recognition module is used to acquire the observed modality set and the missing modality set in the emotion sample, use multiple encoders to extract features from the observed modality set and the missing modality set, and map all the extracted features to the latent representation space to obtain the observed modality feature set and the missing modality feature set; The modality reconstruction module is used to construct the conditional distribution of the observed modality set and the missing modality set, train a diffusion module using the conditional distribution, the observed modality feature set, and the missing modality feature set, and use the diffusion module to recover the missing modality set to obtain a complete modality set; The loss construction module is used to assign corresponding modal weights to different types of modes in the complete modality set, and to construct multiple different loss terms using all the modal weights; The model optimization and result visualization module is used to adaptively modulate the reverse update process of each single modality using the constructed error transformation function when optimizing the diffusion module with all the loss terms, so as to output the classification result of each modality feature, and to visualize the classification result, so as to finally obtain all the visualization results of the sentiment sample.
7. A terminal, characterized in that, The terminal includes: a memory, a processor, and a sentiment feature visualization program based on modal mixture and dynamic evidence stored in the memory and executable on the processor. When the sentiment feature visualization program based on modal mixture and dynamic evidence is executed by the processor, it implements the steps of the sentiment feature visualization method based on modal mixture and dynamic evidence as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a sentiment feature visualization program based on modal mixture and dynamic evidence, which, when executed by a processor, implements the steps of the sentiment feature visualization method based on modal mixture and dynamic evidence as described in any one of claims 1-5.
Citation Information
Patent Citations
Hybrid mode expert emotion recognition method and system
CN119089259A
Multimodal emotion recognition method and system based on hypergraph diffusion and evidence fusion, terminal and storage medium
CN121009512A