Cognitive-driven progressive alignment mode self-adaptive multi-mode emotion recognition method and system

Through the cognitive-driven progressive alignment modal adaptation method, the problem of dynamic adaptation and cross-scene semantic mismatch of multimodal emotion recognition in an open environment is solved, and the robustness of multimodal emotion recognition and cross-scene migration capabilities are achieved, and zero-sample reasoning of any modal combination is supported.

CN120408405APending Publication Date: 2025-08-01HANGZHOU DIANZI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510471075.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional multimodal emotion recognition methods are difficult to adapt to the dynamic changes of modal combinations in an open environment, and cross-scene semantic mismatch and fragmentation of cognitive frameworks lead to limited cross-scene migration capabilities.

Method used

Using a cognitively driven progressive alignment modal adaptive method, a breakthrough in the robustness of multimodal emotion recognition and cross-scene migration capabilities are achieved by constructing a dynamic decoupled feature fusion architecture and a hierarchical semantic alignment mechanism, including an adaptive progressive self-attention classification module and an progressive alignment paradigm module.

Benefits of technology

It significantly improves the robustness and accuracy of multimodal emotion recognition, supports zero-sample reasoning capabilities for any modal combination, and enhances the generalization performance of emotion recognition in an open environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408405A_ABST
    Figure CN120408405A_ABST
Patent Text Reader

Abstract

The invention discloses a cognitive-driven progressive alignment mode self-adaptive multi-mode emotion recognition method and a cognitive-driven progressive alignment mode self-adaptive multi-mode emotion recognition system. The method comprises the following steps: establishing a multi-modal emotion recognition model comprising an adaptive progressive self-attention classification module and a progressive alignment normal form module; and the multi-modal features are output to a classifier after passing through two layers of cross-modal attention units and one layer of adaptive feature weighted fusion unit in the adaptive progressive self-attention classification module. The features output by the two layers of cross-modal attention units respectively pass through an adaptive feature weighted fusion unit to obtain fusion features; and optimizing a multi-modal emotion recognition model through a three-level progressive semantic alignment mechanism in a progressive alignment normal form module according to the two obtained fusion features and a classification result output by the classifier. According to the method, a collaborative architecture of adaptive feature fusion and hierarchical semantic alignment is provided, and the robustness and accuracy of multi-modal emotion recognition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multi-modal emotion recognition, and particularly relates to a cognitive-driven progressive alignment modal adaptive multi-modal emotion recognition method and system. Background Art

[0002] As a core technology for human-computer interaction and mental health monitoring, multi-modal emotion recognition has made significant progress in fields such as depression screening and autism intervention by integrating text, speech, vision, and physiological signals. However, traditional methods are limited by the closed-scene assumption and predefined evaluation system, and it is difficult to cope with the challenges of dynamics and continuity in open environments. The specific technical bottlenecks are as follows:

[0003] 1. Dimensional tearing of the cognitive framework and evaluation system: Existing methods adopt a discrete emotion label system (such as 6 basic emotions or 26 fine-grained classifications), but the emotional state is essentially a continuous and gradual spectrum. Different scenarios (clinical diagnosis, social media) adopt heterogeneous evaluation criteria (discrete classification, continuous dimension scoring, generative description), resulting in a mismatch in the distribution of the feature space. This essential conflict between the discrete cognitive system and the continuity of emotions limits the construction of cross-scene semantic representations.

[0004] 2. Dynamic heterogeneity of modal combinations: In practical applications, the modal configuration has spatio-temporal heterogeneity: the visual modality may rely on facial micro-expressions (CAER) or body language (BoLD), the audio modality focuses on prosody (IEMOCAP) or speech content (MSP-IMPROV), and the text modality needs to consider both semantics (GoEmotions) and pragmatics (Sarcasm Corpus). Traditional models adopt a fixed encoder and a static fusion strategy. When a modality is missing, the cross-modal association mechanism fails, and the zero-shot transfer ability is limited.

[0005] 3. Structural limitations of cross-scene generalization: The interaction of the above problems forms a systematic bottleneck: a static architecture (such as late fusion) cannot adapt to the dynamic changes of modal combinations, and the parameterized interaction mechanism faces the problem of combinatorial explosion in an open environment. For example, the assumption of relying on the synchronization of expressions and speech in a video scenario fails when the audio is missing, resulting in the failure of reconstructing the cross-modal association pattern. Although current advanced models introduce dynamic fusion, they still assume that all modalities are available and lack the ability to adapt to the heterogeneity of the evaluation system. Summary of the Invention

[0006] The object of the present invention is to provide a cognition-driven progressive alignment modality adaptive multi-modal emotion recognition method and system for the problems existing in the prior art, such as difficult dynamic adaptation of modal combinations, cross-scenario semantic mismatch, and fragmentation of cognitive frameworks. By constructing a dynamic decoupling feature fusion architecture and a hierarchical semantic alignment mechanism, the robustness improvement of multi-modal emotion recognition in an open environment and the breakthrough of cross-scenario migration ability are realized.

[0007] In a first aspect, the present invention provides a cognition-driven progressive alignment modality adaptive multi-modal emotion recognition method, which includes: establishing a multi-modal emotion recognition model; the multi-modal emotion recognition model includes an adaptive progressive self-attention classification module and a progressive alignment paradigm module;

[0008] The multi-modal features are output to a classifier after passing through two layers of cross-modal attention units and one layer of adaptive feature weighted fusion unit in the adaptive progressive self-attention classification module.

[0009] The features output by the two layers of cross-modal attention units respectively obtain fused features through the adaptive feature weighted fusion unit; the two obtained fused features and the classification result output by the classifier are in the progressive alignment paradigm module, and the multi-modal emotion recognition model is optimized through a three-level progressive semantic alignment mechanism.

[0010] The measured multi-modal data is input into the adaptive progressive self-attention classification module in the multi-modal emotion recognition model with optimized input parameters, and the adaptive progressive self-attention classification module outputs an emotion recognition result.

[0011] Preferably, the weighted fusion process of the adaptive feature weighted fusion unit is: respectively performing linear transformation on the input modal features and generating non-negative saliency scores through an activation function; normalizing the non-negative saliency scores of each modality along the modality dimension to obtain a dynamic weight tensor; respectively performing channel-wise multiplication of each dynamic weight tensor with the input features of the corresponding modality, and then performing cross-modal summation normalization on the obtained features to obtain a normalized fusion feature.

[0012] Preferably, the unified loss function of the three-level progressive semantic alignment mechanism includes a primitive alignment loss, a triplet contrast loss, and a classification loss.

[0013] Preferably, the output features of the first - layer cross - modal attention unit pass through the adaptive feature weighted fusion unit. The obtained modality - invariant features are combined with the label semantic features to calculate the cosine distance, which is used as the primitive alignment loss. The output features of the second - layer cross - modal attention unit are fused through the adaptive feature weighted fusion unit. The obtained modality - invariant features are combined with the label semantic prototype matrix to calculate the triplet contrast loss. The triplet contrast loss includes the cross - modal semantic alignment loss, the multi - modal feature cohesion loss, and the label semantic distillation loss, which are weighted and summed. The classification result output by the classifier is input to the decision generalization layer to calculate the classification loss.

[0014] Preferably, the cross - modal attention unit adopts a double - dynamic layer normalization architecture; adjustable layer normalization modules are inserted before and after the fully - connected layers of the query mapping and the key - value mapping in the cross - modal attention unit; the normalization parameters in the adjustable layer normalization module are dynamically generated according to the current modality combination.

[0015] Preferably, the multi - modal features include one or more of text modality features, visual modality features, and audio modality features.

[0016] Preferably, when some modalities are missing in the multi - modal data input to the multi - modal emotion recognition model, the dynamic weight tensors of the corresponding channels in the adaptive feature weighted fusion unit are suppressed and reduced, and the recognition results are still output through two layers of cross - modal attention units, one layer of adaptive feature weighted fusion unit, and the classifier.

[0017] Preferably, when adding new modalities, there is no need to adjust the network structure, and seamless expansion is achieved through online weight assignment.

[0018] In a second aspect, the present invention provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The memory stores the computer program; the processor executes the foregoing multi - modal emotion recognition method.

[0019] In a third aspect, the present invention provides a readable storage medium, which stores a computer program; when the computer program is executed by a processor, it is used to implement the foregoing multi - modal emotion recognition method.

[0020] In a fourth aspect, the present invention provides a multi - modal emotion recognition system with cognitive - driven progressive alignment modality adaptation, which includes a multi - modal data acquisition module and a multi - modal emotion recognition module. The multi - modal data acquisition module is used to acquire one or more of text modality features, visual modality features, and audio modality features, and input them into the multi - modal emotion recognition module.

[0021] The multi-modal emotion recognition module includes an adaptive progressive self-attention classification module and a progressive alignment paradigm module. The adaptive progressive self-attention classification module includes a data encoding module, a two-layer cross-modal attention module, an adaptive feature weighted fusion unit, and a classifier. The progressive alignment paradigm module includes a label encoding module, a primitive alignment layer, a feed-forward layer, a concept abstraction layer, and a decision generalization layer.

[0022] The data encoding module extracts features from the input multi-modal data to obtain multi-modal features. The label encoding module is used to extract label semantic features. The cross-modal attention module includes a cross-modal attention unit and an adaptive feature weighted fusion unit. The multi-modal features output by the data encoding module pass through the cross-modal attention unit in the two-layer cross-modal attention module, and then through one layer of the adaptive feature weighted fusion unit, and are input into the classifier for classification.

[0023] The fusion features output by the first-layer cross-modal attention module are input into the primitive alignment layer, and the primitive alignment loss is calculated with the label semantic features output by the label encoding module; the fusion features output by the second-layer cross-modal attention module are input into the concept abstraction layer, and the triplet contrast loss is calculated with the label semantic prototype matrix generated by the label semantic features through the feed-forward layer. The output result of the classifier is input into the decision generalization layer to calculate the cross-entropy loss as the classification loss. The primitive alignment loss, the triplet contrast loss, and the weighted sum of the classification functions form the unified loss function of the multi-modal emotion recognition model.

[0024] Advantages of the present invention:

[0025] 1. The present invention proposes a collaborative architecture of adaptive feature fusion and hierarchical semantic alignment, which significantly improves the robustness and accuracy of multi-modal emotion recognition; through a three-level progressive alignment mechanism, it simulates the "perception - abstraction - decision" progressive processing mechanism of human emotion cognition, constructs a cross-scenario unified semantic space, and realizes the task compatibility of discrete classification and continuous prediction.

[0026] 2. The present invention innovates the adaptive attention mechanism, proposes a modality decoupled dynamic fusion paradigm, breaks through the dependence of traditional models on preset modality combinations, thus supporting the zero-shot inference ability of arbitrary modality combinations and reducing the model expansion cost; adopts a task-adaptive boundary calibration technology to enhance the generalization performance of emotion recognition in an open environment.

[0027] 3. The present invention designs a vector space mapping model across evaluation systems to achieve the alignment and association of heterogeneous annotation systems in a unified vector space. Description of the Drawings

[0028] Figure 1 It is the overall flowchart of the embodiment of the present invention.

[0029] Figure 2 This is the structural framework diagram of the multi-modal emotion recognition model in the embodiments of the present invention.

[0030] Figure 3 This is the structural framework diagram of the adaptive progressive self-attention classification module in the embodiments of the present invention.

[0031] Figure 4 This is the structural framework diagram of the adaptive feature weighted fusion unit in the embodiments of the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0034] Embodiment

[0035] A cognition-driven progressive alignment modality adaptive multi-modal emotion recognition method, and the multi-modal emotion recognition model used includes an adaptive progressive self-attention classification module APF and a progressive alignment paradigm module PAP.

[0036] The adaptive progressive self-attention classification module APF includes a data encoding module, a two-layer cross-modal attention module APT, an adaptive feature weighted fusion unit AFWF, and a classifier. The progressive alignment paradigm module PAP includes a label encoding module, a primitive alignment layer, a feed-forward layer, a concept abstraction layer, and a decision generalization layer.

[0037] The data encoding module includes encoders corresponding to three modalities and a linear projection layer. The cross-modal attention module APT includes a cross-modal attention unit MA and an adaptive feature weighted fusion unit AFWF. The multi-modal features output by the data encoding module are extracted by the cross-modal attention unit MA in the two-layer cross-modal attention module APT, and then passed through one layer of the adaptive feature weighted fusion unit AFWF and input to the classifier for classification.

[0038] In the first - layer cross - modal attention module APT, the output features of the cross - modal attention unit MA are input to the adaptive feature weighted fusion unit AFWF for processing. The obtained features are input to the primitive alignment layer, and the alignment loss is calculated with the label features output by the label encoding module.

[0039] In the second - layer cross - modal attention module APT, the output features of the cross - modal attention unit MA are input to the adaptive feature weighted fusion unit AFWF for processing. The obtained features are input to the concept abstraction layer, and the triplet contrast loss is calculated with the label features output by the label encoding module.

[0040] The output result of the classifier is input to the decision generalization layer to calculate the cross - entropy loss function.

[0041] The alignment loss in the primitive alignment layer, the triplet contrast loss in the concept abstraction layer, and the cross - entropy loss function in the decision generalization layer are weighted and summed to form the unified loss function of the multi - modal sentiment recognition model.

[0042] Based on the multi - modal sentiment recognition model, the multi - modal sentiment recognition method provided in this embodiment includes the following steps:

[0043] S1: Multi - modal data pre - processing and feature extraction

[0044] S1.1: Input data acquisition: Receive any non - empty subset in the dynamic modal combination space as input, where the text modality T is a natural - language sentence, the visual modality V is an RGB image sequence, and the audio modality A is a time - domain waveform signal.

[0045] S1.2: Frozen encoder processing: The method of text feature extraction is: Use the Qwen2.5 - 7B - Instruct - GPTQ - Int8 encoder to process the input text, and output a semantic vector x T with a dimension of d T = 4069; The method of visual feature extraction is: Use the CLIP - ViT - L / 14 model to process an image with a resolution of 336×336, and output where d V = 1024; The method of audio feature extraction is: Process the 16kHz sampled audio through the HuBERT - Large model, and output d A = 1024.

[0046] S1.3: Linear projection layer: Uniformly map the features of each modality to the shared dimension d = 1024 to obtain the features h of each modality adapted to this task m as follows:

[0047]

[0048] where is a learnable parameter matrix; x m is the original feature output by the feature extractor; b m is the bias term.

[0049] S2: Adaptive Progressive Transformer (APT) processing

[0050] The adaptive progressive self-attention classification module adopts a hierarchical progressive feature refinement strategy. In each layer of processing, first, cross-modal attention (ModalityAttention) is used to perform interactive modeling on the input multi-modal features: ① Embed the modality dimension into the attention feature sequence dimension to construct a universal calculation framework independent of the number of modalities; ② Introduce dynamic layer normalization (LayerNorm) into the standard multi-head attention mechanism, and apply adaptive adjustable normalization operations before and after the fully connected layer of the query / key-value mapping. This design ensures the consistency of the model structure under any modality combination input. The features fused by cross-modal attention in the cross-modal attention module APT are transmitted along two paths: ① Longitudinal depth refinement: Transmit backward for deep feature abstraction, and achieve progressive evolution of the cognitive level through hierarchical stacking; ② Transverse semantic alignment: Generate the feature vector (such as f sample or H multi ) corresponding to the sample at the corresponding level through weighted fusion by the adaptive feature weighted fusion unit AFWF, and calculate the cosine distance or contrast loss with the label feature (f label or H label ) at the corresponding level. This two-way information flow design realizes the collaborative optimization of feature representation in terms of depth and breadth, and constructs a multi-level semantic foundation for cross-scene migration.

[0051] The feature extraction and fusion process of the cross-modal attention module APT is as follows:

[0052] S2.1. Cross-modal attention mechanism:

[0053] Embed the modality dimension into the attention sequence to construct the query matrix Q = [h T ; h V ; h A ; Adopt a double dynamic layer normalization architecture: Q' = LayerNorm1(QW Q ), K' = LayerNorm2(QW K ), where the normalization parameters are dynamically generated according to the current modality combination; Calculate the improved attention weight:

[0054]

[0055] S2.2. Adaptive Feature Weighted Fusion (AWAF):

[0056] As the core component of the APT architecture, the Adaptive Feature Weighted Fusion Unit (AFWF) aims to achieve modality - independent fusion through dynamic weight allocation. The operation process of this module is as follows: First, the input text, visual, and audio modality features are respectively passed through the fully - connected layers corresponding to each modality for linear transformation, and then processed by the ReLU activation function to generate non - negative saliency scores s for each modality m ; Subsequently, Softmax normalization is performed along the modality dimension to obtain the dynamic weight tensors α representing the relative importance of each modality m . After aligning each dynamic weight tensor α m to the original feature dimension through the dimension broadcasting mechanism, channel - by - channel multiplication is performed with each modality feature to achieve importance weighting. Finally, cross - modality summation and L2 normalization are performed on the weighted features to obtain the dimension - normalized robust fusion feature z as follows:

[0057]

[0058] S3: Progressive Alignment Paradigm (PAP) processing

[0059] S3.1. Primitive Alignment Layer: The Primitive Alignment Layer aims to solve the fine - grained semantic gap problem of cross - modality sentiment representation by performing atomic - level feature matching between the original multi - modality signals (audio - visual - text modality) of the same sample and the label semantic vector. Its core mechanism is to construct a modality - independent normalized semantic space, and by minimizing the cosine distance between the sample feature vector f sample (extracted from the multi - modality fusion feature, i.e., the output of the Adaptive Feature Weighted Fusion Unit AFWF in the cross - modality attention module APT of the first layer) and the corresponding label prototype f label (generated by the label text encoder through the linear projection layer), precise coupling of each modality within the sample with the label sentiment primitive is achieved. The optimization objective of this process is defined by the primitive alignment loss function L Align :

[0060]

[0061] where N is the number of samples; the feature vectors f sample , f label are both processed by L2 normalization

[0062] cos(·) represents the cosine similarity calculation (7), and the expression is as follows:

[0063]

[0064] S3.2. Concept Abstraction Layer: The concept abstraction layer realizes the adaptive aggregation and hierarchical representation of cross-scenario emotional semantics by constructing a structured contrastive learning framework in the shared semantic space. This layer drives the topological optimization of the semantic space with a ternary contrastive loss system, and its overall objective function is defined as:

[0065]

[0066] where α, β, γ are weight hyperparameters, is the label semantic prototype matrix (generated by f label through the feedforward layer), is the multimodal category feature matrix (the output content of the adaptive feature weighted fusion unit AFWF in the cross-modal attention module APT of the second layer).

[0067] The overall objective function of the concept abstraction layer 's three sub-loss terms respectively achieve the following functions:

[0068] Cross-modal semantic alignment By constraining the cosine distance between H label and H multi force the label prototypes of the same category and the multimodal features to form a unified representation with close distances in the shared space: where the contrastive loss ψ(·,·) calculates the distance penalty for pairs of samples of the same category and the margin constraint for pairs of samples of different categories, effectively eliminating the semantic drift problem between modalities.

[0069] Multimodal feature cohesion Drive the autonomous aggregation of multimodal features of the same category, break through the hard boundary limitations of traditional classification, and form a dynamic clustering structure of the emotional continuous spectrum:

[0070] Label semantic distillation Strengthen the semantic clustering between label prototypes, and achieve the autonomous aggregation of label prototypes of the same category through self-contrastive learning: This process maps discrete labels to dense semantic vectors and constructs a linearly separable category decision boundary through the geometric repulsion between prototypes.

[0071] The contrastive loss ψ(·,·) is used to calculate the distance penalty for pairs of samples of the same category and the distance exclusion constraint for pairs of samples of different categories, realizing the adaptive aggregation and hierarchical representation of cross-scenario emotional semantics. The mathematical form of the contrastive loss function ψ(·,·) is defined as:

[0072]

[0073] where L ∈ {0, 1} N×N is the label clustering mask matrix, satisfying:

[0074]

[0075] Among them, is the number of homogeneous pairs, is the number of heterogeneous pairs.

[0076] S3.3. Decision Generalization Layer: The decision generalization layer realizes the rapid migration of downstream tasks through a task-adaptive boundary calibration mechanism. This layer only needs to fine-tune the lightweight classification head, and can dynamically reconstruct the decision hyperplane. On the premise of ensuring the stability of the core semantic representation, it completes the task adaptation from discrete classification (such as six-category sentiment) to continuous regression (such as dimension prediction). Taking a typical multi-classification scenario as an example, the cross-entropy loss function is used for optimization:

[0077]

[0078] where C is the total number of sentiment categories, is the one-hot encoded true label, is the softmax-normalized predicted probability.

[0079] S3.4. Three-Stage Loss Merging and Technical Implementation:

[0080] Integrate the three-stage optimization objectives into a unified loss function through a weighted fusion mechanism:

[0081]

[0082] Among them, the weight λ1 of the primitive alignment loss dominates the primitive semantic coupling, the weight λ2 of the triplet contrast loss regulates the concept abstraction intensity, and the weight λ3 of the classification loss strengthens the task boundary calibration. Dynamically balance the optimization focus of each stage through hyperparameters.

[0083] After training the multi-modal sentiment recognition model using the unified loss function, input the measured multi-modal data into the data encoding module; the multi-modal features output by the data encoding module pass through two layers of cross-modal attention units MA and one layer of adaptive feature weighted fusion unit AFWF, and then input into the classifier; the classifier outputs the sentiment recognition result.

[0084] At the technical implementation level, the Progressive Alignment Paradigm module PAP uniformly processes multi-source heterogeneous annotation information through a structured semantic mapping framework. In view of the differences in evaluation systems such as discrete sentiment classification (e.g., "angry"), continuous dimension scoring (Valence = 0.7), and open text description ("a bit sarcastic and helpless"), a semantic conversion mechanism with adaptive capabilities is designed: First, various types of annotation information are standardized into a unified sentiment state description template - "The sentiment state of the person in the video is [Label], the valence score is [Value], and the specific manifestation is [Description]". With the help of the Qwen language model's ability to represent complex semantics, such structured descriptions are encoded into semantic vectors f with consistent dimensions label , and benchmark anchor points across evaluation criteria are constructed in the shared semantic space. Through the mediation of natural language, this solution maps the traditional discrete symbol space, continuous numerical space, and open semantic space to a geometrically continuous semantic manifold, providing a cross-domain compatible semantic reference system for multi-modal feature alignment.

[0085] Finally, the model proposed in this embodiment was tested on three publicly available datasets related to multi-modal sentiment and recognized in the industry, and relevant experiments were conducted. The datasets are as follows:

[0086] In sentiment recognition, this embodiment uses the CH-SIMSv2 and MELD datasets to verify the effectiveness of this embodiment. CH-SIMSv2: A Chinese sentiment benchmark dataset with 5-level annotations and culture-specific expressions, and 10,161 unannotated videos emphasizing non-verbal cues and inconsistencies between annotations and text. MELD: An English multi-dialogue speech dataset containing 13,706 sentences annotated with 7 emotions (anger, joy, etc.) and sentiment polarity, capturing dynamic group interactions.

[0087] To ensure comprehensive model evaluation, we adopted task-specific metrics on each dataset. MELD uses both WAF and accuracy (ACC) to evaluate discrete sentiment classification and sentiment polarity. For CH-SIMSv2, the F1 value and ACC are used for classification tasks (binary classification / three-class classification / five-class classification), while the mean absolute error (MAE) and Pearson correlation coefficient (Corr) are used to measure fine-grained consistency for regression tasks (sentiment intensity prediction). This multi-metric framework ensures robustness in various sentiment recognition scenarios.

[0088] In the experiments of the multi-modal emotion recognition task, the method proposed in this study demonstrated significant advantages on two authoritative datasets. In the evaluation of the fine-grained emotion recognition dataset CH-SIMS v2, the method achieved breakthrough performances in core indicators such as emotion classification accuracy (binary classification: 91.11%, ternary classification: 81.14%, quinary classification: 63.25%) and F1 value (binary classification: 91.12%). Its classification accuracy was improved by more than 7 percentage points compared to the baseline level. At the same time, the method showed excellent performance in emotion intensity prediction, with the mean absolute error reduced to 0.237 and the correlation coefficient increased to 0.858, verifying the model's ability to accurately capture emotion intensity.

[0089] In the test of the dialogue emotion recognition dataset MELD, the method further demonstrated its powerful modeling ability in the dialogue context. Its weighted accuracy (63.91%) and overall accuracy (65.02%) both reached the current highest level, achieving a significant improvement of 3.6 percentage points compared to mainstream methods. Experimental data showed that the method effectively solved key problems such as cross-modal feature alignment and context-dependent modeling through an innovative multi-modal information fusion mechanism, and demonstrated excellent generalization ability in both discrete emotion classification and continuous emotion regression tasks, providing a new technical solution for multi-modal emotion recognition in complex scenarios.

[0090] In this embodiment, ablation experiments were conducted on the Chinese MER2023 dataset for the modality missing scenario. Without any special training for modality missing, zero-shot modality missing tests were directly carried out after tri-modal joint training. Under the single-modal condition, the model demonstrated a powerful strong representation ability for the dominant modality (F1 = 85.574), achieving relative performance improvements of 114.8% and 61.8% compared to the text modality (39.804) and the visual modality (52.889) respectively. In the bimodal combinations, both the audio-visual (88.647) and text-audio (86.221) combinations exceeded the official baseline model of the MER2023 dataset (audio-visual: 86.75), with the audio-visual combination improving by 1.9 percentage points. Notably, the F1 value (90.098) of the method with complete tri-modal input was improved by 3.7 percentage points compared to the MER2023 official baseline (audio + video + text: 86.4), indicating that the method achieved robust feature integration even in the case of modality missing through multi-modal complementarity enhancement and adaptive weighting mechanisms.

[0091] The framework provided by this embodiment establishes a unified semantic space through cross-task sentiment ontology mapping, realizes elastic reasoning for any modal combination through a dynamic modality decoupling mechanism, and simulates human hierarchical cognition ("perception - concept - decision") through step-by-step alignment of the evaluation system. The framework proposed in this embodiment transcends the traditional static fusion paradigm and provides an interpretable solution for emotion computing in open environments by combining dynamic modality adaptation with cognition-driven hierarchical alignment.

[0092] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A cognitive-driven progressive alignment modality adaptive multi-modal sentiment recognition method, characterized in that: Build a multi-modal emotion recognition model; the multi-modal emotion recognition model includes an adaptive progressive self-attention classification module and a progressive alignment paradigm module; The multi-modal features are output to the classifier after passing through two layers of cross-modal attention units and one layer of adaptive feature weighted fusion unit in the adaptive progressive self-attention classification module; The features output by the two layers of cross-modal attention units respectively obtain fused features through the adaptive feature weighted fusion unit; In the progressive alignment paradigm module, the two obtained fused features and the classification result output by the classifier optimize the multi-modal emotion recognition model through a three-level progressive semantic alignment mechanism.

2. The multimodal emotion recognition method according to claim 1, wherein: The weighted fusion process of the adaptive feature weighted fusion unit is as follows: perform linear transformation on each input modal feature respectively, and generate non-negative saliency scores through an activation function; normalize the non-negative saliency scores of each modality along the modality dimension to obtain a dynamic weight tensor; perform channel-wise multiplication of each dynamic weight tensor with the input feature of the corresponding modality respectively, and then perform cross-modal summation normalization on the obtained features to obtain a normalized fused feature.

3. The multimodal emotion recognition method according to claim 1, characterized in that: The unified loss function of the three-level progressive semantic alignment mechanism includes a primitive alignment loss, a triplet contrast loss, and a classification loss.

4. The multimodal emotion recognition method according to claim 3, characterized in that: The output feature of the first layer of cross-modal attention unit passes through the adaptive feature weighted fusion unit, and the obtained modality-invariant feature combines with the label semantic feature to calculate the cosine distance as the primitive alignment loss; The output feature of the second layer of cross-modal attention unit passes through the adaptive feature weighted fusion unit for fusion, and the obtained modality-invariant feature combines with the label semantic prototype matrix to calculate the triplet contrast loss; the triplet contrast loss includes a cross-modal semantic alignment loss, a multi-modal feature cohesion loss, and a label semantic distillation loss calculated by weighted summation; the classification result output by the classifier is input to the decision generalization layer to calculate the classification loss.

5. The multimodal emotion recognition method according to claim 1, wherein: The cross-modal attention unit adopts a dual dynamic layer normalization architecture; An adjustable layer normalization module is inserted before and after the fully connected layers of the query mapping and key-value mapping in the cross-modal attention unit; The normalization parameters in the adjustable layer normalization module are dynamically generated according to the current modality combination.

6. The multimodal emotion recognition method according to claim 1, wherein: The multi-modal features include one or more of text modality features, visual modality features, and audio modality features.

7. The multimodal emotion recognition method according to claim 6, wherein: When some modalities are missing in the multi-modal data input to the multi-modal emotion recognition model, the dynamic weight tensor of the corresponding channel in the adaptive feature weighted fusion unit is suppressed and reduced.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The memory stores a computer program; the processor executes the multi-modal emotion recognition method according to any one of claims 1-7.

9. A readable storage medium stores a computer program; characterized in that: When the computer program is executed by the processor, it is used to implement the multi-modal emotion recognition method according to any one of claims 1-7.

10. A multi-modal emotion recognition system driven by cognitive progressive alignment modal adaptation, comprising a multi-modal data acquisition module and a multi-modal emotion recognition module; characterized in that: The multi-modal data acquisition module is used to acquire one or more of text modality features, visual modality features, and audio modality features, and input them into the multi-modal emotion recognition module; The multi-modal emotion recognition module includes an adaptive progressive self-attention classification module and a progressive alignment paradigm module; the adaptive progressive self-attention classification module includes a data encoding module, a two-layer cross-modal attention module, an adaptive feature weighted fusion unit, and a classifier; the progressive alignment paradigm module includes a label encoding module, a primitive alignment layer, a feed-forward layer, a concept abstraction layer, and a decision generalization layer; The data encoding module extracts features from the input multi-modal data to obtain multi-modal features; the label encoding module is used to extract label semantic features; the cross-modal attention module includes a cross-modal attention unit and an adaptive feature weighted fusion unit; the multi-modal features output by the data encoding module pass through the cross-modal attention unit in the two-layer cross-modal attention module, and then pass through a layer of adaptive feature weighted fusion unit, and are input to the classifier for classification; The fused features output by the first-layer cross-modal attention module are input to the primitive alignment layer to calculate the primitive alignment loss with the label semantic features output by the label encoding module; The fused features output by the second-layer cross-modal attention module are input to the concept abstraction layer to calculate the triplet contrast loss with the label semantic prototype matrix generated by the label semantic features through the feed-forward layer; the output result of the classifier is input to the decision generalization layer to calculate the cross-entropy loss as the classification loss; the weighted sum of the primitive alignment loss, the triplet contrast loss, and the classification function constitutes the unified loss function of the multi-modal emotion recognition model.

Citation Information

Cited By

  • Information processing method and device in intelligent question-answering system, equipment and medium

    CN121233736A