An environment condition weak observation potential causal sentiment recognition method and system for cross-environment low-label scenarios

CN122734512APending Publication Date: 2026-09-11WUHAN TEXTILE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610780572.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0008](3)弱标注机制在不同环境下会发生变化

Benefits of technology

本发明针对现有技术假设弱标签是可靠/固定噪声的直接监督信号的缺陷,重新定义弱标签的生成机制:将弱标签从“训练目标”重构为“潜在真实情感的环境条件有偏观测”,扩展弱标签生成机制,显式引入环境、弱源作为影响弱标签可靠性的变量,提出潜在因果情感后验,融合多模态模型预测、多源弱观测、环境弱源可靠性、样本级模态可靠性,生成软监督信号替代弱标签硬目标,解决了弱标签直接作为训练目标引入错误监督、弱标签跨环境失效的问题,降低了对人工标签的依赖,可在较低的人工标签比例下稳定训练,使得弱标签利用可靠性提升,错误弱标签的干扰被大幅抑制,可适配弱标注机制随环境变化的问题,跨数据集、跨平台迁移性能显著提升;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734512A_ABST
    Figure CN122734512A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for weakly observed latent causal sentiment recognition in cross-environment, low-label scenarios, relating to the field of multimodal sentiment recognition technology. The method mainly includes: training a basic sentiment recognition model using labeled samples; estimating the reliability of each weakly labeled source in various environments; and selecting highly reliable weakly supervised samples. It also involves calculating the latent causal sentiment posterior and sample-level weakly supervised weights; using the total loss from joint weakly supervised training to update the pre-trained basic sentiment recognition model; and using sample-level debiasing coefficients to correct the original prediction results, obtaining sentiment category, sentiment polarity, and reliability explanation information. Implementing the method and system provided by this invention for weakly observed latent causal sentiment recognition in cross-environment, low-label scenarios can avoid the negative impact of environment-related weak label bias and erroneous pseudo-labels on model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal emotion recognition technology, and more specifically, to a method and system for weakly observed latent causal emotion recognition in cross-environment low-label scenarios. Background Technology

[0002] Multimodal emotion recognition aims to integrate information from multiple sources, including text, audio, facial expressions, and gestures, to determine the polarity and category of emotions in video, audio interactions, or human-computer dialogues. Compared to single-text emotion recognition, multimodal emotion recognition can leverage the complementary relationships between language content, acoustic prosody, and visual expressions, making it valuable for applications in scenarios such as online education, smart cockpits, psychological health assessment assistance, intelligent customer service, human-computer interaction, and remote conferencing analysis.

[0003] In real-world deployment environments, multimodal data typically does not satisfy the ideal assumption that "each modality is complete and the training and testing modes are identically distributed." On one hand, sensor occlusion, frame loss, audio noise, network transmission interruptions, poor microphone quality, camera shutdown, or privacy protection strategies can cause partial or complete loss of text, audio, and visual modalities. On the other hand, differences in vocabulary distribution, speaker groups, collection environment, emotional expression habits, and prior category distribution exist between different datasets, platforms, courses, or business scenarios, leading to performance degradation of the model across datasets, platforms, or distributional scenarios.

[0004] Existing research typically focuses on two directions: robust learning of missing modalities and distribution shift correction. For example, some methods address missing modalities through feature reconstruction, modality translation, self-distillation, gated fusion, or modality completion. Other methods reduce the model's dependence on non-causal linguistic cues or category priors through causal inference, counterfactual inference, inverse probability weighting, domain adversarial methods, or bias-corrected classifiers. Robust multimodal sentiment recognition methods, such as CIDer, have attempted to integrate missing modality learning and distribution shift generalization within a unified framework.

[0005] However, the methods described above generally assume that the training set has sufficient and reliable human-generated sentiment labels. In real-world scenarios, manual annotation of multimodal sentiment data is costly. A video often requires simultaneous understanding of textual semantics, speech tone, facial expressions, and context, and sentiment labels also have a degree of subjectivity. Therefore, many business scenarios can only obtain a small number of high-quality labeled samples. For example, only a small number of samples have human-generated labels, while the remaining large number of samples can only obtain unlabeled data, weak labels, coarse-grained labels, rule-based labels, or pseudo-labels generated by teacher models. In one implementation, the proportion of human-generated labels can be 1% to 3%, but this invention is not limited to this proportion.

[0006] Furthermore, in low-label scenarios across environments, weak labels are not fixed and reliable supervisory signals, but rather biased observations of potential genuine emotions. The reliability of the same weak label source may vary in different environments. For example, in online education scenarios, text dictionary rules are relatively reliable when students directly express "I don't understand," but may fail when there is irony, colloquialism, or spelling errors; audio rules may misjudge when microphone quality is poor; and visual expression rules may be unreliable when the camera is off, obstructed, or when students are looking down to take notes. Therefore, low labeling, modality loss, distribution shift, and weak label observation bias often coexist. It should be noted that the low-label weak supervision scenarios here include semi-supervised situations where a small number of labeled samples are used for training with a large number of unlabeled samples, as well as situations where some samples only have weak labels, coarse-grained labels, noisy labels, multi-source pseudo-labels, or weak observations related to the environment. Cross-environment scenarios include, but are not limited to, different datasets, different courses, different platforms, different devices, different data collection conditions, IID (Independent and Identically Distributed) / OOD (Out-of-Distribution) partitioning within the same dataset, and cross-dataset migration scenarios. Overall, existing technologies have the following shortcomings: (1) A small number of labeled samples are insufficient to support stable training. Existing missing modality distillation and causal bias removal methods usually require sufficient labeled data to train the classifier and distillation constraints. When the proportion of manually labeled data is low, the classification boundary is easily affected by a small number of samples, and the model has difficulty learning a stable sentiment semantic representation.

[0007] (2) Using weak labels directly as training targets will introduce incorrect supervision. Weakly supervised training usually uses rule labels, teacher model outputs, or pseudo-labels, but these weak labels are affected by prediction confidence, modality missingness, cross-modal conflict, text bias, and class prior bias. Using only prediction confidence to select pseudo-labels can easily lead to the inclusion of high-confidence but biased erroneous samples into the training.

[0008] (3) Weak labeling mechanisms can change in different environments. Existing weak supervision methods usually assume that the weak label noise pattern is fixed, that is, the error distribution of weak labels relative to the true labels remains stable in the training and deployment environments. However, in cross-dataset, cross-distribution, or cross-platform scenarios, the reliability, confusion pattern, and class prior of the same weak labeling source may change with the environment, and a fixed global weak label calibration is difficult to adapt.

[0009] (4) Fixed distillation weights and fixed counterfactual debiasing coefficients are difficult to adapt to different samples. The missing locations, missing proportions, noise intensity, modal confidence, and bias strengths of different samples are not the same. Uniform distillation strength may lead to over-constraint or under-compensation; uniform debiasing strength may also lead to under-correction or over-correction.

[0010] (5) Existing methods lack a linkage mechanism between “multi-source weak observations – environmental reliability – potential sentiment posterior – distillation and debiasing”. When low labeling, multimodal missing and distribution bias coexist, the model must not only determine whether a certain modality is reliable, but also whether different weak labeling sources are reliable in the current environment, and infer stable potential real sentiment from multiple biased observations.

[0011] Therefore, in low-label scenarios with few labeled samples across environments, how to utilize a large amount of unlabeled or weakly labeled multimodal data, while simultaneously handling modality loss, OOD distribution shift, and cross-dataset migration, and avoiding the negative impact of environment-related weak label bias and erroneous pseudo-labels on model training, is an urgent problem to be solved. Summary of the Invention

[0012] The purpose of this invention is to provide a method and system for identifying potential causal emotions in weakly observed environmental conditions in low-labeled cross-environment scenarios. This method can utilize a large amount of unlabeled or weakly labeled multimodal data, while simultaneously handling modality loss, OOD distribution shift, and cross-dataset migration, and avoids the negative impact of environmentally relevant weak label bias and erroneous pseudo-labels on model training.

[0013] This invention provides a method for identifying latent causal sentiment in weakly observed environmental conditions in low-labeled cross-environment scenarios, comprising the following steps: S1: Based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers and weakly labeled source identifiers, construct a validity mask, a missing location mask and a missing degree statistic to obtain a sample-level missing state vector; S2: Utilize a multimodal fusion network to extract the local representations and initial joint representations of each modality from the multimodal data; S3: Based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, calculate the reliability score of each modality based on the degree of missing data, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. S4: Use labeled samples to train the basic sentiment recognition model to obtain a pre-trained basic sentiment recognition model and initial weak source calibration parameters; S5: For unlabeled or weakly labeled samples, call multiple weak labeling sources based on the weak labeling source identifier to generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; S6: Based on labeled samples, weak observation labels, and environmental identifiers, estimate the reliability, transition matrix, and category prior of each weak labeling source under each environment. Use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. S7: Based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weak labeling source in each environment, and the reliability scores of each modality, the latent causal sentiment posterior and sample-level weak supervision weights are obtained. S8: Based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, the pre-trained basic sentiment recognition model is updated by weakly supervised adaptive causal distillation training using the total loss of weakly supervised joint training, resulting in the updated basic sentiment recognition model. S9: Based on the updated base sentiment recognition model after training, the initial joint representation of the multimodal model is decomposed into sentiment factors and environmental bias factors. Environmental invariant constraints and orthogonal constraints are used to reduce the impact of environmental pseudo-correlation. Counterfactual views of specific weak sources, text words, audio segments or visual segments are constructed by deleting, replacing or masking them, and bias predictions caused by weak observation bias, language bias, label bias or modal pseudo-correlation are estimated. S10: Based on the sentiment factor and bias prediction, the original prediction results are corrected using the sample-level debiasing coefficient to obtain sentiment category, sentiment polarity, and reliability interpretation information.

[0014] This invention also provides a system for recognizing potential causal emotions in weakly observed environmental conditions in low-labeled cross-environment scenarios, comprising the following modules: The cross-environment low-label multimodal data receiving and missing state modeling module is used to: construct a validity mask, a missing location mask, and a missing degree statistic based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers, and weakly labeled source identifiers, and obtain a sample-level missing state vector; The multimodal feature extraction and preliminary fusion module is used to: extract the local representations of each modality and the initial joint representation of multimodal data using a multimodal fusion network; The sample-level modal reliability assessment module is used to: calculate the reliability score of each modality based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, and based on the degree of missing state, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. The labeled sample supervised initialization module is used to: train the basic sentiment recognition model using labeled samples to obtain the pre-trained basic sentiment recognition model and initial weak source calibration parameters; The module for generating multiple weak observations for unlabeled samples is used to: for unlabeled or weakly labeled samples, call multiple weak labeling sources according to the weak labeling source identifier, and generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; The environmental condition weak observation reliability assessment and pseudo-label screening module is used to: estimate the reliability, transition matrix and category prior of each weak labeling source under each environment based on labeled samples, weak observation labels and environmental labels, and use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. The module for generating latent causal sentiment posterior inference and sample-level weakly supervised weights is used to: obtain latent causal sentiment posterior and sample-level weakly supervised weights based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weakly labeled source in each environment, and the reliability scores of each modality. The weakly supervised adaptive causal distillation training module is used to: update the pre-trained basic sentiment recognition model by weakly supervised adaptive causal distillation training using the weakly supervised joint training total loss based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, and obtain the trained and updated basic sentiment recognition model. The module for decoupling sentiment factors and environmental bias factors and estimating counterfactual biases is used to: decompose the initial joint representation of multimodal data into sentiment factors and environmental bias factors based on the updated base sentiment recognition model after training; reduce the impact of environmental spurious correlations by using environmental invariant constraints and orthogonal constraints; construct counterfactual views of specific weak sources, text words, audio segments, or visual segments by deleting, replacing, or masking them; and estimate the bias predictions caused by weak observation bias, language bias, label bias, or modal spurious correlations. The dynamic prediction correction and sentiment result output module is used to: correct the original prediction results based on sentiment factors and bias predictions using sample-level debiasing coefficients to obtain sentiment category, sentiment polarity, and reliability interpretation information.

[0015] The method and system for weak observation of latent causal sentiment recognition in cross-environment low-label scenarios provided by this invention have the following beneficial effects: This invention addresses the shortcomings of existing technologies that assume weak labels are reliable / fixed noise direct supervision signals. It redefines the weak label generation mechanism by reconstructing weak labels from "training targets" to "biased observations of environmental conditions with potential real emotions." This expands the weak label generation mechanism by explicitly introducing the environment and weak sources as variables affecting the reliability of weak labels. It proposes a potential causal emotion posterior and integrates multimodal model prediction, multi-source weak observations, environmental weak source reliability, and sample-level modal reliability to generate soft supervision signals to replace the hard targets of weak labels. This solves the problems of introducing incorrect supervision by directly using weak labels as training targets and the failure of weak labels across environments. It reduces the dependence on manual labels and enables stable training with a lower proportion of manual labels, improving the reliability of weak label utilization. The interference of erroneous weak labels is significantly suppressed. It can adapt to the problem of weak labeling mechanisms changing with the environment and significantly improves cross-dataset and cross-platform transfer performance. This invention addresses the shortcomings of existing technologies that fail to model sample-level differences and employ globally uniform parameters. It designs a sample-level modality reliability assessment stage, dynamically evaluating the reliability of a single sample and single modality across six dimensions: missing data level, cross-modal consistency, feature quality, prediction uncertainty, reconstruction error, and bias strength. Based on the confidence level of the potential causal sentiment posterior, the reliability of each modality, the reliability of weak environmental conditions, and the bias strength, it dynamically generates sample-specific weights, including pseudo-label training weights, hierarchical distillation weights, environmentally invariant constraint weights, and counterfactual debiasing coefficients. This solves the problem that fixed distillation / debiasing strengths cannot adapt to different sample differences, avoiding over-constraint / under-constraint and under-correction / over-correction. The model's adaptability to samples with different missing data patterns and bias strengths is improved, resulting in increased training stability and suppressing the impact of pseudo-label fluctuations on training. This invention addresses the shortcomings of existing technologies, which suffer from isolated steps and a lack of linkage mechanisms. It constructs an end-to-end linkage framework: Environmental Label E → Multi-source Weak Observation Generation → Modal Reliability Assessment → Environmental Weak Source Reliability Modeling → Latent Sentiment Posterior Inference → Sample-level Weight Generation → Adaptive Causal Distillation → Dynamic Debias Correction. The outputs of all steps serve as inputs for subsequent steps, forming a closed loop. Three core steps are added: decoupling of sentiment factors and environmental bias factors, weak source counterfactual consistency, and reliability memory and prior calibration. These steps integrate the processing results of each sub-problem into a unified control system, solving the problem that existing methods cannot simultaneously handle complex scenarios with multiple overlapping issues. It can simultaneously address complex scenarios with low annotation, missing modalities, OOD bias, and weak label bias, significantly improving robustness. It can output explanatory information such as modal reliability, weak source reliability, and posterior confidence, improving model interpretability. Its modular design allows for flexible integration with existing multimodal sentiment recognition backbone networks, enhancing scalability. This invention addresses the shortcomings of existing technologies that do not consider the diversity of modal missing modes by designing a missing state modeling stage. It performs refined modeling of the location, proportion, and degree of modal missing modes, supporting not only whole modal missing modes but also complex missing modes such as local frame missing modes, fragment missing modes, and random feature missing modes. The hierarchical distillation mechanism adaptively adjusts the distillation intensity according to different missing degree modes, solving the problem that existing methods only support whole modal missing modes and the performance degradation under complex missing modes. The robustness of the model to complex missing modes is greatly improved, and it can adapt to various missing situations in real-world scenarios such as sensor occlusion, frame loss, noise, and modal closure.

[0016] Compared to existing technologies, this invention enables training with a large amount of unlabeled or weakly labeled multimodal data under conditions of only a small number of labeled samples, reducing the cost of multimodal sentiment data labeling and decreasing reliance on manual labels in low-label scenarios across environments. This invention does not use all weak or pseudo-labels equally, but rather filters or weights them based on the reliability of weak sources under environmental conditions, modal reliability, potential sentiment posterior confidence, and counterfactual bias strength, reducing the impact of erroneous weak labels on training and improving the reliability of weak label utilization. and By considering environmental conditions and parameters, this invention can describe and correct the inconsistent reliability of the same weak source under different datasets, platforms, devices, or OOD scenarios, adapting to the problem of weak annotation mechanisms changing with the environment. This invention is applicable not only to whole-modal missing data but also to local frame missing data, fragment missing data, random feature missing data, and multimodal mixed missing data, enhancing the model's robustness to complex missing patterns. Through latent sentiment factor invariant learning, counterfactual bias estimation, and sample-level debiasing coefficients, this invention applies different correction intensities to samples with different bias strengths, avoiding undercorrection or overcorrection caused by uniform debiasing, and mitigating erroneous predictions caused by distribution shift and text bias. The sample-level training weights in this invention ensure high reliability. Weakly supervised samples play a greater role, while the impact of low-reliability samples is suppressed. Reliability memory and prior calibration further reduce single-round pseudo-label fluctuations, improving the stability of the training process. In this invention, modality reliability score, weak source reliability score, latent sentiment posterior, pseudo-label reliability score, and dynamic correction coefficient can serve as auxiliary explanatory information to help determine why the model uses a certain weak observation, why it relies more on a certain modality, and why it performs prediction correction, thus improving model interpretability. The environmental condition weak observation modeling, latent sentiment posterior inference, sample-level weight generation, and dynamic bias removal mechanism of this invention can be integrated as independent modules into existing multimodal sentiment recognition networks, without limiting the specific backbone model structure, and have strong system scalability. Attached Figure Description

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the method for weak observation of potential causal sentiment in cross-environment low-label scenarios provided by the present invention; Figure 2 This is a schematic diagram of the framework of the weak observation latent causal emotion recognition system for cross-environment low-label scenarios provided by the present invention. Detailed Implementation

[0018] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0019] Figure 1 This diagram illustrates a method for identifying latent causal sentiment under weak observations of environmental conditions in low-labeled cross-environment scenarios, as shown in this embodiment. In this embodiment, the method for identifying latent causal sentiment under weak observations of environmental conditions in low-labeled cross-environment scenarios includes the following steps: S1: Based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers and weakly labeled source identifiers, construct a validity mask, a missing location mask and a missing degree statistic to obtain a sample-level missing state vector; In one exemplary embodiment, the formula for calculating the missing quantity statistic is:

[0020]

[0021] in, For the sample modality The degree of missing data, The larger the value, the more severe the mode loss. Indicates sample modality Number of positions; Indicates sample modality In the The validity mask of each position, if This indicates that the position is valid. This indicates that the position is missing; , , These represent text, audio, and visual modalities, respectively.

[0022] As an exemplary embodiment, in step S1, position-level validity masks are generated for the three types of modalities, the missing positions and missing proportions of each modality are counted, and sample-level missing state vectors are generated.

[0023] S2: Utilize a multimodal fusion network to extract the local representations and initial joint representations of each modality from the multimodal data; As an exemplary embodiment, in step S2, the multimodal fusion network is a pre-trained language model, a speech feature extractor, a visual feature extractor, or a cross-modal attention layer and a Transformer fusion layer; text / audio / visual feature extractors (such as pre-trained language models, speech feature extractors, and visual feature extractors) are used to extract the sequence features of each modality; the features are mapped to a unified semantic space through the multimodal fusion network (such as a cross-modal attention layer and a Transformer fusion layer) to obtain the local representation of each modality and the initial multimodal joint representation.

[0024] S3: Based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, calculate the reliability score of each modality based on the degree of missing data, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. In one exemplary embodiment, the formula for calculating the intermodal consistency is:

[0025] in, For the sample modality Average consistency with other available modes, i.e., intermodal consistency; Indicates sample In addition to mode The set of available modes other than those mentioned above; and Representing samples respectively modality and modality Feature representation; and Indicates modality and modality Transformation functions that map to a unified semantic space.

[0026] In one exemplary embodiment, the reliability score is calculated using the following formula:

[0027] in, Let m be the reliability score of mode m in sample i; This represents the Sigmoid function or other normalization functions; This is the learnable weight vector or weight matrix corresponding to mode m; For mode m, there is the learnable bias term; This represents a reliability feature vector composed of missing validity, intermodal consistency, modal quality, prediction determinism, reconstruction reliability, and bias suppression. Let m be the missing value statistic for modality m of sample i; To ensure intermodal consistency between mode m of sample i and other available modes; Let m be the modal quality index of mode m of sample i; Let m be the prediction uncertainty of mode m for sample i; Let m be the reconstruction error or offset of mode m of sample i; Let be the bias intensity related to mode m of sample i.

[0028] As an exemplary embodiment, in step S3, sample-level reliability scores for text, audio, and vision are calculated from six dimensions: degree of absence (proportion of absence of each modality), intermodal consistency (calculation of cosine similarity of features of a certain modality with other available modalities), feature quality, prediction uncertainty, reconstruction error, and bias strength, through weighted summation and Sigmoid normalization.

[0029] S4: Use labeled samples to train the basic sentiment recognition model to obtain a pre-trained basic sentiment recognition model and initial weak source calibration parameters; As an exemplary embodiment, in step S4, a basic emotion recognition model is trained using a small number of manually labeled samples to obtain initial classification ability, initial multimodal representation, and initial weak source calibration parameters.

[0030] Specifically, the input consists of a small number of labeled samples of three modalities and corresponding artificial real sentiment labels output from step S1. Using cross-entropy as the loss function, a basic sentiment recognition model is trained using a small number of labeled samples to obtain initial classification ability, initial multimodal representation parameters, and initial weak source calibration parameters. The output is: the initially trained basic sentiment recognition model and the initial weak source calibration parameters.

[0031] As an exemplary embodiment, the basic sentiment recognition model can directly adopt robust multimodal sentiment recognition methods such as CIDer as its basic backbone. These methods are optimized for scenarios with missing modalities and distribution shifts, resulting in more stable initial performance and meeting the scenario requirements of this invention. Alternatively, a pre-trained multimodal large-scale model fine-tuning class can be used, selecting a general multimodal pre-trained model as the basic structure. Only a small number of labeled samples are needed to fine-tune the basic sentiment recognition model, such as CLIP, ALBEF, and BLIP. These models have already completed cross-modal semantic alignment and can quickly obtain stable basic classification capabilities with a small number of samples, outputting initial multimodal joint representations and sentiment prediction probabilities. Alternatively, a dedicated multimodal sentiment recognition model class can be used, selecting mature multimodal classification models designed for sentiment tasks, such as Emotion-Transformer and Multimodal Sentiment Gated Fusion Network (MSP-Net). These models are designed for modal complementarity and temporal dependence in sentiment tasks. The performance of the model has been optimized, and the initial performance is better suited to the needs of emotion recognition. Alternatively, a lightweight concatenation and fusion model can be used, which adopts a simple structure of "single-modal feature extractor + feature fusion layer + classification head". For example, BERT / RoBERTa can be used to extract text semantic features, Wav2Vec2.0 can be used to extract audio prosodic features, and ResNet / VideoMAE can be used to extract visual expression features. The three types of features are concatenated and connected to a fully connected classification layer. A basic emotion recognition model can be obtained by training with a small number of labeled samples. This model is simple to implement, highly interpretable, and suitable for cold start scenarios. Alternatively, a traditional machine learning classifier can be used as the basic implementation. For example, TF-IDF / emotion dictionary features of text, MFCC / prosodic features of audio, and HAAR / LBP expression features of vision can be extracted first. Then, support vector machine (SVM) and decision tree classifier (such as DTREG tool) can be used as classification layers. Basic model training can be completed with a small number of labeled samples. The above models are merely optional implementation examples and do not constitute a limitation on the scope of protection of this invention. The innovation of this invention does not rely on a specific basic model structure. Any of the above models that meet the requirements can be used as the implementation carrier of the basic emotion recognition model.

[0032] S5: For unlabeled or weakly labeled samples, call multiple weak labeling sources based on the weak labeling source identifier to generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; As an exemplary embodiment, in step S5, according to the type of weak annotation source, the corresponding source is called to generate weak observation labels, including text sentiment dictionary rules, unimodal / multimodal teacher model prediction, audio / visual rules, business behavior rules, large model pseudo-labels, etc.; at the same time, the generation confidence, applicable conditions, and modality missing status of the current sample of each weak source are recorded; the means of generating multiple weak observation labels include using dictionary rules, unimodal teacher models, multimodal teacher models, audio / visual rules or large models.

[0033] S6: Based on labeled samples, weak observation labels, and environmental identifiers, estimate the reliability, transition matrix, and category prior of each weak labeling source under each environment. Use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. In one exemplary embodiment, the formula for calculating the transition matrix of each weakly labeled source in each environment is as follows:

[0034] in, This is the transition matrix for each weakly labeled source in each environment, used to describe the transitions in the environment labeling. Below, weakly labeled sources Category of potential true emotions Observation as label The probability of; Indicates the first Weak observation labels generated by a weak annotation source; Indicates the potential true emotional category, The value representing the potential true sentiment category. Indicates the category of weak observation label. A collection of environmental labels, This is a set of weakly labeled sources.

[0035] In one exemplary embodiment, the formula for calculating the reliability of each weakly labeled source in each environment is as follows:

[0036] in, as a weak source In the environment Reliability under these conditions; Represents the normalization function; Indicates environment Embedded representation; and as a weak source The corresponding learnable parameters.

[0037] In one exemplary embodiment, the formula for calculating the reliability of each weakly labeled source in each environment is as follows:

[0038] in, as a weak source In the environment Reliability under these conditions; Indicates the number of sentiment categories; The transition matrix is ​​represented by the first... The entropy of a line.

[0039] As an exemplary embodiment, in step S6, an environmental condition weak observation transition matrix is ​​constructed based on the environment E and the weak source S. The reliability, confusion mode and category prior of different weak sources under different environments are learned by Bayesian estimation or a learnable neural network. Combined with the sample-level modal reliability score, weak supervised samples are screened: those with reliability higher than the threshold are marked as high-reliability weak supervised samples, and those with reliability lower than the threshold only participate in consistency constraints or do not participate in training.

[0040] S7: Based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weak labeling source in each environment, and the reliability scores of each modality, the latent causal sentiment posterior and sample-level weak supervision weights are obtained. As an exemplary embodiment, in step S7, the model prediction probability, multi-source weak observation labels, environmental condition reliability, and modal reliability are fused to obtain the potential true sentiment posterior. ; and then according to The uncertainty, modal reliability, weak source reliability, and bias strength are used to generate pseudo-label training weights, hierarchical distillation weights, environment-invariant constraint weights, and counterfactual debiasing coefficients.

[0041] Specifically, by fusing multimodal prediction probabilities, multi-source weak observation labels, environmental category priors, and weak source reliability, the latent causal sentiment posterior (i.e., the probability distribution of the true sentiment Y, serving as a soft supervision signal) is calculated; The confidence level (entropy value) is combined with modal reliability, weak source reliability, and bias strength to dynamically generate sample-level weak supervision weights, including pseudo-label training weights, hierarchical distillation weights (feature distillation, attention distillation, joint representation distillation), environment-invariant constraint weights, and counterfactual debiasing coefficients.

[0042] In one exemplary embodiment, the formula for calculating the potential causal emotional posterior is:

[0043] in, For the corrected latent causal sentiment posterior, let represent the sample Belongs to the emotion category The latent posterior probability; This indicates a normalization operation; For multimodal models of samples Predicted output For the emotion category The predicted probability; Environmental labeling Next category The prior probability; Indicates sample Available set of weak sources; Indicated in environmental labeling Below, weakly labeled sources Category of potential true emotions Observation as label The probability of; Indicates weak source For the sample The given weak observation labels and This is the adjustment coefficient; as a weak source In the environment Reliability under these conditions.

[0044] In one exemplary embodiment, the formula for calculating the sample-level weak supervision weights is:

[0045]

[0046] in, It is a set of sample-level weakly supervised weights. This indicates the loss weights for samples without pseudo-labels. , These represent the weights for feature distillation, attention distillation, and joint representation distillation, respectively. Indicates the weight of the causal consistency constraint. Indicates the counterfactual debiasing coefficient; This represents a regular function, a neural network, a Bayesian estimation function, or a nonparametric function. For the sample Weak supervision reliability, For the sample Chinese text modality Reliability score; For the sample Mid-frequency modality Reliability score; sample Mid-visual modality Reliability score; Indicates sample The multimodal missingness vector; For environmental labeling; Represents the posterior distribution of latent emotion Entropy; Indicates the number of sentiment categories; This indicates taking the average value; For modal type variables, , , Representing text, audio, and visual modalities respectively; For the sample intermediate mode Reliability score; Indicates sample Available set of weak sources; as a weak source In the environment Reliability under these conditions; Indicates sample Bias strength estimation.

[0047] S8: Based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, the pre-trained basic sentiment recognition model is updated by weakly supervised adaptive causal distillation training using the total loss of weakly supervised joint training, resulting in the updated basic sentiment recognition model. In one exemplary embodiment, the formula for calculating the total loss of the weakly supervised joint training is:

[0048]

[0049]

[0050] in, Total losses in weakly supervised joint training; This indicates supervised loss for labeled samples; This is a set of unlabeled or weakly labeled samples; The loss weights are assigned to unlabeled samples with pseudo-labels. Indicates the posterior of potential causal emotions For soft targets, the distribution is predicted by the model. The soft-label cross-entropy of the prediction results; These represent the feature layer distillation loss, attention layer distillation loss, and joint representation layer distillation loss, respectively. Indicates loss of causal consistency. Indicates loss if the environment remains unchanged. This represents the decoupling loss between emotional factors and environmental bias factors. This indicates the loss of consistency with counterfactual sources. Represents a reliability regularization term. , , , , , , , These represent the weight coefficients of the feature layer distillation loss, attention layer distillation loss, joint representation layer distillation loss, causal consistency loss, environmental invariance loss, decoupling loss between sentiment factors and environmental bias factors, weak source counterfactual consistency loss, and reliability regularization term, respectively. Indicates sentiment category The model predicts the distribution. Represents the latent emotional posterior soft target, and represents the sample. Belongs to the emotion category The latent posterior probability; Indicates sample The reliability of weak supervision; Indicates KL divergence; For the sample Belongs to the category of predicting sentiment The posterior probability; Indicates blocking, deleting, or replacing the first. The latent affective posterior obtained by recalculating after identifying each weak source.

[0051] As an exemplary embodiment, in step S8, multiple types of losses are jointly optimized: labeled sample supervision loss, weak observation posterior loss with potential true sentiment posterior as the soft objective, hierarchical distillation loss (applying distillation constraints to missing modalities at the feature / attention / joint representation layers to enable missing modalities to learn from complete / reliable modalities), environment-invariant loss, sentiment / environment factor decoupling loss, weak source counterfactual consistency loss, and reliability regularization term; the contribution of each loss is dynamically adjusted according to the sample-level weights, and the training weights of low-reliability samples are reduced or they only participate in consistency constraints, outputting the trained and updated multimodal sentiment recognition model.

[0052] S9: Based on the updated base sentiment recognition model after training, the initial joint representation of the multimodal model is decomposed into sentiment factors and environmental bias factors. Environmental invariant constraints and orthogonal constraints are used to reduce the impact of environmental pseudo-correlation. Counterfactual views of specific weak sources, text words, audio segments or visual segments are constructed by deleting, replacing or masking them, and bias predictions caused by weak observation bias, language bias, label bias or modal pseudo-correlation are estimated. In one exemplary embodiment, the expression for decomposing the initial joint representation of the multimodalities into an emotion factor and an environmental bias factor is as follows:

[0053] in, As an emotional factor, Environmental deviation factor and For projection function, This is the initial joint representation of the multimodal expressions.

[0054] In one exemplary embodiment, the expression for the orthogonal constraint is:

[0055] in, For orthogonal constraints, The sentiment factor representation matrix represents a batch of samples. The environmental deviation factor representation matrix represents the batch of samples. This represents the Frobenius norm.

[0056] As an exemplary embodiment, in step S9, during the weakly supervised training process, the multimodal joint representation is decomposed into sentiment factors and environmental bias factors, and the influence of environmental pseudo-correlation is reduced through environmental invariance constraints and orthogonal constraints; at the same time, a counterfactual view is constructed by deleting, replacing or masking specific weak sources, text words, audio segments or visual segments, and the bias prediction caused by weak observation bias, language bias, label bias or modal pseudo-correlation is estimated.

[0057] First, the multimodal joint representation is decomposed into sentiment factor representation and environmental bias factor representation through a projection function. Orthogonal constraints are used to reduce the information overlap between the two, while domain adversarial constraints make it difficult for the sentiment factor representation to predict the environment, while the environmental bias factor representation retains bias information related to the environment. Then, a counterfactual view is constructed (e.g., by masking biased words in the text, deleting audio noise segments, and removing weak observation labels of unreliable weak sources). This view is input into the model to obtain the counterfactual prediction score. The weak source counterfactual consistency loss is calculated using KL divergence to estimate the sample-level bias strength.

[0058] S10: Based on the sentiment factor and bias prediction, the original prediction results are corrected using the sample-level debiasing coefficient to obtain sentiment category, sentiment polarity, and reliability interpretation information.

[0059] In one exemplary embodiment, the expression for the correction is:

[0060] in, This is the final correction result; The original multimodal prediction score; Indicates the sample-level counterfactual debiasing coefficient; The biased prediction score is obtained from the counterfactual view; Indicates the retention coefficient of stable sentiment factors; The invariant predictor is obtained from the potential sentiment factor.

[0061] This embodiment provides a system for identifying potential causal emotions in weakly observed environmental conditions in low-label cross-environment scenarios, including the following modules: The cross-environment low-label multimodal data receiving and missing state modeling module is used to: construct a validity mask, a missing location mask, and a missing degree statistic based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers, and weakly labeled source identifiers, and obtain a sample-level missing state vector; The multimodal feature extraction and preliminary fusion module is used to: extract the local representations of each modality and the initial joint representation of multimodal data using a multimodal fusion network; The sample-level modal reliability assessment module is used to: calculate the reliability score of each modality based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, and based on the degree of missing state, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. The labeled sample supervised initialization module is used to: train the basic sentiment recognition model using labeled samples to obtain the pre-trained basic sentiment recognition model and initial weak source calibration parameters; The module for generating multiple weak observations for unlabeled samples is used to: for unlabeled or weakly labeled samples, call multiple weak labeling sources according to the weak labeling source identifier, and generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; The environmental condition weak observation reliability assessment and pseudo-label screening module is used to: estimate the reliability, transition matrix and category prior of each weak labeling source under each environment based on labeled samples, weak observation labels and environmental labels, and use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. The module for generating latent causal sentiment posterior inference and sample-level weakly supervised weights is used to: obtain latent causal sentiment posterior and sample-level weakly supervised weights based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weakly labeled source in each environment, and the reliability scores of each modality. The weakly supervised adaptive causal distillation training module is used to: update the pre-trained basic sentiment recognition model by weakly supervised adaptive causal distillation training using the weakly supervised joint training total loss based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, and obtain the trained and updated basic sentiment recognition model. The module for decoupling sentiment factors and environmental bias factors and estimating counterfactual biases is used to: decompose the initial joint representation of multimodal data into sentiment factors and environmental bias factors based on the updated base sentiment recognition model after training; reduce the impact of environmental spurious correlations by using environmental invariant constraints and orthogonal constraints; construct counterfactual views of specific weak sources, text words, audio segments, or visual segments by deleting, replacing, or masking them; and estimate the bias predictions caused by weak observation bias, language bias, label bias, or modal spurious correlations. The dynamic prediction correction and sentiment result output module is used to: correct the original prediction results based on sentiment factors and bias predictions using sample-level debiasing coefficients to obtain sentiment category, sentiment polarity, and reliability interpretation information.

[0062] In some embodiments, the above-described method for identifying potential causal emotions in weakly observed environmental conditions in low-labeled cross-environment scenarios can also be implemented in the following ways.

[0063] In this embodiment, the focus is on the following aspects: (1) Receive a small number of labeled multimodal samples, a large number of unlabeled or weakly labeled multimodal samples, environmental labels and weakly labeled source labels; (2) Model missing states for text, audio and visual modalities to obtain sample-level missing state vectors; (3) Obtain or generate weak observation labels for multiple weak annotation sources, and record the confidence and applicability of each weak source; (4) Estimate the sample-level modal reliability score based on the degree of missing data, intermodal consistency, feature quality, prediction uncertainty, reconstruction error, and bias strength; (5) Estimate the weak observation transition matrix or weak source reliability score of environmental conditions based on environmental identifiers and weak source identifiers; (6) Integrate multimodal model prediction, multi-source weak observation, environmental condition weak source reliability and modal reliability to infer the potential causal sentiment posterior Z_y; (7) Generate sample-level weak supervision weights based on the confidence, modal reliability, weak source reliability and bias strength of Z_y. The sample-level weak supervision weights include at least unlabeled sample training weights, hierarchical distillation weights, environmental invariant constraint weights and counterfactual debiasing coefficients. (8) Use labeled samples and highly reliable weakly supervised samples for weakly supervised adaptive causal distillation training; (9) Estimate bias prediction through weak source counterfactual, modal counterfactual or linguistic counterfactual, and dynamically correct sentiment recognition results based on sample-level debiasing coefficients; (10) Output sentiment category or sentiment polarity, and output modal reliability, weak source reliability or potential posterior confidence as explanatory information.

[0064] Among them, the joint generation of environmental condition weak observation reliability modeling, potential causal sentiment posterior inference, sample-level weak supervision weight generation, weak source counterfactual consistency, hierarchical distillation weight and counterfactual debiasing coefficient is suggested as the main protection content of this invention.

[0065] It should be noted that in this embodiment, "weak observation" is a collective term for weak labels, rule labels, teacher model output, pseudo labels, and business behavior labels; "pseudo labels" mainly refer to the predicted labels generated by the current model or teacher model for unlabeled samples.

[0066] In some embodiments, the above-described system for identifying potential causal emotions with weak observations of environmental conditions in low-labeled cross-environment scenarios can also be implemented in the following ways.

[0067] In this embodiment, the system accepts three modalities of input: text, audio, and vision. It receives a small number of labeled samples, a large number of unlabeled or weakly labeled samples, environmental identifiers associated with the samples, and multi-source weak observation labels. The system first models missing states and extracts features for each modality; secondly, it estimates sample-level modal reliability and environmental weak source reliability; then, it fuses multi-modal predictions, multi-source weak observations, and environmental reliability into a latent causal sentiment posterior; subsequently, it generates weakly supervised training weights, hierarchical distillation weights, and counterfactual bias correction coefficients based on this latent sentiment posterior; finally, it completes the sentiment recognition output after weakly supervised training and causal correction.

[0068] Unlike simple pseudo-labeling methods, this embodiment does not directly treat weak labels as training targets. Instead, it considers weak labels, teacher model output, rule labels, and modal predictions as biased observations of potential true sentiment variables. The system explicitly models the impact of environmental variables E and weak labeling sources S on the reliability of weak labels; that is, the weak label generation mechanism is extended from P(~Y|Y) to P(~Y|Y,E,S). Here, Y represents true sentiment, ~Y represents weak observation labels, E represents the environment, and S represents weak labeling sources.

[0069] The latent causal sentiment posterior is not a new artificial label, but rather a probability estimate of the true sentiment category of a sample, derived by the system after integrating multimodal prediction, multi-source weak observation, environmental condition weak source reliability, and modal reliability. This posterior can serve as a soft supervision signal for unlabeled or weakly labeled samples.

[0070] The core of this embodiment lies in: estimating the reliability and confusion patterns of different environments and weak sources through a weak observation mechanism based on environmental conditions; inferring stable sentiment semantics from multimodal inputs and multi-source weak observations through the latent sentiment posterior Z_y; and jointly regulating the training weights of unlabeled samples, the distillation intensity of missing views, and the counterfactual debiasing intensity during the inference stage through sample-level reliability.

[0071] This embodiment can use robust multimodal emotion recognition methods such as CIDer as the basic backbone or comparison baseline, but it is not limited to a specific backbone network structure. The focus of protection in this embodiment is not on a specific Transformer, reconstruction network, or counterfactual text construction method, but on the modeling of weak observation reliability under environmental conditions, the posterior inference of potential causal emotion, and the joint generation mechanism of sample-level training, distillation, and debiased weights.

[0072] like Figure 2 As shown in the figure, the environmental condition weak observation latent causal sentiment recognition system for cross-environment low-label scenarios in this embodiment mainly includes the following modules: cross-environment low-label multimodal data receiving and state modeling module; multimodal feature extraction and preliminary fusion module; multi-source weak observation generation and management module; sample-level modal reliability assessment module; environmental condition weak observation reliability modeling module; latent causal sentiment posterior inference module; sentiment factor and environmental bias factor decoupling module; reliability memory and prior calibration module; weakly supervised adaptive causal distillation module; weak source / modal counterfactual bias estimation module; dynamic prediction correction and sentiment output module; and weakly supervised joint training constraint module.

[0073] The system modules and their connections are as follows: The cross-environment low-label multimodal data receiving and state modeling module is used to receive a small number of labeled multimodal samples, a large number of unlabeled or weakly labeled multimodal samples, and the environment identifier of the sample. The environment identifier can represent the dataset source, training / test distribution, course type, platform, equipment quality, collection conditions, or business scenario. For each sample, this module receives text, audio, and visual modal data and generates a modality validity mask, a missing location mask, and a missing degree statistic, obtaining a sample-level missing state vector.

[0074] The multimodal feature extraction and preliminary fusion module is used to extract feature representations from text, audio, and visual modalities respectively, and map the features of different modalities to a unified semantic space to obtain local representations of each modality and an initial joint multimodal representation. This module can be implemented using Transformer, recurrent networks, convolutional networks, graph neural networks, pre-trained language models, speech feature extractors, visual feature extractors, or other multimodal backbone networks.

[0075] The multi-source weak observation generation and management module is used to receive or generate weak observation labels from multiple weak annotation sources. Weak annotation sources include, but are not limited to, text sentiment dictionary rules, text teacher models, audio rules, visual rules, multimodal teacher models, manually added rules with limited annotation, business behavior rules, or large model pseudo-labels. This module records the label value, confidence level, missing status, and applicable conditions for each weak source.

[0076] The sample-level modality reliability assessment module estimates the reliability score of each modality in the current sample from dimensions such as missing value, feature quality, cross-modal consistency, reconstruction error or representation offset, prediction entropy, counterfactual variation, and text bias strength. This reliability score does not simply indicate the existence of the modality, but rather whether the modality is suitable as a basis for sentiment judgment in the current sample and environment.

[0077] The environmental condition weak observation reliability modeling module is one of the core modules in this embodiment. It is used to learn the reliability, confusion patterns, and category priors of different weakly labeled sources under different environments. This module can be implemented using environmental condition transition matrices, environmental embeddings, source reliability parameters, Bayesian estimation, or learnable neural networks. Its output includes at least a weak source reliability score and a weak label correction probability. This module can avoid using the same set of fixed weak label calibration parameters in different environments, thereby reducing the risk of weak label misleading training in cross-dataset, cross-platform, or OOD scenarios.

[0078] The latent causal sentiment posterior inference module integrates multimodal model predictions, multi-source weak observation labels, environmental condition weak source reliability, and sample-level modal reliability to obtain the posterior distribution Z_y of potential true sentiment. This posterior distribution serves as a soft supervision signal for unlabeled or weakly labeled samples, rather than directly using a single weak label or hard pseudo-label.

[0079] The sentiment factor and environmental bias factor decoupling module is used to decompose the multimodal joint representation into sentiment factor representation and environmental bias factor representation. The sentiment factor representation is used to predict potential true sentiment and try to remain invariant across environments, while the environmental bias factor representation is used to incorporate non-causal variations such as dataset, platform, device, topic, or style.

[0080] The reliability memory and prior calibration module is used to maintain historical statistics of reliability at the sample level, weak source level, or environment level during training iterations. It performs moving average or confidence calibration on weak source reliability, category prior, and potential sentiment posterior to reduce the impact of single-round prediction fluctuations on pseudo-label selection.

[0081] The weakly supervised adaptive causal distillation module is used to train simultaneously using labeled samples and highly reliable weakly supervised samples during the training phase. For missing views, this module imposes constraints on the feature layer, attention layer, joint representation layer, and causal consistency layer based on sample-level distillation weights, enabling the missing view to learn from the complete view, reliable modal view, and stable sentiment posterior.

[0082] The Weak Source / Modal Counterfactual Bias Estimation Module constructs linguistic counterfactual, modal counterfactual, weak source counterfactual, or joint counterfactual views to estimate bias predictions caused by non-causal vocabulary, category priors, weak source bias, modal noise, or cross-modal conflicts. For unlabeled samples, this module can also determine whether current weak observations or false labels may be caused by bias.

[0083] The dynamic prediction correction and sentiment output module dynamically corrects the original prediction results based on the sample-level debiasing coefficient, counterfactual bias estimation results, and latent sentiment posterior. For samples with reliable and unbiased text, the system retains more textual contributions; for samples with strong text bias or significant audio / visual conflicts with the text, the system enhances the counterfactual deduction strength or increases the contributions of non-textual modalities.

[0084] The weakly supervised joint training constraint module is used to jointly optimize the labeled sample supervision loss, weak observation posterior supervision loss, consistency loss, hierarchical distillation loss, environment invariant loss, sentiment / environment factor decoupling loss, weak source counterfactual consistency loss, and reliability regularization term, so that the model can maintain stable training when low labeling, modality missing, and distribution bias coexist.

[0085] In some embodiments, the above-described method for identifying potential causal emotions in weakly observed environmental conditions in low-labeled cross-environment scenarios can also be implemented in the following ways.

[0086] In this embodiment, the method for weak observation of potential causal sentiment in cross-environment low-label scenarios includes the following steps: Step S1. Receiving and Modeling Missing States of Low-Label Multimodal Data Across Environments: Receive a small number of labeled samples, a large number of unlabeled or weakly labeled samples, environmental identifiers and weakly labeled source identifiers, generate validity masks, missing location masks and missing degree statistics for text, audio and visual modalities, and obtain sample-level missing state vectors.

[0087] Step S2. Multimodal feature extraction and preliminary fusion: Extract text, audio and visual features respectively, and obtain local representations and initial joint representations of each modality through a multimodal fusion network.

[0088] Step S3. Sample-level modal reliability assessment: Calculate the reliability scores for text, audio, and visual modalities based on missing mask, feature quality, intermodal consistency, reconstruction uncertainty, prediction confidence, and bias strength.

[0089] Step S4. Supervised initialization with labeled samples: Train the basic emotion recognition model using a small number of manually labeled samples to obtain initial classification ability, initial multimodal representation and initial weak source calibration parameters.

[0090] Step S5. Generation of weak observations from multiple sources for unlabeled samples: Generate multiple weak observation labels using dictionary rules, unimodal teacher models, multimodal teacher models, audio / visual rules, or large models, and record the confidence, applicability, and missing status of each weak source.

[0091] Step S6. Reliability Assessment and Pseudo-Label Screening for Weak Observations under Environmental Conditions: Based on a small number of labeled samples, multi-source weak observations, and environmental labels, estimate the reliability, transition matrix, and category prior of each weakly labeled source under each environment. Combine this with modal reliability to determine whether weakly supervised samples participate in pseudo-label supervision. For unreliable samples, they only participate in consistency constraints or are temporarily excluded from pseudo-label supervision.

[0092] Step S7. Potential causal sentiment posterior inference and sample-level weakly supervised weight generation: Fusing model prediction probabilities, multi-source weakly observed labels, environmental condition reliability, and modal reliability to obtain the potential true sentiment posterior. ; and then according to The uncertainty, modal reliability, weak source reliability, and bias strength are used to generate pseudo-label training weights, hierarchical distillation weights, environment-invariant constraint weights, and counterfactual debiasing coefficients.

[0093] Step S8. Weakly supervised adaptive causal distillation training: Training is performed using labeled samples and highly reliable weakly supervised samples, and various supervision, distillation, consistency and debiasing constraints are adjusted according to sample-level weights.

[0094] Step S9. Decoupling of sentiment factors and environmental bias factors and counterfactual bias estimation: During weakly supervised training, the multimodal joint representation is decomposed into sentiment factors and environmental bias factors, and the influence of environmental spurious correlations is reduced through environmental invariance constraints and orthogonal constraints; at the same time, counterfactual views are constructed by deleting, replacing or masking specific weak sources, text words, audio segments or visual segments, and the bias predictions caused by weak observation bias, language bias, label bias or modal spurious correlations are estimated.

[0095] Step S10. Dynamic prediction correction and sentiment result output: Correct the original prediction results based on the sample-level debiasing coefficient, and output sentiment category, sentiment polarity and reliability explanation information.

[0096] It should be noted that this embodiment plays a role in both the training and inference phases. During training, the system determines whether weakly supervised samples participate in training and the distillation constraint strength based on the weak observation reliability of environmental conditions and the sample-level modal reliability. During inference, the system determines the subtraction strength of the counterfactual bias prediction based on modal reliability and bias strength. Although training and inference operate at different stages, both are based on the latent sentiment posterior. Driven by sample-level reliability results, a unified weakly supervised joint control mechanism is formed.

[0097] It should be noted that the core algorithm module and calculation formula of this embodiment are as follows: Let the set of labeled samples be: The unlabeled or weakly labeled sample set is: Each sample It includes three modalities: text, audio, and visual. ;in, These represent text, audio, and visual modalities, respectively. The environment to which the sample belongs is identified as follows: The set of weakly labeled sources is: ;No. The weak observation labels generated by each weak annotation source are: The category of potential genuine emotions is denoted as: The potential causal emotional posterior notation is: ; (1) Modeling the degree of missing information For any mode The degree of its absence Defined as: (1) in, Indicates sample modality In the A validity mask for each position. If This indicates that the position is valid; if This indicates that the position is missing. The larger the value, the more severe the mode is missing.

[0098] (2) Modeling of intermodal consistency For mode Its average consistency with other available modes It can be defined as: (2) in, Indicates sample In addition to mode Other available modal sets, and These represent the feature representations of the corresponding modes. and This represents the transformation function that maps different modalities to a unified semantic space. The consistency score can be normalized to... Interval.

[0099] (3) Sample-level modal reliability estimation The sample is obtained by comprehensively considering the degree of missing data, intermodal consistency, quality indicators, prediction uncertainty, reconstruction error, and bias strength. intermediate mode Reliability score: (3) in, Represents the modal reliability score. Indicates modal quality index, Indicates uncertainty in forecasting. This represents the reconstruction error or the offset. This indicates the bias strength associated with this mode. This represents the Sigmoid function or other normalization functions.

[0100] (4) Observation transition matrix under weak environmental conditions For the environment and weakly labeled sources Establish the weak observation transition matrix under environmental conditions: (4) in, Indicates the potential true emotional category, The value representing the potential true sentiment category. Indicates the category of weak observation label. Indicates the first Weak observation labels generated by a weak annotation source. Used to describe the environment Below, weak source Category of potential true emotions Observation as label The probability of.

[0101] (5) Environmental conditions and the reliability of weak sources Weak source In the environment reliability It can be computed by embedding the environment: (5-1) It can also be calculated from the uncertainty of the transition matrix: (5-2) in, Indicates environment Embedded representation, and as a weak source The corresponding learnable parameters, The transition matrix is ​​represented by the first... line entropy, This indicates the number of emotion categories. The larger the value, the more trustworthy the vulnerability is in the current environment.

[0102] (6) Possible causal affective posterior inference Assume a multimodal model for samples The predicted probability is:

[0103] The corrected latent causal affect posterior can then be expressed as: (6) in, Indicates sample Belongs to the emotion category The latent posterior probability, Indicates environment Next category The prior probability, Indicates sample Available set of weak sources Indicates weak source For the sample The given weak observation labels and For adjustment coefficients, This indicates a normalization operation.

[0104] (7) Latent posterior confidence and spurious label reliability The confidence level of the latent sentiment posterior can be calculated using entropy, and combined with modal reliability, weak source reliability, and bias strength to generate sample-level weakly supervised reliability: (7) in, Indicates sample Weak supervision reliability, Represents the posterior distribution of latent emotion entropy, This indicates the sample-level bias strength estimation. The larger the value, the more suitable the sample is for weakly supervised training.

[0105] (8) Sample-level weight generation Based on sample-level weakly supervised reliability, modal reliability, and bias strength, dynamically generate the weights required for training and inference: (8) in, This indicates the loss weights for samples without pseudo-labels. , These represent the weights for feature distillation, attention distillation, and joint representation distillation, respectively. Indicates the weight of the causal consistency constraint. This represents the counterfactual debiasing coefficient. This represents a regular function, a neural network, a Bayesian estimation function, or a nonparametric function. Indicates sample The multimodal missingness vector.

[0106] (9) Decoupling of emotional factors and environmental deviation factors Multimodal joint representation Decomposed into emotional factors and environmental deviation factors : (9-1) And information overlap between the two is reduced by orthogonal constraints: (9-2) in, and For projection function, The sentiment factor representation matrix represents a batch of samples. The environmental deviation factor representation matrix represents the batch of samples. This represents the Frobenius norm. The system can be further improved through gradient inversion or domain adversarial methods. Unpredictable environment At the same time Preserve environmental information.

[0107] (10) Weak source counterfactual consistency Counterfactual intervention against weak points, including deletion, blocking, or replacement of the first... After identifying the weaknesses, the latent sentiment posterior is recalculated: (10) in, Indicates blocking, deleting, or replacing the first. The latent sentiment posterior obtained by recalculating after identifying each weak source This represents the KL divergence. This constraint is used to avoid the model from becoming overly dependent on a single weak source, such as relying solely on textual rules or a single teacher model. This indicates the loss of consistency with counterfactual sources.

[0108] (11) Weakly supervised joint training objectives When jointly training with labeled samples and weakly supervised samples, the total loss can be expressed as: (11-1) in, This indicates supervised loss with labeled samples. Indicates the model's predicted distribution. This represents a potential emotional posterior soft goal. Indicates the underlying emotional posterior For soft targets, the distribution is predicted by the model. The soft-label cross-entropy of the prediction results: (11-2) These represent the distillation losses of the feature layer, attention layer, and joint representation layer, respectively. Indicates loss of causal consistency. Indicates loss if the environment remains unchanged. This represents the decoupling loss between emotional factors and environmental bias factors. This indicates the loss of consistency with counterfactual sources. This represents a reliability regularization term.

[0109] (12) Counterfactual bias estimation and dynamic prediction correction set up The original multimodal prediction scores, The biased prediction score is obtained from the counterfactual view. If the invariant predictor is the latent sentiment factor, then the final correction result is: (12) in, This represents the sample-level counterfactual debiasing coefficient. This represents the retention coefficient of stable emotional factors. and The reliability is determined jointly by sample-level reliability, bias strength, and environmental condition weak source reliability. The three predicted scores lie in the same category space, and temperature normalization or scale calibration can be performed first if necessary.

[0110] The above formula system uniformly adopts the following core symbol definitions: : Indicates a weak observation label; : Indicates the potential true emotional category; : Indicates a sample The potential causal emotional posterior; : Indicates environment and weak sources The corresponding weak observation transition matrix; : Indicates environment Weak source Reliability; : Indicates a sample intermediate mode Reliability.

[0111] It should be noted that, firstly, this invention models weak labels as biased observations of environmental conditions that may indicate genuine sentiment, rather than directly using weak labels as training targets. This is achieved through a weak label generation mechanism... Expand to First, this invention can describe the reliability differences of the same weak source in different environments. Second, it introduces environmental condition-based weak observation reliability modeling, which can estimate the reliability, confusion patterns, and category priors of different environments and different weak labeling sources, thereby avoiding the use of fixed global weak label calibration parameters. Third, this invention utilizes latent causal sentiment posteriors. This invention integrates multimodal prediction, multi-source weak observation, environmental condition weak source reliability, and sample-level modal reliability, enabling unlabeled or weakly labeled samples to participate in training in the form of soft posterior. Fourth, this invention unifies weakly supervised training weights, feature distillation weights, attention distillation weights, joint representation distillation weights, environment-invariant constraint weights, and counterfactual bias removal coefficients through a sample-level weakly supervised weight joint generation mechanism. Fifth, this invention links weakly supervised distillation in the training phase, environment-invariant learning, and dynamic counterfactual bias removal in the inference phase, enabling the model to maintain robustness even when low labeling, missing modalities, OOD distribution shifts, and cross-dataset transfer coexist.

[0112] In some embodiments, the above-described system for identifying potential causal emotions with weak observations of environmental conditions in low-labeled cross-environment scenarios can also be implemented in the following ways.

[0113] This embodiment uses an online education scenario as an example. The system receives students' text statements, voice tone, and facial expressions within a class time window. Due to privacy protection, annotation costs, and cross-platform deployment requirements, the system struggles to obtain a large number of artificial emotion labels and can only use teacher feedback, classroom interaction behavior, text rules, voice rules, expression rules, or pseudo-labels from teacher models as weak observations.

[0114] In this scenario, environment E can represent course type, instructor, teaching platform, equipment quality, student group, or school origin. Weak sources S can represent text rules, audio rules, visual rules, classroom behavior rules, or pseudo-labels from a large model. The reliability of weak sources varies in different environments. For example, the reliability of audio weak labels decreases when network quality is poor, the reliability of visual weak labels decreases when the camera is obstructed, and the reliability of text weak labels decreases when students use sarcastic or colloquial expressions.

[0115] This embodiment, through environmental condition weak source reliability and latent sentiment posterior inference, can output the student's emotion category or emotion polarity, while providing explanations for the reliability of text, audio, and visual modalities, as well as the reliability of weak sources. For example, when visual modality is severely lacking, the system reduces the visual contribution; when counterfactual changes in the text indicate strong text bias, the system increases the counterfactual debiasing strength; when multiple weak sources are unreliable in the current environment, the system reduces the weakly supervised training weight of that sample.

[0116] The present invention can also be implemented by an electronic device, which includes a processor, a memory, and a program stored in the memory and executable by the processor. When the program is executed by the processor, it implements the steps of the above-described method for weak observation of potential causal emotion recognition in low-labeled cross-environment scenarios.

[0117] The present invention may also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method steps. The aforementioned electronic device may be a server, personal computer, edge computing device, smart terminal, in-vehicle computing device, or other device with multimodal data processing capabilities.

[0118] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for identifying latent causal sentiment in weakly observed environmental conditions in cross-environment, low-label scenarios, characterized in that... Includes the following steps: S1: Based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers and weakly labeled source identifiers, construct a validity mask, a missing location mask and a missing degree statistic to obtain a sample-level missing state vector; S2: Utilize a multimodal fusion network to extract the local representations and initial joint representations of each modality from the multimodal data; S3: Based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, calculate the reliability score of each modality based on the degree of missing data, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. S4: Use labeled samples to train the basic sentiment recognition model to obtain a pre-trained basic sentiment recognition model and initial weak source calibration parameters; S5: For unlabeled or weakly labeled samples, call multiple weak labeling sources based on the weak labeling source identifier to generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; S6: Based on labeled samples, weak observation labels, and environmental identifiers, estimate the reliability, transition matrix, and category prior of each weak labeling source under each environment. Use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. S7: Based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weak labeling source in each environment, and the reliability scores of each modality, the latent causal sentiment posterior and sample-level weak supervision weights are obtained. S8: Based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, the pre-trained basic sentiment recognition model is updated by weakly supervised adaptive causal distillation training using the total loss of weakly supervised joint training, resulting in the updated basic sentiment recognition model. S9: Based on the updated base sentiment recognition model after training, the initial joint representation of the multimodal model is decomposed into sentiment factors and environmental bias factors. Environmental invariant constraints and orthogonal constraints are used to reduce the impact of environmental pseudo-correlation. Counterfactual views of specific weak sources, text words, audio segments or visual segments are constructed by deleting, replacing or masking them, and bias predictions caused by weak observation bias, language bias, label bias or modal pseudo-correlation are estimated. S10: Based on the sentiment factor and bias prediction, the original prediction results are corrected using the sample-level debiasing coefficient to obtain sentiment category, sentiment polarity, and reliability interpretation information.

2. The method according to claim 1, characterized in that, The formula for calculating the missing degree statistic is: in, For the sample modality The degree of missing data, The larger the value, the more severe the mode loss. Indicates sample modality Number of positions; Indicates sample modality In the The validity mask of each position, if This indicates that the position is valid. This indicates that the position is missing; , , These represent text, audio, and visual modalities, respectively.

3. The method according to claim 1, characterized in that, The formula for calculating the intermodal consistency is: in, For the sample modality Average consistency with other available modes, i.e., intermodal consistency; Indicates sample In addition to mode The set of available modes other than those mentioned above; and Representing samples respectively modality and modality Feature representation; and Indicates modality and modality Transformation functions that map to a unified semantic space; The formula for calculating the reliability score is: in: Let m be the reliability score of mode m in sample i; This represents the Sigmoid function or other normalization functions; This is the learnable weight vector or weight matrix corresponding to mode m; For mode m, there is the learnable bias term; This represents a reliability feature vector composed of missing validity, intermodal consistency, modal quality, prediction determinism, reconstruction reliability, and bias suppression. Let m be the missing value statistic for modality m of sample i; To ensure intermodal consistency between mode m of sample i and other available modes; Let m be the modal quality index of mode m of sample i; Let m be the prediction uncertainty of mode m for sample i; Let m be the reconstruction error or offset of mode m of sample i; Let be the bias intensity related to mode m of sample i.

4. The method according to claim 1, characterized in that, The formula for calculating the transition matrix of each weakly labeled source in each environment is as follows: in, This is the transition matrix for each weakly labeled source in each environment, used to describe the transitions in the environment labeling. Below, weakly labeled sources Category of potential true emotions Observation as label The probability of; Indicates the first Weak observation labels generated by a weak annotation source; Indicates the potential true emotional category, The value representing the potential true sentiment category. Indicates the category of weak observation label. A collection of environmental labels, A set of weakly labeled sources; The formulas for calculating the reliability of each weakly labeled source under each environment are as follows: in, as a weak source In the environment Reliability under these conditions; Represents the normalization function; Indicates environment Embedded representation; and as a weak source The corresponding learnable parameters; The formulas for calculating the reliability of each weakly labeled source under each environment are as follows: in, as a weak source In the environment Reliability under these conditions; Indicates the number of sentiment categories; The transition matrix is ​​represented by the first... The entropy of a line.

5. The method according to claim 1, characterized in that, The formula for calculating the potential causal emotional posterior is as follows: in, For the corrected latent causal sentiment posterior, let represent the sample Belongs to the emotion category The latent posterior probability; This indicates a normalization operation; For multimodal models of samples Predicted output For the emotion category The predicted probability; Environmental labeling Next category The prior probability; Indicates sample Available set of weak sources; Indicated in environmental labeling Below, weakly labeled sources Category of potential true emotions Observation as label The probability of; Indicates weak source For the sample The given weak observation labels and This is the adjustment coefficient; as a weak source In the environment Reliability under these conditions.

6. The method according to claim 1, characterized in that, The formula for calculating the sample-level weak supervision weights is as follows: in, It is a set of sample-level weakly supervised weights. This indicates the loss weights for samples without pseudo-labels. , These represent the weights for feature distillation, attention distillation, and joint representation distillation, respectively. Indicates the weight of the causal consistency constraint. Indicates the counterfactual debiasing coefficient; This represents a regular function, a neural network, a Bayesian estimation function, or a nonparametric function. For the sample Weak supervision reliability, For the sample Chinese text modality Reliability score; For the sample Mid-frequency modality Reliability score; sample Mid-visual modality Reliability score; Indicates sample The multimodal missingness vector; For environmental labeling; Represents the posterior distribution of latent emotion Entropy; Indicates the number of sentiment categories; This indicates taking the average value; For modal type variables, , , Representing text, audio, and visual modalities respectively; For the sample intermediate mode Reliability score; Indicates sample Available set of weak sources; as a weak source In the environment Reliability under these conditions; Indicates sample Bias strength estimation.

7. The method according to claim 1, characterized in that, The formula for calculating the total loss of the weakly supervised joint training is as follows: in, Total losses in weakly supervised joint training; This indicates supervised loss for labeled samples; This is a set of unlabeled or weakly labeled samples; The loss weights are assigned to unlabeled samples with pseudo-labels. Indicates the posterior of potential causal emotions For soft targets, the distribution is predicted by the model. The soft-label cross-entropy of the prediction results; These represent the feature layer distillation loss, attention layer distillation loss, and joint representation layer distillation loss, respectively. Indicates loss of causal consistency. Indicates loss if the environment remains unchanged. This represents the decoupling loss between emotional factors and environmental bias factors. This indicates the loss of consistency with counterfactual sources. Represents a reliability regularization term. , , , , , , , These represent the weight coefficients of the feature layer distillation loss, attention layer distillation loss, joint representation layer distillation loss, causal consistency loss, environmental invariance loss, decoupling loss between sentiment factors and environmental bias factors, weak source counterfactual consistency loss, and reliability regularization term, respectively. Indicates sentiment category The model predicts the distribution. Represents the latent emotional posterior soft target, and represents the sample. Belongs to the emotion category The latent posterior probability; Indicates sample The reliability of weak supervision; Indicates KL divergence; For the sample Belongs to the category of predicting sentiment The posterior probability; Indicates blocking, deleting, or replacing the first. The latent affective posterior obtained by recalculating after identifying each weak source.

8. The method according to claim 1, characterized in that, The expression for decomposing the initial joint representation of multimodalities into sentiment factors and environmental bias factors is as follows: in, As an emotional factor, Environmental deviation factor and For projection function, This is the initial joint representation of the multimodal expressions; The expression for the orthogonal constraint is: in, For orthogonal constraints, The sentiment factor representation matrix represents a batch of samples. The environmental deviation factor representation matrix represents the batch of samples. This represents the Frobenius norm.

9. The method according to claim 1, characterized in that, The expression for the correction is: in, This is the final correction result; The original multimodal prediction score; Indicates the sample-level counterfactual debiasing coefficient; The biased prediction score is obtained from the counterfactual view; Indicates the retention coefficient of stable sentiment factors; The invariant predictor is obtained from the potential sentiment factor.

10. A system for recognizing potential causal emotions in weakly observed environmental conditions in low-labeled cross-environment scenarios, characterized in that, Includes the following modules: The cross-environment low-label multimodal data receiving and missing state modeling module is used to: construct a validity mask, a missing location mask, and a missing degree statistic based on labeled samples, unlabeled or weakly labeled samples, environmental identifiers, and weakly labeled source identifiers, and obtain a sample-level missing state vector; The multimodal feature extraction and preliminary fusion module is used to: extract the local representations of each modality and the initial joint representation of multimodal data using a multimodal fusion network; The sample-level modal reliability assessment module is used to: calculate the reliability score of each modality based on the sample-level missing state vector, the local representation of each modality, and the initial joint representation, and based on the degree of missing state, intermodal consistency, quality index, prediction uncertainty, reconstruction error, and bias strength. The labeled sample supervised initialization module is used to: train the basic sentiment recognition model using labeled samples to obtain the pre-trained basic sentiment recognition model and initial weak source calibration parameters; The module for generating multiple weak observations for unlabeled samples is used to: for unlabeled or weakly labeled samples, call multiple weak labeling sources according to the weak labeling source identifier, and generate multiple weak observation labels, confidence, applicability and missing status of each weak labeling source; The environmental condition weak observation reliability assessment and pseudo-label screening module is used to: estimate the reliability, transition matrix and category prior of each weak labeling source under each environment based on labeled samples, weak observation labels and environmental labels, and use the reliability scores of each modality to screen weak supervised samples to obtain highly reliable weak supervised samples. The module for generating latent causal sentiment posterior inference and sample-level weakly supervised weights is used to: obtain latent causal sentiment posterior and sample-level weakly supervised weights based on the initial joint representation, the predicted probability of the basic sentiment recognition model, weak observation labels, the reliability of each weakly labeled source in each environment, and the reliability scores of each modality. The weakly supervised adaptive causal distillation training module is used to: update the pre-trained basic sentiment recognition model by weakly supervised adaptive causal distillation training using the weakly supervised joint training total loss based on labeled samples, highly reliable weakly supervised samples, potential causal sentiment posterior, sample-level weakly supervised weights and sample-level missing state vectors, and obtain the trained and updated basic sentiment recognition model. The module for decoupling sentiment factors and environmental bias factors and estimating counterfactual biases is used to: decompose the initial joint representation of multimodal data into sentiment factors and environmental bias factors based on the updated base sentiment recognition model after training; reduce the impact of environmental spurious correlations by using environmental invariant constraints and orthogonal constraints; construct counterfactual views of specific weak sources, text words, audio segments, or visual segments by deleting, replacing, or masking them; and estimate the bias predictions caused by weak observation bias, language bias, label bias, or modal spurious correlations. The dynamic prediction correction and sentiment result output module is used to: correct the original prediction results based on sentiment factors and bias predictions using sample-level debiasing coefficients to obtain sentiment category, sentiment polarity, and reliability interpretation information.