Emotion domain adaptation method based on information bottleneck theory decoupling
Through the decoupled emotional domain adaptation method based on information bottleneck theory, emotion-specific features are extracted and cross-domain alignment is performed, the deviation and aberration problems of CLIP model in the application of the emotional domain are solved, and more efficient emotion representation and recognition performance is achieved.
Patent Information
- Application Number
- CN202510211031.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
There are significant deviations in the application of existing CLIP models in the emotional field, especially in terms of cross-domain alignment and emotional representation accuracy, and there is a misalignment between CLIP space and emotional space, resulting in redundancy and inaccuracy in the emotional recognition task.
A decoupled emotional domain adaptation method based on information bottleneck theory is proposed. Through emotion extraction, structural optimization and information reconstruction modules, redundant information in CLIP embedding is decoupled, specific emotions are extracted, and the accuracy and alignment effect of emotional representation is improved through pseudo-label generation and cross-domain alignment mechanisms.
The performance of emotional domain adaptation was significantly improved, and compared with the existing methods, it improved by 38.31% in the ArtPhoto->FI task, and also showed superior effects in other emotional domain adaptation settings, enhancing the model's emotion detection and distinction ability.
Smart Images

Figure CN120146187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a visual emotion analysis technology, in particular an emotion domain adaptation technology based on a large-scale vision-language model; and a feature disentanglement technology, in particular an emotion feature extraction technology based on information bottleneck. Background Art
[0002] With the rapid development of social networks, people are increasingly inclined to express emotions through online picture sharing. This trend has promoted the wide attention of visual emotion recognition (VER) technology, which has important application values in fields such as public opinion monitoring and depression detection. However, the abstractness and subjectivity of emotions have brought huge challenges to the annotation of large-scale datasets.
[0003] In response to the above problems, the emotion domain adaptation (EDA) technology has emerged. Its goal is to use the labeled source domain data to process the unlabeled target domain data. However, the existing methods mainly focus on single-modal and low-level features and cannot effectively capture the complex emotional semantics contained in visual content. For example, in the domain adaptation task, for the mutual adaptation and transfer between the ArtPhoto dataset and the FI dataset (Art<->FI), the average accuracy of the current state-of-the-art method is only 36.26%, which indicates that the EDA technology urgently needs to break through.
[0004] Recently, large-scale vision-language models represented by CLIP have demonstrated excellent zero-shot transfer capabilities through large-scale image-text pre-training and achieved remarkable results in unsupervised domain adaptation (UDA) tasks. However, when dealing with abstract concepts such as emotions, the average accuracy of these methods in the Art<->FI setting has dropped significantly to 40%.
[0005] Through in-depth analysis, we found that there is a significant misalignment between the CLIP space and the emotion space. Since CLIP is pre-trained on millions of image-text pairs to learn general feature representations, there may be redundancy in its text and image embeddings. For the emotion recognition task, only a small number of feature dimensions are effectively activated, while the remaining dimensions are redundant and irrelevant. This problem can be numerically verified through experiments: taking the text embeddings encoded by CLIP as an example, we encoded the prompt text "a photo seems to express some feelings like [emotion]" for different emotion categories, and the calculated cosine similarity shows that the text embeddings between different emotions are highly similar and lack discriminative information.
[0006] The Information Bottleneck (IB) theory provides a principled approach to solving the above problems. This theory extracts task-related information through data compression while eliminating redundant information. Specifically, our goal is to disentangle specific emotional features from CLIP embeddings, that is, to maximize the mutual information between the features and the emotional space while minimizing the mutual information between the features and the CLIP space.
[0007] However, directly applying the Information Bottleneck theory to high-dimensional CLIP embeddings faces significant computational challenges. First, calculating the mutual information in high-dimensional spaces is computationally infeasible. Second, the high-dimensional nature of CLIP embeddings results in an overly sparse feature space, making it difficult to accurately estimate probability distributions. Third, traditional Information Bottleneck methods often require a large number of iterations to converge, which incurs significant computational overhead when dealing with large-scale vision-language models.
[0008] Based on the above problems and challenges, this invention faces two key scientific issues, namely disentangling specific emotional information from redundant CLIP embeddings and performing emotional semantic decoupling using the Information Bottleneck framework while maintaining computational feasibility. These challenges drive researchers to further explore new technical solutions to improve the performance of visual emotion recognition. Summary of the Invention
[0009] The objective of this invention is to address the significant biases in the application of existing CLIP models in the emotional domain, especially in cross-domain alignment and the accuracy of emotional representation. The CLIP model mainly embeds text and images through the semantic space, but the features it embeds are less adaptable to emotional analysis tasks. Especially in emotional representation and emotional classification tasks, there are problems of information redundancy and inaccuracy. A method for emotional domain adaptation based on Information Bottleneck theory decoupling is proposed, aiming to achieve alignment between the CLIP model and the emotional space by effectively separating emotion-specific features and eliminating redundant information.
[0010] This invention solves the above technical problems through an Information Bottleneck-based Emotional Decoupling network (EmoD). This framework includes three main modules: Emotional Extraction (EoT), Structure Optimization (SoT), and Information Reconstruction (IoT). The Emotional Extraction module (EoT) ensures that the extracted features can effectively reflect emotional information and remove irrelevant redundant information by maximizing the mutual information between the separated features and the emotional space; the Structure Optimization module (SoT) restricts the bottleneck capacity of the decoupled features by imposing a normalization penalty, thereby reducing redundant information and improving the discriminative ability of emotional features; the Information Reconstruction module (IoT) reduces information loss through a reconstruction loss mechanism to ensure the semantic integrity of emotion-specific features.
[0011] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0012] An emotion domain adaptation method based on decoupling by the information bottleneck theory, the method is:
[0013] Step 1: Emotion-specific feature disentanglement
[0014] Although the rich semantic representation of CLIP helps to connect low-level pixel information and high-level emotion semantics, due to the training on millions of captions and concepts, its embedding contains redundant information, which may affect the similarity classification performance. To solve this problem, our goal is to disentangle the CLIP embedding into emotion-specific features.
[0015] (1) Emotion extraction (EoT)
[0016] The present invention aims to use the embedding encoder E m to extract the emotion-specific feature z = E m (t) ∈ R k×d , where E m (t) represents the output after the text embedding passes through the encoder; k and d respectively represent the number of emotion categories and the feature dimension; in order to ensure that the extracted features can effectively capture emotion information, it is necessary to maximize the mutual information I(z; e) between the emotion-specific feature z and the emotion space e, which is defined as:
[0017] I(z; e) = E z,e [logp(e|z)] - E e [logp(e)] (3)
[0018] where E z,e represents the expectation calculated for the two variables z and e; p(e|z) represents the posterior distribution of e with respect to z; E e represents the expectation calculated for e; p(e) represents the prior distribution of e; the second term E e [logp(e)] remains unchanged during the optimization process; the first term represents the expected log-likelihood of emotion prediction, which naturally corresponds to the negative value of the cross-entropy loss and captures the predictive ability of the extracted features;
[0019] Based on this insight, we formulate the cross-entropy loss between the emotion-specific feature z and the corresponding label y k from K emotion categories as:
[0020]
[0021] where f() represents the emotion label classifier, represents the k-th element in the softmax output of the K-dimensional vector a, a kDenotes the k-th dimension of the K-dimensional vector a, a i Denotes the i-th dimension of the K-dimensional vector a, y k Is `1' for the correct class and `0' for the remaining classes; The emotion extraction (EoT) optimization process maximizes the mutual information between the extracted emotion-specific features and the emotion space.
[0022] (2) Structure optimization (SoT)
[0023] From the perspective of the information bottleneck, it is necessary to minimize the mutual information between the emotion-specific feature z and the input embedding t to reduce redundant information; Due to the computational complexity of directly calculating the mutual information, inspired by β-VAE, it uses the KL divergence of the standard normal distribution to limit the bottleneck capacity and implements an L on the emotion-specific feature z 2 Regularization term:
[0024] L bottle = E[|z| 2 (5)
[0025] where E represents taking the expectation, and L bottle Represents the loss function under the bottleneck constraint;
[0026] To further enhance the discriminative ability of the disentangled classification embeddings, an orthogonal regularization is proposed to enforce the independence between different classification dimensions:
[0027]
[0028] where F represents taking the F-norm; L ortho Represents the loss function of the orthogonal regularization; p = z T z ∈ R d×d Represents the emotion-specific projection matrix, I ∈ R d×d Is the identity matrix; This regularization enforces the orthogonality constraint between the row vectors of z, ensuring the linear independence between the classification dimensions; The overall optimization objective of the SoT process combines the bottleneck constraint and the orthogonal regularization, and the formula is as follows:
[0029] L SoT = L bottle + L ortho (7)
[0030] (3) Information reconstruction (IoT)
[0031] Considering that excessive feature compression may lead to the loss of semantic information, we introduce a reconstruction loss to retain the information richness of the emotion-specific feature z by enabling the model to recover the original CLIP text embedding;
[0032] Specifically, we model the reconstruction process through the posterior probability p(t|z), which represents the likelihood of reconstructing the original embedding t through the decoder D m given the emotion-specific feature z; the reconstruction quality is quantified by minimizing the mean squared error (MSE) between the original and reconstructed embeddings:
[0033] L IoT = E[|t - D m (z)| 2 (8)
[0034] where D m (z) represents the CLIP embedding reconstructed from the emotion-specific feature z; this reconstruction objective ensures that z captures the features of a specific emotion while preserving the basic semantic properties of the original embedding; through this balanced approach, EmoD achieves strong emotion disentanglement while maintaining the rich semantic context inherent in the CLIP embedding;
[0035] Step 2: Pseudo-label generation
[0036] After obtaining the emotion-specific representation, a projection matrix p = z T z maps all target image embeddings and text embeddings t ∈ R K×d into a shared emotion-specific subspace:
[0037]
[0038] where N T is the number of emotional images in the target domain, v t is the feature embedding of all target images, v i t is the feature embedding of each image, and d is the feature dimension representing the emotion category;
[0039] Given that the text embeddings are explicitly aligned with the emotion space through EoT optimization and their categories are predefined, use the emotion-aligned t proj as the semi-supervised propagation center to generate pseudo-labels for unlabeled target images, as shown in the Figure 1 right panel. For text and image alignment, aggregate the propagation center and all projected target image embeddings:
[0040]
[0041] Then, calculate a sparse similarity matrix and generate pseudo-labels based on the sample distances:
[0042]
[0043] where A = sim(x · xT ) Calculate the cosine similarity between all samples in the subspace; A n Extract the first n values from each row of matrix A. I represents the indicator function; i and j represent the i-th row and j-th column of the matrix. The pseudo-label matrix is initialized as:
[0044]
[0045] After normalizing the similarity matrix A n perform label propagation using conjugate gradient (CG) to obtain the pseudo-labels of the unlabeled target images, where α is a smoothing parameter used to control the propagation speed:
[0046]
[0047] Subsequently, optimize the text encoder by minimizing the pseudo-label-guided target domain loss function and the true supervised source domain loss function :
[0048] L CE = L pl + L gt (14)
[0049] where L CE represents the loss function obtained from Equation (14);
[0050] Step 3: Align the target domain and the source domain
[0051] To achieve effective sentiment domain adaptation and further bridge the domain gap, we introduce cross-domain alignment based on the maximum mean discrepancy (MMD):
[0052] L align = MMD 2 (F s , F t ) (15)
[0053] where F t and F s represent the CLIP-encoded image distributions from the source domain and the target domain, respectively; the Gaussian RBF kernel with bandwidth parameter β is used to calculate the unbiased MMD estimate. During the adaptation process, we jointly optimize the text encoder using the labeled source data and the pseudo-labeled target data to learn domain-invariant sentiment representations. This mechanism ensures the consistency of cross-domain sentiment representations.
[0054] Furthermore, the preparations for the method include the following three points:
[0055] (1) Sentiment domain adaptation
[0056] This paper focuses on EDA, that is, adapting from the labeled source domain to the unlabeled target domain;
[0057] The source domain is represented as where is the source image, and y k is the corresponding label from K emotion categories, and N s is the total number of source images; K is the number of emotion categories; i represents the i-th sample; k represents the k-th category;
[0058] The target domain is represented as consisting of N unlabeled T emotion images ; for simplicity, the subscript i is omitted hereafter; N T represents the number of emotion images in the target domain (the number of target domain samples);
[0059] The source domain and the target domain share the same label space, but the distribution P s (x) of the source domain and the distribution P t (x) of the target domain are different;
[0060] (2) CLIP large vision-language model
[0061] Our framework is based on CLIP, which is one of the most popular pre-trained vision-language models; the main components of CLIP are the image encoder E I and the text encoder E T , which accept image and text inputs respectively; in the EDA task, CLIP incorporates the k-th emotion category into the pre-designed text prompt p k , which is expressed as "A [domain] photo seems to express some feelings like [emotion]". The text prompt p k and the image x i are then encoded as t k = E T (p k ), v i = E I (x i ), sharing the same feature dimension d v ; classification is based on the highest cosine similarity between t k and v i :
[0062]
[0063] Our goal is to maximize the inter-class distance between text embeddings to enhance the discriminative ability of similarity classification;
[0064] (3) Information Bottleneck (IB)
[0065] To separate emotion-specific features from redundant CLIP embeddings, the view of the information bottleneck theory is combined, and its principle is as follows:
[0066] max[I(z; e) - βI(t; z)] (2)
[0067] Where I(·;·) represents mutual information, and β represents the Lagrange multiplier; from the perspective of IB, we formalize a constrained optimization problem aiming to maximize the mutual information between the emotion-specific feature z and the emotion space e, while minimizing the emotion-irrelevant information in the CLIP text embedding t. We implement this optimization through the EoT and SoT modules; however, during the compression process, high-level semantic information may be lost, potentially affecting the model's emotion recognition ability. We further introduce the IoT module to help recover key information.
[0068] Figure 1 Summarizes the framework EmoD of the present invention. The main objective of the present invention is to solve the inconsistency problem between CLIP and the emotion space. The present invention proposes an emotion disentanglement framework (including EoT, SoT, and IoT) to eliminate redundant information and extract emotion-specific features from the text embeddings encoded by CLIP. Subsequently, the present invention uses the extracted features to construct an emotion projection matrix to project the text and image embeddings of CLIP into a shared subspace. Guided by the emotion-aligned text embeddings, the present invention uses label propagation to generate pseudo-labels for the image embeddings. In addition, the present invention calculates the maximum mean discrepancy (MMD) between the source domain and the target domain to achieve cross-domain feature alignment.
[0069] The beneficial effects of the present invention compared with the prior art are as follows: The effects of the present invention are significant. Through the emotion decoupling (EmoD) framework, compared with the existing state-of-the-art emotion domain adaptation (EDA) methods, in the domain adaptation task, the model adapts and migrates from the ArtPhoto dataset to the FI dataset (ArtPhoto->FI), and EmoD improves by 38.31%, mainly due to the fact that CLIP can effectively align low-level pixel semantics with high-level emotion semantics. In addition, compared with the previous unsupervised domain adaptation (UDA) method UniMoS based on CLIP, EmoD improves by 3.96% in the ArtPhoto->FI task and by 3.43% in the EmotionROI->FI task. This effect mainly comes from the emotion-specific feature decoupling module, which enhances the model's ability to detect and distinguish emotion features. Further, EmoD improves by more than 10% compared with the CLIP-based prompting methods (PDA, DAPrompt, DAMP) in the ArtPhoto->FI task, demonstrating its superior ability in emotion semantic extraction. Brief Description of the Drawings
[0070] Figure 1 The EmoD framework proposed for disentangling emotion-specific features, where the dashed line represents the optimization path involving only the target domain data (left), and the schematic diagram of label propagation (right). Specific implementation mode
[0071] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.
[0072] Visual emotion recognition refers to the technology of using computer vision technology to understand human emotional responses to different visual stimuli. First, due to the inherent ambiguity of emotional features, this poses significant challenges to data annotation in the supervised learning paradigm. Second, emotion domain adaptation technology solves this problem by transferring the knowledge of the labeled source domain to the unlabeled target domain. Third, although large-scale vision-language models such as CLIP have shown excellent transfer performance in traditional unsupervised domain adaptation tasks, when generalized to abstract concepts such as emotion, the misalignment between the CLIP space and the emotion space seriously affects the model performance. Finally, the emotion disentanglement framework based on the information bottleneck theory extracts specific emotion features through a disentanglement network, while removing redundant emotion-irrelevant information, and uses feature alignment to reduce inter-domain emotion differences, thereby effectively improving the model performance.
[0073] In terms of emotional semantic representation, the present invention constructs a projection matrix to map the text and image embeddings encoded by CLIP into a shared subspace using emotion-specific features. In this way, the projected text embeddings can be aligned with the emotion space and show a greater separation between emotion categories (especially between positive and negative emotional polarities). In addition, the present invention also uses these aligned emotional text embeddings as supervision guidance and adopts the label propagation algorithm to generate pseudo-labels for the target image embeddings, thereby improving the accuracy and robustness of emotion classification.
[0074] To achieve cross-domain alignment, the present invention adopts the maximum mean discrepancy (MMD) method to achieve alignment between the source domain and the target domain under the condition of unsupervised learning. This method ensures that the emotional representations of the source domain and the target domain are consistent between different domains, so that consistent performance can be shown in different datasets and application scenarios.
[0075] The main technical contributions of the present invention include: First, it reveals the limitations of the CLIP model in emotional domain adaptation, especially the deviations in emotional space alignment and semantic representation; Second, it proposes an Emotion Disentanglement (EmoD) framework, which can effectively disentangle and extract emotion-specific features based on the information bottleneck theory, thereby improving the accuracy of emotion representation; Third, through extensive experiments, the effectiveness of this framework in four Emotional Domain Adaptation (EDA) settings is verified, with an average performance improvement of 2.49% compared to the existing state-of-the-art methods.
[0076] Example 1
[0077] An emotion domain adaptation method based on disentanglement of the information bottleneck theory, the method is as follows:
[0078] Step 1: Preparation:
[0079] (1) Emotional domain adaptation
[0080] This article focuses on EDA, that is, adapting from the labeled source domain to the unlabeled target domain;
[0081] The source domain is represented as where is the source image, y k is the corresponding label from K emotion categories, N s is the total number of source images; K is the number of emotion categories; i represents the i-th sample; k represents the k-th category;
[0082] The target domain is represented as consisting of N T unlabeled emotional images ; for simplicity, the subscript i is omitted hereafter; N T represents the number of emotional images in the target domain (the number of target domain samples);
[0083] The source domain and the target domain share the same label space, but the distribution P s (x) of the source domain and the distribution P t (x) of the target domain are different;
[0084] (2) CLIP large vision-language model
[0085] Our framework is based on CLIP, which is one of the most popular pre-trained vision-language models; the main components of CLIP are the image encoder E I and the text encoder E T , which accept image and text inputs respectively; in the EDA task, CLIP combines the k-th emotion category into a pre-designed text prompt p k which is expressed as "A [domain] photo seems to express some feelings like [emotion]". The text prompt pk and image x i are then encoded as t k = E T (p k ), v i = E I (x i ), sharing the same feature dimension d v ; classification is based on the highest cosine similarity between t k and v i :
[0086]
[0087] Our goal is to maximize the inter-class distance between text embeddings to enhance the discriminative ability of similarity classification;
[0088] (3) Information Bottleneck (IB)
[0089] To separate emotion-specific features from redundant CLIP embeddings, we incorporate the perspective of the Information Bottleneck theory, whose principle is:
[0090] max[I(z; e) - βI(t; z)] (2)
[0091] where I(·; ·) represents mutual information and β represents the Lagrange multiplier; from the perspective of IB, we formulate a constrained optimization problem aiming to maximize the mutual information between emotion-specific features z and the emotion space e while minimizing the emotion-irrelevant information in the CLIP text embedding t. We implement this optimization through the EoT and SoT modules; however, during the compression process, high-level semantic information may be lost, potentially affecting the emotion recognition ability of the model. We further introduce the IoT module to help recover key information.
[0092] Figure 1 summarizes the framework EmoD of the present invention. The main objective of the present invention is to address the inconsistency problem between CLIP and the emotion space. The present invention proposes an emotion disentanglement framework (including EoT, SoT, and IoT) to eliminate redundant information and extract emotion-specific features from the text embeddings encoded by CLIP. Subsequently, the present invention constructs an emotion projection matrix using the extracted features to project the text and image embeddings of CLIP into a shared subspace. Guided by the emotion-aligned text embeddings, the present invention uses label propagation to generate pseudo-labels for the image embeddings. In addition, the present invention calculates the Maximum Mean Discrepancy (MMD) between the source domain and the target domain to achieve cross-domain feature alignment.
[0093] Step 2: Emotion-Specific Feature Disentanglement
[0094] Although the rich semantic representation of CLIP helps to connect low-level pixel information and high-level emotion semantics, due to the training on millions of captions and concepts, its embeddings contain redundant information, which may affect the performance of similarity classification. To address this issue, our goal is to disentangle the CLIP embeddings into emotion-specific features.
[0095] (1) Emotion Extraction from Text (EoT)
[0096] The present invention aims to use an embedding encoder E m to extract emotion-specific features z = E m (t) ∈ R k×d from the text embeddings t encoded by CLIP, where E m (t) represents the output after the text embeddings pass through the encoder; k and d represent the number of emotion categories and the feature dimension, respectively; to ensure that the extracted features can effectively capture emotional information, the mutual information I(z; e) between the emotion-specific features z and the emotion space e needs to be maximized, which is defined as:
[0097] I(z; e) = E z,e [log p(e|z)] - E e [log p(e)] (3)
[0098] where E z,e denotes the expectation calculated for the two variables z and e; p(e|z) represents the posterior distribution of e with respect to z; E e denotes the expectation calculated for e; p(e) represents the prior distribution of e; the second term E e [log p(e)] remains constant during the optimization process; the first term represents the expected log-likelihood of emotion prediction, which naturally corresponds to the negative value of the cross-entropy loss and captures the predictive ability of the extracted features;
[0099] Based on this insight, we formulate the cross-entropy loss between the emotion-specific features z and the corresponding labels y k from K emotion categories as:
[0100]
[0101] where f() represents the emotion label classifier, represents the k-th element in the softmax output of the K-dimensional vector a, a k represents the k-th dimension of the K-dimensional vector a, a i represents the i-th dimension of the K-dimensional vector a, y k is `1' for the correct class and `0' for the remaining classes; the optimization process of Emotion Extraction from Text (EoT) maximizes the mutual information between the extracted emotion-specific features and the emotion space.
[0102] (2) Structure Optimization (SoT)
[0103] From the perspective of the information bottleneck, it is necessary to minimize the mutual information between the emotion-specific feature z and the input embedding t to reduce redundant information; due to the computational complexity of directly calculating the mutual information, inspired by β-VAE, which uses the KL divergence of the standard normal distribution to limit the bottleneck capacity, an L 2 regularization term is implemented on the emotion-specific feature z:
[0104] L bottle = E[|z| 2 (5)
[0105] where E represents taking the expectation, and L bottle represents the loss function under the bottleneck constraint;
[0106] To further enhance the discriminative ability of the disentangled classification embeddings, an orthogonal regularization is proposed to enforce the independence between different classification dimensions:
[0107]
[0108] where F represents taking the F-norm; L ortho represents the loss function of the orthogonal regularization; p = z T z ∈ R d×d represents the emotion-specific projection matrix, and I ∈ R d×d is the identity matrix; this regularization enforces the orthogonality constraint between the row vectors of z, ensuring the linear independence between the classification dimensions; the overall optimization objective of the SoT process combines the bottleneck constraint and the orthogonal regularization, and the formula is as follows:
[0109] L SoT = L bottle + L ortho (7)
[0110] (3) Information Reconstruction (IoT)
[0111] Considering that excessive feature compression may lead to the loss of semantic information, we introduce a reconstruction loss to retain the information richness of the emotion-specific feature z by enabling the model to recover the original CLIP text embedding;
[0112] Specifically, we model the reconstruction process through the posterior probability p(t|z), which represents the likelihood of reconstructing the original embedding t through the decoder D m given the emotion-specific feature z; the reconstruction quality is quantified by minimizing the mean squared error (MSE) between the original and the reconstructed embeddings:
[0113] L IoT = E[|t - D m(z)| 2 (8)
[0114] Among them, D m (z) represents the CLIP embedding reconstructed from the specific emotion feature z; this reconstruction objective ensures that z captures the features of a specific emotion while retaining the basic semantic properties of the original embedding; through this balanced approach, EmoD achieves strong emotion disentanglement while maintaining the rich semantic context inherent in the CLIP embedding;
[0115] Step 3: Pseudo-label generation
[0116] After obtaining the emotion-specific representation, construct a projection matrix p = z T z projects all target image embeddings and the text embedding t ∈ R K×d into a shared emotion-specific subspace:
[0117] v proj = v t · p ∈ R NT×d , t proj = t · p ∈ R K×d (9)
[0118] where, N T is the number of emotion images in the target domain, v t is the feature embedding of all target images, is the feature embedding of each image, and d is the feature dimension representing the emotion category;
[0119] Given that the text embedding is explicitly aligned with the emotion space through EoT optimization and its category is predefined, use the emotion-aligned t proj as the semi-supervised propagation center to generate pseudo-labels for unlabeled target images, as shown in the right panel. For text and image alignment, aggregate the propagation center and all projected target image embeddings: Figure 1
[0120]
[0121] Then, calculate a sparse similarity matrix and generate pseudo-labels based on the sample distances:
[0122]
[0123] where, A = sim(x · x T ) calculates the cosine similarity between all samples in the subspace; A n extracts the top n values from each row of the matrix A, I represents the indicator function; i and j represent the i-th row and j-th column of the matrix, and the pseudo-label matrix is initialized as:
[0124]
[0125] After normalizing the similarity matrix A n conjugate gradient (CG) is applied for label propagation to obtain pseudo-labels for unlabeled target images, where α is a smoothing parameter for controlling the propagation speed:
[0126]
[0127] Subsequently, the text encoder is optimized by minimizing the pseudo-label-guided target domain loss function and the true supervised source domain loss function :
[0128] L CE = L pl + L gt (14)
[0129] where L CE represents the loss function obtained from Equation (14);
[0130] Step 4: Align the target domain and the source domain
[0131] To achieve effective sentiment domain adaptation and further bridge the domain gap, we introduce cross-domain alignment based on the maximum mean discrepancy (MMD):
[0132] L align = MMD 2 (F s , F t ) (15)
[0133] where F t and F s represent the CLIP-encoded image distributions from the source domain and the target domain, respectively; a Gaussian RBF kernel with bandwidth parameter β is used to compute an unbiased MMD estimate. During the adaptation process, we jointly optimize the text encoder using labeled source data and pseudo-labeled target data to learn domain-invariant sentiment representations. This mechanism ensures the consistency of cross-domain sentiment representations.
[0134] The method of the present invention is based on CLIP and fine-tunes the layer normalization weights of the text encoder while keeping the image encoder frozen. For fair comparison, the present invention uses the same backbone (ViT-B / 16) as other CLIP-based UDA methods. The encoder E m , the decoder D m and the classifier f are implemented as two-layer MLPs with a hidden dimension of 512. All models are optimized using source domain and target domain data.
[0135] The hyperparameters are set as follows: α = 0.99, β = 1, and n = 1. The loss weights are set to λ = 1.2 and γ = 0.8. All images are resized and center-cropped to 224×224. The model is trained for 30 epochs on an NVIDIA A100 GPU with a batch size of 64. We use the text encoder E T and the learning rates of MLP(E m , D m , f) are both set to 1×10 -3 , and the weight decay is 1×10 -4 .
Claims
1. A method for emotional domain adaptation based on information bottleneck theory decoupling, characterized by: The method is: Step 1: Untangling Emotion-Specific Features (1) Emotion Extraction (EoT) Using the embedding encoder E m Extracting emotion-specific features z=E from CLIP-encoded text embedding t m (t)∈R k×d , where E m (t) represents the output of the encoder after the text embedding; k and d represent the number of emotion categories and feature dimensions, respectively; in order to ensure that the extracted features can effectively capture the emotional information, it is necessary to maximize the mutual information I(z; e) between the emotion-specific feature z and the emotion space e, which is defined as: I(z;e)=E z,e [logp(e|z)]-E e [logp(e)] (3) Among them, E z,e represents the expected value of two variables z and e; p(e|z) represents the posterior distribution of e relative to z; E e represents the expected value of e; p(e) represents the prior distribution of e; the second term E e [logp(e)] remains unchanged during the optimization process; the first term represents the expected log-likelihood of sentiment prediction, which naturally corresponds to the negative value of the cross-entropy loss, capturing the predictive power of the extracted features; Based on this insight, emotion-specific features z and corresponding labels y from K emotion categories k The cross entropy loss between is: Among them, f() represents the emotion label classifier, represents the kth element in the softmax output of the K-dimensional vector a, a k represents the kth dimension of a K-dimensional vector a, a i represents the i-th dimension of the K-dimensional vector a, y k is `1' for the correct category and `0' for the rest of the categories; (2) Structure optimization (SoT) An L2 regularization term is implemented on the sentiment-specific feature z using the KL divergence of the standard normal distribution to limit the bottleneck capacity: L bottle =E[|z| 2 ] (5) Among them, E represents the expectation, L bottle represents the loss function under the bottleneck constraint; To further enhance the discriminative power of the disentangled classification embeddings, an orthogonal regularization is proposed to enforce the independence between different classification dimensions: Among them, F means to find the F norm; L ortho represents the loss function of orthogonal regularization; p = z T z∈R d×d represents the emotion-specific projection matrix, I∈R d×d is the identity matrix; this regularization enforces the orthogonality constraint between the row vectors of z and ensures the linear independence between the classification dimensions; the overall optimization goal of the SoT process combines the bottleneck constraint and orthogonality regularization, as follows: L SoT =L bottle +L ortho (7) (3) Information reconstruction (IoT) The reconstruction process is modeled by the posterior probability p(t|z), which represents the probability of passing through the decoder D given the emotion-specific feature z. m The likelihood of reconstructing the original embedding t; the reconstruction quality is quantified by minimizing the mean squared error (MSE) between the original and reconstructed embeddings: L IoT =E[|t-D m (z)| 2 ] (8) Among them, D m (z) represents the CLIP embedding reconstructed from the emotion-specific feature z; this reconstruction objective ensures that z captures the characteristics of the specific emotion while preserving the essential semantic properties of the original embedding; Step 2: Pseudo label generation After obtaining the emotion-specific representation, we construct a projection matrix p = z T z embeds all target images and text embedding t∈R K×d Mapped into a shared emotion-specific subspace: Among them, N T is the number of emotional images in the target domain, v t is the feature embedding of all target images, v i t is the feature embedding of each image, and d is the feature dimension representing the emotion category; Given that text embeddings are explicitly aligned to the sentiment space via EoT optimization and their categories are predefined, using sentiment-aligned t proj As a semi-supervised propagation center, generate pseudo labels for unlabeled target images. For text and image alignment, aggregate the propagation center and all projected target image embeddings: Then, a sparse similarity matrix is calculated to generate pseudo labels based on the sample distances: Where A = sim(x·x T ) Calculate the cosine similarity between all samples in the subspace; A n Extract the first n values from each row of matrix A, I represents the indicator function; i and j represent the i-th row and j-th column of the matrix, the pseudo label matrix Initialized to: In the similarity matrix A n After normalization, conjugate gradient (CG) is applied for label propagation to obtain the pseudo labels of the unlabeled target image, where α is a smoothing parameter used to control the propagation speed: Then, the target domain loss function is guided by minimizing the pseudo-labels And the real supervised source domain loss function To optimize the text encoder: L CE =L pl +L gt (14) Among them, L CE It represents the loss function obtained by formula 14; Step 3: Align the target domain and the source domain A cross-domain alignment based on maximum mean difference (MMD) is introduced: L align =MMD 2 (F s ,F t ) (15) Among them, F t and F s denote the distribution of CLIP encoded images from the source and target domains, respectively; a Gaussian RBF kernel with bandwidth parameter β is used to compute the unbiased MMD estimate.
2. The method for emotional domain adaptation based on information bottleneck theory decoupling according to claim 1, characterized in that: The preparation work of the method includes the following three points: (1) Adaptation in the emotional domain The source domain is represented as in, is the source image, y k are the corresponding labels from K sentiment categories, N s is the total number of source images; K is the number of emotion categories; i represents the i-th sample; k represents the k-th category; The target domain is represented as By unmarked N T Emotional images Composition; for simplicity, the subscript i is omitted hereafter; N T Represents the number of emotional images in the target domain (number of target domain samples); The source and target domains share the same label space, but the distribution of the source domain P s (x) and the distribution P of the target domain t (x) different; (2) CLIP Large-Scale Visual Language Model The main component of CLIP is the image encoder E I and text encoder E T , accepting image and text inputs respectively; in the EDA task, CLIP incorporates the kth emotion category into the pre-designed text prompt p k In the text prompt p k and image x i It is then encoded as t k =E T (p k ), v i =E I (x i ), share the same feature dimension d v ; Classification based on t k and v i The highest cosine similarity between: Our goal is to maximize the text embedding The inter-class distance between them is used to enhance the discriminative ability of similarity classification; (3) Information bottleneck (IB) In order to separate emotion-specific features from redundant CLIP embeddings, the information bottleneck theory is combined, the principle is: max[I(z;e)-βI(t;z)] (2) Where I(·; ·) represents mutual information, and β represents the Lagrange multiplier.