Double emotion classification method based on diffusion anti-fact data enhancement

Through the dual emotion classification method enhanced by diffusion counterfactual data, antisense samples with diversity and semantic consistency are generated. Combined with reinforcement learning and contrast learning techniques, the pseudo-association problem of extraterritorial data in existing sentiment analysis is solved, and the accuracy and robustness of emotion classification are improved.

CN120386867AActive Publication Date: 2025-07-29NORTHEAST FORESTRY UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311646573.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-07-29
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

The existing emotion analysis methods fail to effectively solve the pseudo-association problem of extradomain data when training data in the domain, resulting in poor emotional classification effects. The traditional synonym data enhancement and word embedding against perturbation methods fail to effectively improve the generalization ability of the model.

Method used

A dual emotion classification method based on diffusion counterfactual data augmentation is adopted, and a network model and a dual discriminator are generated by constructing a counterfactual diffusion to generate antisense samples with diversity and semantic consistency, and optimized using reinforcement learning and contrast learning techniques, combining dual emotion classifiers for emotion classification.

Benefits of technology

It improves the generalization ability of the model on extradomain data, solves the pseudo-association problem, and improves the accuracy and robustness of emotional classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386867A_ABST
    Figure CN120386867A_ABST
Patent Text Reader

Abstract

The invention relates to the field of natural language processing, in particular to a double sentiment classification method based on diffusion anti-fact data enhancement. The invention provides a double sentiment classification method based on diffusion anti-fact data enhancement, and aims at solving the problem that the sentiment classification effect is poor due to the fact that the problem of false association in data is not solved when the performance of a model is improved through synonymous sentence data enhancement or word embedding disturbance resistance in the prior art. Establishing and obtaining a trained double sentiment classification model DCA based on diffusion anti-fact data enhancement; and inputting the emotion data to be classified into the well-trained double emotion classification model DCA based on diffusion anti-fact data enhancement to complete double emotion classification. And two predictors are adopted as emotion classifiers to carry out double emotion classification. And the classifier with higher confidence is selected to output the result, so that the problem of poor sentiment classification effect caused by the false correlation problem in the data in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular, to a dual sentiment classification method based on diffusion counterfactual data augmentation. Background Art

[0002] With the wide popularity on channels such as social media websites, news portals, and forums, people can freely and conveniently express their opinions. These views and opinions deeply and truly reflect people's lives related to the interests of the people, social production, and subjective social and political attitudes related to the emotional will of the people. It is of great significance for research fields such as public opinion surveys, market trend analysis, recommendation systems, personalized recommendations for users, public health monitoring, information retrieval, and rumor detection. At the same time, automatically obtaining semantic information from natural language texts has always been an important research issue and has many practical application fields, which can solve problems that meet actual needs such as sentiment analysis, sarcasm, controversy, irony, rumors, fake news detection, stance detection, and argument mining. In recent years, in the field of sentiment analysis research, the use of pre-trained language models (such as BERT, GPT, XLNet, RoBERTa, ALBERT, T5, and ERNIE) has significantly improved the performance of natural language processing tasks. The inherent complexity of human emotions poses continuous challenges in actual sentiment analysis, leading to problems such as overfitting. When encountering minor modifications in real-world examples, this challenge can lead to model prediction failures. Researchers have tried to address these challenges by using data augmentation and adversarial perturbations, such as enhancing training data by generating synonymous sentences or introducing random noise during the word embedding stage to enhance the robustness of the model. However, models trained on independent and identically distributed (IID) data show a decreasing trend in generalization ability on out-of-domain (OOD) data. Then, in order to meet the requirements of complex application scenarios in the real world, how to improve the sentiment analysis task of model generalization ability has become an insurmountable major research topic in front of us.

[0003] Previous sentiment analysis research has mainly been divided into three categories to enhance the robustness of neural networks: adversarial training, causal reasoning, and data augmentation. The adversarial training method introduces adversarial noise during the training process, adds regularization terms, or uses dropout strategies to help the model produce consistent outputs when facing data perturbations. Causal reasoning methods aim to improve the robustness of the model by identifying the factors that truly affect sentiment in the text. Some studies have emphasized establishing causal relationships between word features and sentiment labels. The main idea of data augmentation is to improve the generalization ability of the model by generating more training data. Some studies have proposed synonym-based augmentation, which randomly replaces corresponding words in the original sample with synonyms, hypernyms, or hyponyms. However, the synonym augmentation method cannot effectively solve the pseudo-association problem existing in the original sample. For this reason, some studies have made minimal modifications to the training data and used manual annotation methods to perform label inversion. However, this manual method is costly and time-consuming. In addition, some studies have adopted automated techniques, such as formulating antonym replacement rules to automatically generate antonym samples under the rules, which alleviates the pseudo-association problem existing in the original data to a certain extent. However, the current counterfactual methods have several limitations: (i) Most studies rely on fixed rules or known negations, synonyms, and antonyms in WordNet to create antonym samples, which limits the diversity and semantic consistency of the generated samples. (ii) During the training process, the generated anonymized samples are merged with the original samples without considering their corresponding relationships. (iii) Data generation and classification are usually treated as different tasks and trained sequentially, resulting in problems such as error accumulation.

[0004] Therefore, in the prior art, the current sentiment analysis methods cannot handle out-of-domain (OOD) data well due to the model being trained on in-domain (IID) data and often improve the performance of the model through synonym sentence data augmentation or by performing adversarial perturbations on word embeddings, without solving the pseudo-association problem within the data, resulting in poor sentiment classification effects. Summary of the Invention

[0005] The object of the present invention is to solve the problem in the prior art that the performance of the model is improved through synonym sentence data augmentation or by performing adversarial perturbations on word embeddings, without solving the pseudo-association problem within the data, resulting in poor sentiment classification effects. We propose a dual sentiment classification method based on diffusion counterfactual data augmentation.

[0006] I. Establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain the trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation;

[0007] Second, input the emotional data to be classified into the trained dual emotional classification model DCA based on diffusion counterfactual data augmentation to complete the dual emotional classification.

[0008] In the first step above, establish a dual emotional classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual emotional classification model DCA based on diffusion counterfactual data augmentation; the specific process is as follows:

[0009] Step 1: Obtain a training set;

[0010] The training set includes emotional sample data w o , and the emotional sample data w o in the training set are all labeled data, and the labeled data is emotion s;

[0011] Step 2: Establish a dual emotional classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual emotional classification model DCA based on diffusion counterfactual data augmentation.

[0012] The specific process of establishing a dual emotional classification model DCA based on diffusion counterfactual data augmentation and obtaining a trained dual emotional classification model DCA in Step 2 above is as follows:

[0013] The dual emotional classification model DCA based on diffusion counterfactual data augmentation includes: a text generation layer, a dual discriminator layer, a contrast learning layer, and a dual emotional classification layer;

[0014] S1: Input the original sample w o in the training set into the text generation layer for processing to construct an antonym-guided paradigm w g ;

[0015] S2: Construct a counterfactual diffusion generation network model and obtain a trained counterfactual diffusion generation network model;

[0016] Input the antonym-guided paradigm w g and the original sample w o into the trained counterfactual diffusion generation network model, and output the antonym sample w a ;

[0017] S3: Construct an original edge discriminator C o and obtain a trained original edge discriminator C o ;

[0018] Construct an antonym edge discriminator C a ; obtain a trained antonym edge discriminator C a ;

[0019] The dual discriminator layer uses the trained original edge discriminator C oand the trained antonym edge discriminator C a Optimize the trained counterfactual diffusion generation network model to obtain an optimized counterfactual diffusion generation network model; input the antonym guidance paradigm w g and the original sample w o into the optimized counterfactual diffusion generation network model, and output the updated antonym sample w a1 ;

[0020] S4: Finally optimize the optimized counterfactual diffusion generation network model obtained in S3 to obtain a qualified counterfactual diffusion generation network model. Input the antonym guidance paradigm w g and the original sample w o into the qualified counterfactual diffusion generation network model, and output the qualified antonym sample w a2 ;

[0021] S5: Use the trained original edge discriminator C o and the antonym edge discriminator obtained in S3 as the original edge classifier and the antonym edge classifier of the sentiment classification layer; then transmit the original sample w o and the qualified antonym sample w a2 generated in S4 to the sentiment classification layer in the form of sample pairs for dual sentiment classification.

[0022] The beneficial effects of the present invention are as follows:

[0023] The present invention proposes a dual sentiment classification neural network model DCA based on diffusion counterfactual data augmentation. This framework combines a counterfactual diffusion generation network model, a discriminator, and a dual sentiment classification model, and integrates the above parts into a framework using reinforcement learning and contrast learning techniques. The diffusion model has been applied to the field of generating controlled text. Compared with traditional methods such as GANs, the diffusion model often shows a more stable training process and can generate different contents without retraining. Therefore, content control can be performed and the diversity of content is particularly valuable.

[0024] In the generation stage, first use multi-label learning to design the optimal antonym paradigm for the original sample. Next, use contrast learning between the antonym paradigm and the original sample to guide the sampling process of the diffusion model. By integrating the rewards from the discriminator, diverse antonym samples with controllable emotion labels and coherent semantics are generated. In the emotion classification stage, a dual emotion classifier is established, which uses a trained original sample predictor and an antonym sample predictor from the discriminator, and uses the two predictors as sentiment classifiers for dual sentiment classification. By selecting the output result of the classifier with higher confidence, the problem of pseudo-correlation inside the data in the prior art, which leads to poor sentiment classification effect, is solved. Description of the Drawings

[0025] Figure 1 It is a schematic structural diagram of the DCA model of the present invention. Detailed implementation manners

[0027] Detailed implementation manner 1: In combination with Figure 1 describe the present invention, including

[0028] I. Establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation;

[0029] II. Input the sentiment data to be classified into the trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation to complete the dual sentiment classification.

[0030] Detailed implementation manner 2: The difference between this implementation manner and the first detailed implementation manner is that

[0031] in the first one, establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation; the specific process is as follows:

[0032] Step 1. Obtain a training set;

[0033] The training set includes sentiment sample data w o , and the sentiment sample data w o in the training set are all labeled data, and the label data is sentiment s; the types of labels carried by the sample data are different;

[0034] For example: positive sentiment sample w o : The plot of this movie is full of ups and downs and fascinating. I think it is very interesting. Among them, the sample w o is used as the input of the DCA model, and the positive sentiment s is used as the output of the DCA model;

[0035] To train the dual sentiment classification model DCA based on diffusion counterfactual data augmentation, a validation set and a test set are also used.

[0036] The validation set and the test set also include sentiment sample data w o , and the labels carried by the sample data in the test set are invisible to the DCA model;

[0037] The validation set is used for our final classification process. Before making predictions on the test set, we will use the validation set to globally tune our model parameters using the Adam optimizer to ensure that our sentiment classification model achieves the best classification performance. After parameter tuning, we will input the test set text into the DCA model for text generation and sentiment classification. The labels of the test set are invisible. We will compare the labels predicted by the model with the true labels of the test set text and find that the sentiment classification accuracy of the model has been greatly improved compared to the original sentiment classification method.

[0038] Step 2: Establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation.

[0039] Other steps and parameters are the same as those in the specific implementation method 1.

[0040] The specific process of establishing a dual sentiment classification model DCA based on diffusion counterfactual data augmentation and obtaining a trained dual sentiment classification model DCA in Step 2 is as follows:

[0041] The dual sentiment classification model DCA based on diffusion counterfactual data augmentation includes: a text generation layer, a dual discriminator layer, a contrastive learning layer, and a dual sentiment classification layer;

[0042] S1: Input the original sample w in the training set o into the text generation layer for processing to construct an antonym-guided paradigm w g ;

[0043] S2: Construct a counterfactual diffusion generation network model to obtain a trained counterfactual diffusion generation network model;

[0044] Input the antonym-guided paradigm w g and the original sample w o into the trained counterfactual diffusion generation network model to output an antonym sample w a ;

[0045] S3: Construct an original edge discriminator C o to obtain a trained original edge discriminator C o ;

[0046] Construct an antonym edge discriminator C a ; obtain a trained antonym edge discriminator C a ;

[0047] The dual discriminator layer uses the trained original edge discriminator C o and the trained antonym edge discriminator C aOptimize the trained counterfactual diffusion generation network model to obtain an optimized counterfactual diffusion generation network model; input the antonym guidance paradigm w g and the original sample w o into the optimized counterfactual diffusion generation network model, and output the updated antonym sample w a1 ;

[0048] S4: Perform final optimization on the optimized counterfactual diffusion generation network model obtained in S3 to obtain a qualified counterfactual diffusion generation network model. Input the antonym guidance paradigm w g and the original sample w o into the qualified counterfactual diffusion generation network model, and output the qualified antonym sample w a2 ;

[0049] S5: Use the trained original edge discriminator C o obtained in S3 and the antonym edge discriminator as the original edge classifier and antonym edge classifier of the sentiment classification layer; then transmit the original sample w o and the qualified antonym sample w a2 generated in S4 to the sentiment classification layer in the form of a sample pair for dual sentiment classification.

[0050] The dual sentiment classification is reflected in that if the maximum value of the two sentiment degree probabilities of the original edge is higher than the maximum value of the two sentiment degrees of the antonym edge classifier, then the classified sentiment of the original edge classifier shall prevail. Conversely, since we are predicting the sentiment of the original sample w o , then we shall take the opposite sentiment of the sentiment predicted by the antonym edge classifier as the standard. The opposite sentiment of the antonym sample w a is the sentiment of the original sample w o .

[0051] Other steps and parameters are the same as those in one of the specific implementation manners 1 to 2.

[0052] Specific implementation manner 4: The difference between this implementation manner and the specific implementation manners 1 to 4 is that

[0053] in S1, the specific process of inputting the original sentiment sample sequence w o in the training set into the text generation layer to generate the antonym guidance paradigm w g is as follows:

[0054] S1.1: Set the stop words in the original sample w o as non-replaceable words. The original sample w o is composed of the original sample sequence w o ={w1, w2,..., w n}, where the original sample sequence consists of n words; n is a positive integer; the stop words are words unrelated to emotions to prevent the intrusion of irrelevant information. The stop words are a list of words in the Python NLTK library, including words such as "the", "and", "is", etc. They generally contribute less to the meaning of the text and may introduce noise in text processing tasks

[0055] S1.2: For the adjectives, adverbs, and verbs in the original sample w o that have antonyms, replace them with the corresponding antonyms in the English vocabulary database WordNet; the adjectives, adverbs, and verbs in the original sample w o can affect the emotion of the sentence formed by the original sample sequence;

[0056] S1.3: For the adjectives, adverbs, and verbs in the original sample w o that have no antonyms, as well as non-stop words, replace them with the corresponding synonyms in the English vocabulary database WordNet;

[0057] S1.4: Generate an antonym-guided paradigm and an opposite sentiment label based on the original sample after replacement according to S1.1 to S1.3 and the sentiment label of the original sample (w o , s) where s represents the sentiment label of the original sample w o , represents the opposite sentiment label.

[0058] The other steps and parameters are the same as those in one of the specific embodiments one to three.

[0059] Specific embodiment five: The difference between this embodiment and the specific embodiments one to four is that

[0060] For each word w o = {w1, w2,..., w n} in the original sample sequence, calculate the probability of replacing the word w k with the j-th word in the English vocabulary database WordNet using supervised multi-label classification; the specific process is as follows: k Specifically:

[0061]

[0062] Among them, represents the probability that the j-th word in the English vocabulary database WordNet belongs to the replacement word set of the k-th word w o in the original sample w k , represents the k-th word w o in the original sample w kThe j-th word in the corresponding English vocabulary database WordNet, exp() represents the exponential function, h k is the hidden representation of the word w k , W j is the weight matrix, b j is the bias, and m is the size of the English vocabulary database WordNet;

[0063] Supervised multi-label classification is an algorithm applied to text classification. We design a binary classifier for each word in WordNet. m is the number of dimensions of the words in WordNet, that is, we have m binary classifiers, and its role is to determine whether the word belongs to the replacement word set. Specifically, we obtain the hidden representation hk of each word in w o and perform classification with each of the m words in WordNet to determine whether the j-th word in WordNet belongs to the replacement word. This method is called the supervised multi-label classification algorithm;

[0064] Then, set the replacement probabilities of the words in the original sample sequence w o ={w1, w2,..., w n} that are lower than 0.57 and the words not included in the English vocabulary database WordNet to zero, and obtain the normalized probability distribution w k ~Multinomial(P k ) for determining whether to replace the word w k with the j-th word in the English vocabulary database WordNet;

[0065] The normalized probability distribution w k ~Multinomial(P k ) for determining whether to replace the word w k with the j-th word in the English vocabulary database WordNet; The specific process is as follows:

[0066]

[0067] Among them, P k represents the multinomial word distribution of the replacement probability of the k-th word w o in the original sample w k in the English vocabulary database WordNet. normalize() represents the normalization of the probability belonging to the replacement word set in the English vocabulary database WordNet, represents the probability of the first word in the English vocabulary database WordNet belonging to the replacement word set of the k-th word w o in the original sample w k , and Multinomial represents the original sample w ofrom w1 to w in n The polynomial formed by sampling

[0068] Determine that the probability of replacement for a word with too low a probability is the likelihood of determining whether the j-th word is a replacement word during the multi-label classification process. If this probability is too low, we set the probability to 0.

[0069] Words not included in the English vocabulary database WordNet have their replacement probabilities set to 0 because the replacement probabilities cannot be calculated.

[0070] The other steps and parameters are the same as those in one of the specific implementation manners one to four.

[0071] Specific implementation manner six: The difference between this implementation manner and the specific implementation manners one to five is that

[0072] In S2, the specific process of constructing the counterfactual diffusion generation network model and obtaining the pre-trained counterfactual diffusion generation network model is as follows:

[0073] Construct the counterfactual diffusion generation network model, and use the original sample w o and the antonym-guided paradigm w constructed in S1 g Combined in the form of a sample pair as the training data set of the counterfactual diffusion generation network model; input the antonym-guided paradigm w g and the original sample w o into the counterfactual diffusion generation network model for joint training to obtain the pre-trained counterfactual diffusion generation network model;

[0074] In S2, the specific process of inputting the antonym-guided paradigm w g constructed in S1 and the original sample w o into the diffusion model for joint training to obtain the pre-trained counterfactual diffusion generation network model is as follows:

[0075] S2.1: Concatenate the original sample w o and the antonym-guided paradigm constructed in S1, the antonym-guided paradigm w g and denote it as the discrete sample Perform a forward conditional noise addition process on the obtained discrete sample embedded in the continuous feature space to obtain the word vector x0. The specific process is as follows:

[0076]

[0077] Among them, represents performing a forward process of continuous diffusion on the discrete sample , N() represents a multi-dimensional normal distribution, Emb() represents an embedding function, β0 is the amount of noise added in the time step, and I is the identity matrix;

[0078] Diffusion models can only handle continuous variables. However, the nature of text is discrete as it consists of individual words. Therefore, we emphasize that we embed the discrete samples into a continuous feature space through an embedding function so that the diffusion model can process them. The discrete samples here refer to wo+g after concatenation;

[0079] S2.2: Perform a forward conditional noise addition process on the word vector x0 obtained in S2.1 for T steps, where T is a positive integer; after performing the forward conditional noise addition process for T steps, a series of latent Gaussian noise variables x1,...,x T , x1,...,x T Each Gaussian noise variable in represents an intermediate step output in the forward conditional noise addition process. After performing the T-th forward conditional noise addition process, the complete Gaussian noise x T is obtained. The specific process is as follows:

[0080] When performing the forward conditional noise addition process at each time step t∈(1,2...T), add Gaussian noise with β t-1 to the Gaussian noise x t obtained in the previous step to obtain the Gaussian noise x t at the t-th time step; after performing the T-th forward conditional noise addition process, the complete Gaussian noise x T is obtained. The formula is as follows:

[0081]

[0082] where q(x t |x t-1 ) represents a single continuous diffusion forward process for the Gaussian noise x t-1 . The purpose of dividing it into T steps is that adding a small amount of noise at each step allows the network to learn the noise added at each time step to guide the reverse denoising process. This process learns the noise added at each time step through the network, and the learned parameters are used to guide the reverse denoising process

[0083] S2.3: Perform a reverse conditional denoising process on the complete Gaussian noise x T obtained in S2.2 for T steps to obtain the denoised word vector x0. The specific process is as follows:

[0084] p θ (x t-1 |x t ,t) = N(x t-1 ; μ θ (x t ,t), Σ θ (x t ,t))

[0085] Among them, p θ (x t-1 |x t , t) represents performing a reverse conditional denoising process on the Gaussian noise x t , and μ θ (·) and Σ θ (·) are the Gaussian distribution parameters learned during the reverse conditional denoising process;

[0086] We complete the conditional modeling by replacing t-1 in x with w o .

[0087] In each diffusion sampling step, we apply a rounding operation to the reparameterized x t to project it back into the word embedding space. It can be inferred that the core of our generation process is to learn the data distribution between the original samples and the antonymous paradigms, rather than controlling them through a large number of classifiers

[0088] Different from the above traditional diffusion model, in our reverse denoising process of diffusion, the part of x t-1 that belongs to is replaced with w o (it can be understood that because is concatenated and embedded into the continuous space, then x t-1 is actually composed of the embeddings of and ). Then, if we keep performing the above replacement during reverse denoising, it is equivalent to that reverse denoising only acts on So we call x t a partially Gaussian-noised variable.

[0089] Combined with Figure 1 looking at the lower left part, all the noise points are drawn in the upper half. Therefore, in the present invention, x0 is obtained by reverse denoising from the partially Gaussian-noised xt, and this x0 still consists of two parts. Among them, since the x0 part is always replaced, the x0 part remains unchanged, and finally what is generated is the w g part. The word vector x0 after reverse denoising is not strictly equal to the initial x0, but only close to the original x0.

[0090] S2.4: Calculate the variational lower bound according to the forward noise addition process in S2.2 and the reverse denoising process in S2.3: Use the variational lower bound as the loss function to train the counterfactual diffusion generation network model to obtain a pre-trained counterfactual diffusion generation network model; the specific process is as follows:

[0091]

[0092] Among them, Eq (·) indicates expectation, D KL (q(x T |x0)||p θ (x T )) represents the forward process of diffusion, represents the reverse process of diffusion, Indicates that discrete samples are embedded in continuous space, For rounding operation, D KL Represents the KL divergence calculation formula between Gaussian distributions, x T represents Gaussian noise.

[0093] The variational lower bound as a loss function refers to the mean squared error between the true noise and the model's predicted noise.

[0094] The reason for using the variational lower bound is that the original training loss of the diffusion model is not convenient to optimize. After deriving the original loss, it is deduced and decomposed and converted into a variational lower bound loss function to replace the original loss function.

[0095] ; Other steps and parameters are the same as those in one of the specific implementation methods one to five.

[0096] Specific embodiment 7: This embodiment differs from specific embodiments 1 to 6 in that the original edge discriminator is constructed in S3 to obtain a trained original edge discriminator; the specific process is:

[0097] When the original edge discriminator is an LSTM model, the original sample w o and its corresponding sentiment label s as the training set of the original edge discriminator; each original sample w o The original edge discriminator is trained by using a supervised learning method to obtain a trained original edge discriminator.

[0098] When the original edge discriminator is the pre-trained model Bert-base or Bert-large, the original sample w o and its corresponding sentiment label s as the training set of the original edge discriminator; the original edge discriminator C is trained by fine-tuning o , obtaining a trained original edge discriminator; fine-tuning methods are divided into: fine-tune and prompt-tune. Fine-tune, also known as full-parameter fine-tuning, is a method consistently used in BERT fine-tuning models. All parameter weights are updated to adapt to the domain data, resulting in good results.

[0099] In S3, an antisense edge discriminator is constructed; a trained antisense edge discriminator is obtained; the specific process is:

[0100] When the antonym edge discriminator is an LSTM model, the antonym samples w generated by the pre-trained counterfactual diffusion generation network model obtained in S2 a and their corresponding sentiment labels are used as the training set, and the antonym edge discriminator is trained in a supervised learning manner to obtain a trained antonym edge discriminator;

[0101] Each antonym sample w a is transformed into a word embedding sequence through an embedding function, and the discriminator is enabled to extract sentiment information from the text through supervised learning to finally predict the associated sentiment label s.

[0102] When the antonym edge discriminator is a pre-trained model Bert-base or Bert-large, the antonym samples w generated by the pre-trained counterfactual diffusion generation network model obtained in S2 a and their corresponding sentiment labels s are used as the training set; the antonym edge discriminator C is trained by fine-tuning a to obtain a trained antonym edge discriminator C a ; the fine-tuning methods are divided into: Fine-tune and prompt-tune. Fine-tune, also called full-parameter fine-tuning, is the method always used in the bert fine-tuning model. All parameter weights participate in the update to adapt to the domain data, and the effect is good,

[0103] The original sample w is provided to the model a and its corresponding sentiment label w is converted into an input format suitable for the Bert model through the fine-tuning process. In fine-tuning, we input each original sample w a and its corresponding sentiment label s into the discriminator together to adjust the parameters of the Bert model to better adapt to the sentiment classification task, and the antonym edge discriminator C is trained through the above method a a . Other steps and parameters are the same as one of the specific embodiments 1 to 6.

[0104] Specific embodiment 8: The difference between this embodiment and the specific embodiments 1 to 7 is that

[0105] in S3, the dual discriminant layer uses the trained original edge discriminator and the trained antonym edge discriminator to optimize the trained counterfactual diffusion generation network model, and the specific process of obtaining the optimized counterfactual diffusion generation network model is as follows:

[0106] S3.1: Input the antonym sample w a into the trained original edge discriminator C o , the original edge discriminator C o ​It successively includes multiple hidden layers and an original edge discriminator output layer, and each layer consists of some weights and activation functions; the output of each hidden layer serves as the input of the next hidden layer, and the output h of the last hidden layer o serves as the input of the original edge discriminator output layer, where h o represents the antonym sample w a is the hidden representation within the original edge discriminator C o inside.

[0107] Input the antonym sample w a into the antonym edge discriminator C a , and the antonym edge discriminator C a successively includes multiple hidden layers inside, and each layer consists of weights and activation functions; the output of each hidden layer serves as the input of the next hidden layer, and the output h of the last hidden layer a serves as the input of the antonym edge discriminator output layer, where h a represents the antonym sample w a is the hidden representation within the original edge discriminator C a inside;

[0108] Perform dual sentiment prediction on the obtained h o and h a ;

[0109] They obtain hidden representations through the original edge discriminator and the antonym edge discriminator networks, which are respectively denoted as ho and h a for dual sentiment prediction. When we use LSTM as the discriminator network model, it can capture the temporal information in the text, convert the text into a word embedding sequence through the embedding function, and utilize the cyclic structure of LSTM, enabling it to understand the word order relationship in the sentence and capture the complexity of emotional expression. When we use Bert-base or Bert-large as the discriminator network model, it is based on the Transformer structure, which is pre-trained on a large-scale corpus and can deeply understand the context, providing a powerful modeling ability for sentence-level sentiment analysis. By fine-tuning on the basis of the pre-trained Bert-base or Bert-large model.

[0110] The output layers of both the original edge discriminator and the antonym edge discriminator are softmax layers, and the output is as follows:

[0111] p ori (s|w a ) = softmax(W ori h o + b ori )

[0112] p ant (s|wa ) = softmax(W ant h a + b ant )

[0113] Among them, W ori and b ori represent the parameters of C o ; W ant and b ant represent the parameters of C a ; The parameters of C a are dynamically updated as the antonym samples w a are generated. We prevent the update of the parameters of the original predictor C o by blocking the backpropagation of gradients;

[0114] W ori and b ori are the parameters of the original edge discriminator, which are trained by the original sample set. Therefore, after training, the two parameters will no longer be updated, that is, they will no longer change as the antonym samples are generated. Blocking the backpropagation of gradients means that as explained before, the original edge discriminator is only trained by the original sample set Do. Although it participates in the discrimination of the antonym samples w a , it will not perform backward gradient backpropagation under the antonym samples w a . In this way, the backpropagation of gradients is blocked, and the update of the network parameters of the original edge discriminator C o is stopped. This step is reflected in the process of double sentiment discrimination of the antonym samples w a . The parameters of the antonym edge discriminator are trained by the antonym samples w a , so the gradient backpropagation on C a will not be blocked.

[0115] The parameters W ant and b ant of the antonym edge discriminator, because there are no antonym samples at the beginning, and the antonym samples are generated later. The generated antonym samples w a are used to train an antonym edge discriminator, and this training process makes Want and bant keep updating, so it is described as dynamic update here.

[0116] S3.2: The double discrimination layer performs double discrimination according to the outputs of the original edge discriminator C o and the antonym edge discriminator C a , and generates rewards at the same time. The specific process is as follows:

[0117]

[0118] Among them, α and γ are negatively correlated trade-off parameters whose sum is 1. In Ca During the cold start phase, as C a The performance is improved, α is gradually increased from 0 to 1; the cold start stage means that at the beginning our antisense edge discriminator does not generate antisense samples for training, the antisense samples w a It is generated later, and it is generated one by one by the antisense sample w a Gradually trained. As the number of antisense samples increases, the antisense edge discriminator begins to train gradually, and this process is called cold start.

[0119] S3.3: Use a gradient-based strategy (similar to random sampling) to extract M antisense samples for double identification in S3.2, and obtain M rewards. Calculate the average reward of the M rewards as the baseline reward r avg , and based on the baseline reward r avg Perform reinforcement learning, (the extraction step is performed after the counterfactual diffusion generation network is generated, for an original sample w o The countersense sample w is generated by counterfactual diffusion a It is not unique, but multiple ones are generated. We will calculate the average reward of multiple antonyms as the supervisory signal for the generated antonym samples)

[0120] ; The specific process is:

[0121] When r(w a ) exceeds the average reward r avg When the counterfactual diffusion generation network model is added to generate the countersense sample w a The probability that r(w a ) is less than or equal to the average reward r avg When this decreases, the antisense sample w a The reward and baseline reward r avg The difference is the generated antisense sample w a If the actual reward is negative or 0, the counterfactual diffusion generation network will reduce the amount of such countersense samples w generated. a The production of

[0122] The loss function formula for reinforcement learning is as follows:

[0123] L r =-log(r(w a )-r avg )P G (w a |w o )

[0124] where r(w a )-r avg represents the antonym sample w aThe reward and the baseline reward r avg The difference is used as the antonym sample w a The actual calculated reward, P G (w a |w o ) represents the probability that the counterfactual diffusion generation network model generates the antonym sample w given the original sample w o Generate the antonym sample w a of.

[0125] Other steps and parameters are the same as those in any one of the specific implementation manners one to seven.

[0126] Specific implementation manner nine: The difference between this implementation manner and the specific implementation manners one to eight is that,

[0127] In S4, the optimized counterfactual diffusion generation network model obtained in S3 is finally optimized to obtain a qualified counterfactual diffusion generation network model; the specific process is as follows:

[0128] Map the original sample w o , the antonym guiding paradigm w obtained in S1 g and the antonym sample w generated in S2 a to different-dimensional spaces in the contrast learning layer, and the contrast learning layer uses the embedding function f(·) to learn the contrast representation; obtain the contrast loss to finally optimize the optimized counterfactual diffusion generation network model obtained in S3; the specific process is as follows:

[0129] In S4, map the original sample w o , the antonym guiding paradigm w g and the generated antonym sample w a to different-dimensional spaces in the contrast learning layer, and the contrast learning layer uses the embedding function f(·) to learn the contrast representation; the specific process is as follows:

[0130] S4.1: Use the antonym guiding paradigm w obtained in S1 g as the anchor point, use the updated antonym sample w generated by the counterfactual diffusion generation network model optimized in S3 a1 as the positive sample, and the original sample w o as the negative sample for contrast learning. The loss function formula for contrast learning is as follows:

[0131]

[0132] where ε represents the safety margin, f(·) represents the embedding function, represents the square of the L2 norm, and max(·) represents the maximum value function.

[0133] S4.2: Use the L obtained in S4.1 c and the L obtained in S3.3r and L obtained in S2.4 vlb Add them together and use them as the overall loss function to train the optimized counterfactual diffusion generative network model to obtain a qualified counterfactual diffusion generative network model. The overall loss function formula is as follows:

[0134] L=L vlb +L r +L c

[0135] Among them, L vlb Generate network loss function for diffusion, L r is the reinforcement learning loss function, L c is the contrastive learning loss function.

[0136] According to the original samples and their attached emotions (w o ,s) and the antonym guidance paradigm and its attached emotions We will diffuse the generated network loss L vlb , its role is to make the generative network generate outputs close to the input; reinforcement learning loss L r , which is to jointly optimize the counterfactual diffusion generation network model and the discriminator; and the contrastive learning loss L c , its role is to bring the antonym sample w closer a With the antonym-guided paradigm g distance, pull away the original sample w o Antonym-guided paradigm g The sum of the three is fed back to the counterfactual diffusion generation network model as the total loss, which promotes the counterfactual diffusion generation network model to generate qualified antonym samples and their correct attached emotions.

[0137] The other steps and parameters are the same as those in the first to eighth embodiments.

[0138] Specific embodiment 10: This embodiment differs from specific embodiments 1 to 9 in that:

[0139] In S5, the trained original edge discriminator C obtained in S3 is used o and antisense edge identification as the original edge classifier and antisense edge classifier of the sentiment classification layer; then the original sample w o and S4 generate qualified antisense samples w a2 The samples are transmitted to the sentiment classification layer in the form of pairs for dual sentiment classification. The specific process is as follows:

[0140] Calculate the confidence of the original edge classifier and the antisense edge classifier. When the original edge classifier exceeds the antisense edge classifier confidence threshold μ, the original sample w oThe sentiment label adheres to the result of the original classifier. Conversely, the sentiment label of the original sample w o is regarded as a qualified antonym sample w a2 's sentiment; obtain the sentiment label of the original sample w o , and the judgment formula is as follows:

[0141]

[0142] where μ represents the confidence threshold, p(s|w o ) represents the sentiment prediction of the original sample w o , p ori (s|w o ) represents the sentiment prediction of the original edge classifier for the original sample w o , p ant (s|w a2 ) represents the sentiment prediction of the antonym edge classifier for the qualified antonym sample w a2 .

[0143] The specific process is explained as follows: The original sample w o and the antonym sample w a are in a one-to-one correspondence. We transmit w o to the original edge classifier and transmit the corresponding w o of w a to the antonym edge classifier for sentiment classification respectively. The confidence threshold is a value we set ourselves, ranging from [0,1]. First, it is emphasized that the return form of the predictor result is a set of numerical probability pairs. The original edge classifier and the antonym edge classifier respectively predict the sentiment of the original sample w o and the antonym sample w a in the form of (negative probability, positive probability), where the sum of the two probabilities is 1. Taking (a, b) as an example, the sum of the two values a and b is 1, where a represents the probability of the negative degree and b represents the probability of the positive degree. Then, by comparing the sentiment probability distribution of the original edge classifier for w o with the sentiment probability distribution of the antonym edge classifier for w a , a sentiment with high confidence can be obtained. If this sentiment comes from the original edge classifier, then take the sentiment predicted by the original edge classifier as the standard. If this sentiment comes from the antonym edge classifier, then take the opposite sentiment of the sentiment predicted by the antonym edge classifier.

[0144] For easy understanding, let's take an example. w o →C o , w a .→C a If the prediction given by C o is (0.2, 0.8), C aIf the given prediction is (0.3, 0.7), then the prediction sentiment of C o shall prevail, which is positive. Reverse the prediction probability. If the prediction probability of this C a is higher, then the sentiment of this original sample sentence is negative. The result we want is the sentiment of the original sample w o . This confidence μ threshold is a value set by the user. For the binary classification task, it is set to 0.8 / 0.52 for the LSTM model / Bert model; for the five-classification task, it is set to 0.4 / 0.22 for the LSTM model / Bert model. The purpose of setting the confidence threshold is to control the probability distribution of the sentiment within a suitable range and rely on the result of a certain classifier. Other steps and parameters are the same as those in any one of the specific implementation manners one to nine.

[0145] The above are only descriptions of the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention, can make some changes or modifications to the above-disclosed technical content to obtain equivalent embodiments with equivalent changes. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments still fall within the protection scope of the technical solution of the present invention.

Claims

1. A dual sentiment classification method based on diffusion counterfactual data enhancement, characterized by: It includes the following steps:

1. Establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; Obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation; 2. Input the sentiment data to be classified into the trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation to complete dual sentiment classification.

2. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 1, wherein in step 1, establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation; the specific process is as follows: Step 1: Obtain a training set; The training set includes sentiment sample data w o , and the sentiment sample data w o in the training set are all labeled data, and the label data is sentiment s; Step 2: Establish a dual sentiment classification model DCA based on diffusion counterfactual data augmentation; obtain a trained dual sentiment classification model DCA based on diffusion counterfactual data augmentation.

3. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 2, wherein the specific process of establishing a dual sentiment classification model DCA based on diffusion counterfactual data augmentation and obtaining a trained dual sentiment classification model DCA in step 2 is as follows: The dual sentiment classification model DCA based on diffusion counterfactual data augmentation includes: a text generation layer, a dual discrimination layer, a contrast learning layer, and a dual sentiment classification layer; S1: The original sample w in the training set o Input to the text generation layer for processing and construct the antisense guidance paradigm w g ; S2: Construct a counterfactual diffusion generation network model and obtain a trained counterfactual diffusion generation network model; The antonym-guided paradigm w g and the original sample w o Input the trained counterfactual diffusion generation network model and output the countersense sample w a ; S3: Construct the original edge discriminator C o , obtain the trained original edge discriminator C o ; Construct the antonym edge discriminator C a ; Obtain the trained antonym edge discriminator C a ; The trained original edge discriminator C for the dual discrimination layer o and the trained antonym edge discriminator C a are used to optimize the trained counterfactual diffusion generation network model to obtain an optimized counterfactual diffusion generation network model; the antonym guidance paradigm w g and the original sample w o are input into the optimized counterfactual diffusion generation network model, and the updated antonym sample w a1 is output; S4: Perform final optimization on the optimized counterfactual diffusion generation network model obtained in S3 to obtain a qualified counterfactual diffusion generation network model. The qualified counterfactual diffusion generation network model is input into the antisense guidance paradigm w g and the original sample w o , output qualified antonym sample w a2 ; S5: The original edge discriminator C trained in S3 o and antisense edge identification as the original edge classifier and antisense edge classifier of the sentiment classification layer; then the original sample w o and the qualified antisense sample w generated by S4 a2 The samples are transferred to the sentiment classification layer in the form of sample pairs for dual sentiment classification.

4. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 3, wherein In S1, the original emotion sample sequence w in the training set is o Input to the text generation layer to generate the antonym guidance paradigm w g The specific process is: S1.1: Set the stop words in the original sample w o as non-replacement words, where the original sample w o is composed of an original sample sequence w o ={w1, w2,..., w n}, and the original sample sequence is composed of n words; n is a positive integer; S1.2: The original sample w o Adjectives, adverbs, and verbs that have antonyms are replaced with their corresponding antonyms from the English vocabulary database WordNet; S1.3: The original sample w o Adjectives, adverbs, verbs, and non-stop words that do not have antonyms are replaced with their corresponding synonyms in the English vocabulary database WordNet; S1.4: The original sample after replacement according to S1.1 to S1.3 and the emotion label of the original sample (w o ,s) Generate antonym guidance paradigm and opposite sentiment labels Where s represents the original sample w o emotional labels, Represents opposite sentiment labels.

5. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 4, wherein In S1.2, the original sample w o Adjectives, adverbs, and verbs that have antonyms are replaced with their corresponding antonyms in the English vocabulary database WordNet. The specific process is: For the original sample sequence w o ={w1, w2,..., w n}, for each word w k , the probability of replacing the word w k with the j-th word in the English vocabulary database WordNet is calculated using supervised multi-label classification; the specific process is as follows: Among them, indicates that the j-th word in the English vocabulary database WordNet belongs to the k-th word w o in the original sample w k the probability of the replacement word set, indicates the original sample w o the k-th word w k corresponds to the j-th word in the English vocabulary database WordNet, exp() represents the exponential function, h k is the hidden representation of the word w k W j is the weight matrix, b j is the bias, and m is the size of the English vocabulary database WordNet; Then the original sample sequence w o ={w1,w2,...,w n The replacement probability of words with a replacement probability lower than 0.57 and words not included in the English vocabulary database WordNet is set to zero, and the decision of whether to replace word w with the jth word in the English vocabulary database WordNet is obtained. k The normalized probability distribution w k ~Multinomial(P k ); Determine whether to replace word w with the jth word in the English vocabulary database WordNet k The normalized probability distribution w k ~Multinomial(P k ); the specific process is: Among them, P k represents the polynomial word distribution of the replacement probability for the k-th word w o in the original sample w k in the English vocabulary database WordNet, and normalize() represents the normalization of the probabilities belonging to the replacement word set in the English vocabulary database WordNet. represents the probability that the first word in the English vocabulary database WordNet belongs to the replacement word set of the k-th word w o in the original sample w k , and Multinomial represents the multinomial formed by sampling w1 to w o in the original sample w n .

6. The dual sentiment classification method based on diffusion counterfactual data enhancement according to claim 5 is characterized in that the specific process of constructing a counterfactual diffusion generation network model in S2 and obtaining a pre-trained counterfactual diffusion generation network model is as follows: S2.1: Take the original sample w o and the antisense guiding paradigm w g constructed by S1 and concatenate them as a discrete sample Embed the obtained discrete sample into the continuous feature space for a forward conditional noise addition process to obtain the word vector x0. The specific process is as follows: in, Indicates that the discrete samples Embed the continuous feature space to perform a forward process of continuous diffusion, N() represents the multidimensional normal distribution, Emb() represents the embedding function, β0 is the amount of noise added in the time step, and I is the unit matrix; S2.2: Perform a T-step forward conditional noise addition process on the word vector x0 obtained in S2.1 to obtain T Gaussian noise variables x1,..., x T , x1,..., x T Each of the Gaussian noise variables in x1,..., x represents the output of an intermediate step in the forward conditional noise addition process, and T is a positive integer. The specific process is as follows: When performing the forward conditional noise addition process at each time step \(t\in(1,2,\cdots,T)\), Gaussian noise \(\beta\) is added to the Gaussian noise \(x\) obtained in the previous step to obtain the Gaussian noise \(x_t\) at the \(t\)-th time step. The formula is as follows: t-1 in \(\beta\) t to obtain the Gaussian noise \(x_t\) at the \(t\)-th time step. t The formula is as follows: Among them, q(x t |x t-1 ) represents the Gaussian noise x t-1 Perform a forward process of continuous diffusion; S2.3: Perform a T-step reverse conditional denoising process on the Gaussian noise x obtained in S2.2 to obtain the denoised word vector x0. The specific process is as follows: T ​ p θ (x t-1 |x t ,t) = N(x t-1 ; μ θ (x t ,t), Σ θ (x t ,t)) Among them, p θ (x t-1 |x t , t) represents performing a single reverse conditional denoising process on the Gaussian noise x t , where μ θ (·) and Σ θ (·) are the Gaussian distribution parameters learned during the reverse conditional denoising process; S2.4: Calculate the variational lower bound according to the forward noise addition process in S2.2 and the backward denoising process in S2.3; use the calculated variational lower bound as the loss function for training the counterfactual diffusion generation network model to obtain a pre-trained counterfactual diffusion generation network model; The formula for the loss function of training the counterfactual diffusion generation network model is: Among them, E q (·) represents the expectation function, D KL (q(x T |x0)||p θ (x T )) represents the forward process of diffusion, represents the reverse process of diffusion, represents the discrete sample embedding into the continuous space, is the rounding operation, D KL represents the calculation formula of the KL divergence between Gaussian distributions, x T represents Gaussian noise.

7. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 6, wherein The original edge discriminator C is constructed in S3 o , obtain the trained original edge discriminator C o ; The specific process is: When the original edge discriminator C o is an LSTM model, the original sample w o and its corresponding sentiment label s are used as the training set of the original edge discriminator C o ; obtain the trained original edge discriminator C o ; When the original edge discriminator C o is the pre-trained model Bert-base or Bert-large, the original sample w o and its corresponding sentiment label s are used as the training set of the original edge discriminator C o ; the trained original edge discriminator C is obtained o ; Construct the antonym edge discriminator C in S3 a ; Obtain the trained antonym edge discriminator C a ; The specific process is as follows: When the antonym edge discriminator C a is an LSTM model, the antonym sample w a generated by the pre-trained counterfactual diffusion generation network model obtained in S2 and its corresponding sentiment label s are used as the training set to obtain the trained antonym edge discriminator C a ; When the antonym edge discriminator C a is the pre-trained model Bert-base or Bert-large, the antonym sample w a generated by the pre-trained counterfactual diffusion generation network model obtained in S2 and its corresponding sentiment label a .

8. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 7, wherein The dual discriminator layer in S3 uses the trained original edge discriminator C o and the trained antisense edge discriminator C a The trained counterfactual diffusion generative network model is optimized to obtain an optimized counterfactual diffusion generative network model. The specific process is as follows: S3.1: Input the antonym sample w a into the trained original edge discriminator C o to obtain the output p ori (s|w a ). Input the antonym sample w a into the trained antonym edge discriminator C a to obtain the output p ant (s|w a ); The original edge discriminator C o It contains hidden layer and output layer in sequence, and the output of the hidden layer h o As the original edge discriminator C o The input of the output layer, where h o represents the antonym sample w a In the original edge discriminator C o Hidden representation within; The anti-sense edge discriminator C a sequentially includes a hidden layer and an output layer, and the output of the hidden layer is h a as the input of the output layer of the anti-sense edge discriminator C a where h a represents the hidden representation of the anti-sense sample w a in the original edge discriminator C a inside the hidden representation; Original edge discriminator C o and antisense edge discriminator C a The output layer is the softmax layer, and the output is as follows: p ori (s|w a )=softmax(W ori h o +b ori ) p ant (s|w a ) = softmax(W ant h a + b ant ) Among them, W ori and b ori represent the parameters of C o ; W ant and b ant represent the parameters of C a ; S3.2: The dual discrimination layer performs dual discrimination based on the outputs of the original edge discriminator C o and the antonym edge discriminator C a and generates rewards simultaneously. The specific process is as follows: where α and γ are negatively correlated trade-off parameters whose sum is 1, S3.3: Extract M antisense samples w output by the pre-trained counterfactual diffusion generation network model obtained in S2 a Perform the double discrimination in S3.2, obtain M rewards, and calculate the average reward of the M rewards as the baseline reward r avg , and according to the obtained baseline reward r avg Perform reinforcement learning to obtain an optimized counterfactual diffusion generation network model; the specific process is as follows: When r(w a ) is greater than the average reward r avg When the counterfactual diffusion generation network model is added to generate the countersense sample w a The probability of When r(w a ) is less than or equal to the average reward r avg When the counterfactual diffusion generation network model is reduced to generate countersense samples w a probability; the loss function formula of reinforcement learning is as follows: L r = -log(r(w a ) - r avg )P G (w a |w o ) where \(r(w a ) - r avg represents the difference between the reward of the antonym sample \(w a and the baseline reward \(r avg as the actual calculated reward of the antonym sample \(w a , and \(P G (w a |w o ) represents the probability that the counterfactual diffusion generation network model generates the antonym sample \(w o given the original sample \(w a .

9. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 8, wherein in S4, perform final optimization on the optimized counterfactual diffusion generation network model obtained in S3 to obtain a qualified counterfactual diffusion generation network model; the specific process is as follows: S4.1: Apply the antisense guidance paradigm obtained in S1 to w g As an anchor point, the updated antonym sample w generated by the counterfactual diffusion generation network model optimized by S3 is a1 As a positive sample, the original sample w o As a negative sample, contrastive learning is performed. The loss function formula of contrastive learning is as follows: where ε represents the safety margin, f(·) represents the embedding function, represents the square of the L2 norm, max(·) represents the maximum value function, S4.2: Obtain the reinforcement learning loss function L obtained in S4.1 r and the contrastive learning loss function L obtained in S3.3 r and the counterfactual diffusion generation network loss function L obtained in S2.4 vlb Add them up to train and optimize the counterfactual diffusion generation network model with the overall loss function, and obtain a qualified counterfactual diffusion generation network model. The formula for the overall loss function is as follows: L = L vlb + L r + L c Among them, L vlb is the loss function of the counterfactual diffusion generation network, and L r is the loss function of reinforcement learning, and L c is the loss function of contrastive learning.

10. The dual sentiment classification method based on diffusion counterfactual data augmentation according to claim 9, wherein, In S5, the trained original edge discriminator C obtained in S3 o and the antonym edge are used as the original edge classifier and the antonym edge classifier of the sentiment classification layer; then the original sample w o and the qualified antonym sample w generated by S4 a2 are transmitted to the sentiment classification layer in the form of sample pairs for dual sentiment classification. The specific process is as follows: Calculate the confidence of the original edge classifier and the antisense edge classifier. When the original edge classifier exceeds the antisense edge classifier confidence threshold μ, the original sample w o The sentiment label sticks to the result of the original classifier, whereas the original sample w o The sentiment label of w is regarded as a qualified antonym sample. a2 Emotion; get the original sample w o The emotional label of is determined by the following formula: Among them, μ represents the confidence threshold, p(s|w o ) represents the original sample w o Emotional prediction, p ori (s|w o ) represents the original edge classifier for the original sample w o Emotional prediction, p ant (s|w a2 ) represents the antisense edge classifier for qualified antisense samples w a2 emotional predictions.

Citation Information

Patent Citations

  • Deep learning theme emotion classification method based on data enhancement

    CN110245229A

  • Text generation method and system based on word replacement and storage medium

    CN116090425A

  • Network alignment method based on anti-fact inference

    CN116506302A

  • Attribute-level emotion tetrad prediction system and method based on data enhancement and self-training

    CN116756318A

  • Autonomous mobile robot capable of using elevator

    KR1020250063473A