Data processing method for large model training

By using a psychological model-based correlated emotion distribution model and an improved convolutional neural network method, the problem of insufficient quality of text sentiment analysis large model training data sets is solved, and more accurate data set annotation and model emotion prediction effects are achieved.

CN120216695APending Publication Date: 2025-06-27BEIJING TAOYU DIGITAL TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510279422.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The marked text data set used for text sentiment analysis large model training in the prior art is insufficient in quality, resulting in low model quality, and the cost of obtaining real data is high and there is a risk of infringing personal information.

Method used

The correlation emotion distribution model based on psychological models is used to measure the emotional correlation of product review text expression from a psychological perspective, generate more accurate data set annotations, and perform product text emotional intensity learning through improved convolutional neural networks.

Benefits of technology

The dependence of text sentiment analysis large models on the amount of learning data is reduced, the quality of data set labeling is improved, so that the model can learn product emotions more accurately, and efficient prediction of text emotions is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216695A_ABST
    Figure CN120216695A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model training data processing, and provides a data processing method for large model training, which comprises the following steps of: generating an emotion label set by using a commodity comment text through a feature emotion word combination word segmentation method, and performing word segmentation on the commodity comment text through an emotion dictionary word bank to generate an emotion word set; respectively generating first emotion association distribution and second emotion association distribution from the emotion label set and the emotion word set through an association emotion distribution model based on a psychological model in sequence; performing weighted accumulation on the first emotion association distribution and the second emotion association distribution to calculate comprehensive emotion intensity; and carrying out commodity text sentiment intensity learning on the text sentiment analysis large model based on the improved convolutional neural network by using the sentiment word set and the comprehensive sentiment intensity. According to the method, the emotion relevance expressed by the commodity comment text is measured from the psychological perspective, so that more accurate data set labels are generated, and the dependence of a large text emotion analysis model on the learning data volume is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model training data processing, and in particular, to a data processing method for large model training. Background Art

[0002] The large model for text sentiment analysis is a multi-sentiment analysis model used to process text with emotional ambiguity. Its core idea is to record the expression degree of examples on each emotion through emotional distribution. Different from traditional single-label or multi-label learning, emotional distribution learning can quantitatively model multiple emotions simultaneously.

[0003] The development of artificial intelligence models, especially large models for text sentiment analysis, highly depends on the scale, quality, and type of training data. Traditionally, training data is mainly real data collected from the real world. However, with the rapid development of the artificial intelligence industry, real data resources are being consumed rapidly and tend to be exhausted. In addition, real data not only has a high acquisition cost but also has the risk of infringing personal information when used directly.

[0004] Therefore, the above problems lead to a serious shortage of the labeled text data set for training large models for text sentiment analysis. Therefore, it is necessary to improve the annotation method of the data set to improve the annotation quality of the data set to ensure the quality of the model.

[0005] For example, the invention patent with the application number 202010622293.2 discloses a method, system, device, and medium for generating a training data set based on labeled text. However, its training data set cannot be used for large models for text sentiment analysis.

[0006] Therefore, how to improve the large model for text sentiment analysis and its data set annotation process is a technical problem to be solved currently. Summary of the Invention

[0007] For this reason, the present invention provides a data processing method for large model training. By using the associated emotional distribution model based on the psychological model, it realizes the measurement of the emotional relevance expressed by product review texts from a psychological perspective, and then generates a more accurate data set annotation, reduces the dependence of the large model for text sentiment analysis on the amount of learning data, and also enables the large model for text sentiment analysis to learn more accurate product emotions through the data set.

[0008] To achieve the above object, the present invention proposes a data processing method for large model training, including:

[0009] Generating an emotional label set by combining the feature emotional word segmentation method for product review texts, and segmenting the product review texts through an emotional dictionary thesaurus to generate an emotional word set;

[0010] The set of emotion tags and the set of emotion words are respectively passed through an associated emotion distribution model based on a psychological model to generate a first emotion association distribution and a second emotion association distribution in sequence;

[0011] The first emotion association distribution and the second emotion association distribution are used to calculate the comprehensive emotion intensity through weighted accumulation;

[0012] The set of emotion words and the comprehensive emotion intensity are used for learning the emotion intensity of commodity texts for a large model of text emotion analysis based on an improved convolutional neural network.

[0013] Further, the process of respectively generating the first emotion association distribution and the second emotion association distribution by passing the set of emotion tags and the set of emotion words through the associated emotion distribution model based on the psychological model in sequence includes:

[0014] Select the first main emotion word tag from the set of emotion tags, and the emotion words in the set of emotion words correspond to multiple emotion word tags, and select the second main emotion word tag from the multiple emotion word tags;

[0015] Take the set of emotion tags as the variable of the associated emotion distribution model, and take the first main emotion word tag as the expected value of the associated emotion distribution model to calculate the first emotion association distribution;

[0016] Take the emotion word tags as the variable of the associated emotion distribution model, and take the second main emotion word tag as the expected value of the associated emotion distribution model to calculate the second emotion association distribution;

[0017] Among them, the associated emotion distribution model is constructed based on the discrete lognormal distribution.

[0018] Further, the associated emotion distribution model has a normalization factor for normalizing the first set of emotion words and the second set of emotion words.

[0019] Further, the psychological model is the Russell circular emotion model, the first emotion association distribution includes a first emotion valence association distribution and a first emotion arousal degree association distribution, and the second emotion association distribution includes a second emotion valence association distribution and a second emotion arousal degree association distribution.

[0020] Further, the process of calculating the comprehensive emotion intensity by weighted accumulation of the first emotion association distribution and the second emotion association distribution includes:

[0021] Accumulate multiple second emotion valence association distributions to generate a tag valence, accumulate multiple second emotion valence association distributions to generate a tag emotion valence, and accumulate multiple second emotion arousal degree association distributions to generate a tag arousal degree;

[0022] Generate the comprehensive emotion intensity through weighted summation based on the label emotion valence, the label arousal degree, the first emotion valence correlation distribution, and the second emotion correlation distribution.

[0023] In the above solution, by constructing the emotion valence and emotion arousal degree based on the Russell circular emotion model, it is possible to generate emotion label values that are more in line with human psychological emotions towards commodities, thereby making the emotion label values more accurate and reducing the dependence of the text emotion analysis large model on the amount of learning data.

[0024] Further, the process of using the first emotion word set and the comprehensive emotion intensity for learning the commodity text emotion intensity of the text emotion analysis large model based on the improved convolutional neural network in the input layer, convolutional pooling layer, and fully connected layer of the text emotion analysis large model includes:

[0025] Generate emotion word vectors from the first emotion word set through the word embedding model of the input layer;

[0026] Extract emotion features from the emotion word vectors through the convolutional pooling layer;

[0027] Generate text emotion classification from the emotion features and the comprehensive emotion intensity through the improved loss function of the fully connected layer.

[0028] Further, the improved loss function is constructed based on the cross-entropy metric according to the comprehensive emotion intensity of all labels.

[0029] Further, a channel attention mechanism is set after the convolutional pooling layer, and the channel attention mechanism performs weighted processing on the emotion features.

[0030] In the above solution, it realizes the embedding of emotion label values based on the psychological model into a text emotion analysis large model with a simple and efficient structure, and realizes the efficient prediction of text emotions.

[0031] Further, the process of generating an emotion label set from the commodity review text through the feature emotion word combination word segmentation method includes:

[0032] Match and identify feature words and emotion words from the commodity review text through the commodity feature word library and the commodity emotion word library;

[0033] Select emotion words that meet the feature emotion word combination from the feature words and the emotion words to generate a temporary emotion word set;

[0034] Generate the emotion label set according to the temporary emotion word set.

[0035] Furthermore, the emotion label set includes emotion property labels and emotion intensity labels. The process of generating the emotion label set according to the temporary emotion word set includes:

[0036] Determining the emotion property labels according to the feature words corresponding to the emotion words in the temporary emotion word set and the corresponding scores of the product review texts;

[0037] Calculating the emotion intensity labels according to the occurrence times of the emotion property labels in the product review texts.

[0038] In the above solution, a method for generating emotion labels that conforms to product characteristics is directly constructed without relying on the existing word segmentation dictionary, which realizes that the data set better meets the needs of product emotion classification, further reduces the dependence of the text emotion analysis large model on the learning data volume, and also makes the text emotion analysis large model more conform to product characteristics.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. Through the associated emotion distribution model based on the psychological model, the emotional relevance expressed by the product review texts is measured from a psychological perspective, and then a more accurate data set annotation is generated, reducing the dependence of the text emotion analysis large model on the learning data volume, and also enabling the text emotion analysis large model to learn more accurate product emotions through the data set.

[0041] 2. By constructing the emotion valence and emotion arousal degree based on the Russell circular emotion model, emotion label values that are more in line with human psychological emotions towards products can be generated, making the emotion label values more accurate and reducing the dependence of the text emotion analysis large model on the learning data volume.

[0042] 3. The emotion label values based on the psychological model are embedded into a text emotion analysis large model with a simple and efficient structure, realizing efficient prediction of text emotions.

[0043] 4. A method for generating emotion labels that conforms to product characteristics is directly constructed without relying on the existing word segmentation dictionary, which realizes that the data set better meets the needs of product emotion classification, further reduces the dependence of the text emotion analysis large model on the learning data volume, and also makes the text emotion analysis large model more conform to product characteristics. Description of the Drawings

[0044] Figure 1 It is a general process schematic diagram of the data processing method for large model training according to an embodiment of the present invention;

[0045] Figure 2 It is a schematic diagram of the comprehensive emotion intensity generation process of the data processing method for large model training according to an embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the generation process of the sentiment label set for the data processing method for large model training according to an embodiment of the present invention;

[0047] Figure 4 Schematic diagram of the improved convolutional neural network structure for the data processing method for large model training according to an embodiment of the present invention. Detailed implementation manners

[0048] In order to make the objectives and advantages of the present invention more clear and understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0050] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0051] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0052] As Figures 1 to 4 shown, the present invention provides a data processing method for large model training. By implementing the associated sentiment distribution model based on the psychological model, the emotional relevance expressed in the product review text is measured from a psychological perspective, and then a more accurate dataset annotation is generated, reducing the dependence of the large model for text sentiment analysis on the learning data volume, and also enabling the large model for text sentiment analysis to learn more accurate product sentiment through the dataset.

[0053] As Figures 1 to 4As shown in the figure, this embodiment proposes a data processing method for large model training, including: generating an emotional label set from commodity review texts through the feature emotional word combination word segmentation method, and segmenting the commodity review texts through an emotional dictionary thesaurus to generate an emotional word set;

[0054] Sequentially passing the emotional label set and the emotional word set through an associated emotional distribution model based on a psychological model to generate a first emotional association distribution and a second emotional association distribution respectively;

[0055] Calculating the comprehensive emotional intensity by weighted accumulation of the first emotional association distribution and the second emotional association distribution;

[0056] Using the emotional word set and the comprehensive emotional intensity to perform commodity text emotional intensity learning on a text emotional analysis large model based on an improved convolutional neural network.

[0057] It should be noted that the text emotional analysis large model is used for the classification and judgment of the emotional tendency of human language. The text emotional analysis large model in this embodiment is preferably used to judge its positive emotional tendency level and negative emotional tendency level, and is preferably applied to the classification of the positive tendency level and negative emotional tendency level of mobile phone commodity reviews, so as to determine the content of different degrees of satisfaction or dissatisfaction of users with mobile phone commodities, laying a foundation for realizing more accurate emotional grading of commodity review content.

[0058] Furthermore, as Figure 2 shown in the figure, the process of sequentially passing the emotional label set and the emotional word set through an associated emotional distribution model based on a psychological model to generate a first emotional association distribution and a second emotional association distribution respectively includes:

[0059] Regarding the emotional labels in the emotional label set as the first main emotional word labels, corresponding the emotional words in the emotional word set to multiple emotional word labels, and regarding the multiple emotional word labels as the second main emotional word labels;

[0060] Regarding the emotional label set as the variable of the associated emotional distribution model, and using the first main emotional word label as the expected value of the associated emotional distribution model to calculate the first emotional association distribution;

[0061] Regarding the emotional word labels as the variables of the associated emotional distribution model, and using the second main emotional word label as the expected value of the associated emotional distribution model to calculate the second emotional association distribution;

[0062] Among them, the associated emotional distribution model is constructed based on the discrete lognormal distribution.

[0063] It is understandable that based on a psychological model, the psychological distance between emotions can be defined, and based on this distance, the set of emotion labels and the set of emotion words are respectively discretized into multiple first emotion association distributions / second emotion association distributions that conform to the lognormal distribution, realizing emotion label annotation using the improved emotion distribution labeling enhancement method (Lexicon-based Emotion Distribution Label Enhancement, LLE). LLE is an emotion analysis method based on an emotion lexicon, aiming to enhance the emotion distribution labeling of text by utilizing the prior knowledge in the emotion lexicon, and is used for tasks such as emotion classification and emotion intensity prediction, which can improve the accuracy and robustness of large models in emotion analysis learning. The emotion lexicon contains words and their corresponding emotion property labels and emotion intensity labels. Emotion labels in this embodiment include emotion property labels and emotion intensity labels, and emotion property labels include high negative, medium negative, low negative, neutral, low positive, medium positive, and high positive. For example, the emotion property label corresponding to the word "easy to use" in the emotion lexicon is "positive", and the emotion intensity label is 0.8. However, only using the emotion lexicon for the emotion property labels and emotion intensity labels of words does not consider their context relationship. For example, an ironic product evaluation may be labeled as positive, which is inaccurate and not conducive to the large model of text emotion analysis to accurately learn the emotion intensity of product text based on high-quality labeled data.

[0064] Specifically, for text s i , the process of modeling its set of emotion labels {y i} as a lognormal distribution is as follows:

[0065]

[0066] In the formula, represents the first emotion association distribution, Z1 is the normalization factor, σ is the standard deviation, preferably 1, |y i - a1| represents the distance between the emotion label and the first main emotion word label in the psychological model, where a1 is the first main emotion word label, and y i represents the emotion label.

[0067] Specifically, for text s i , the emotion word label p t generated by it through the emotion lexicon, the process of modeling the emotion word label as a lognormal distribution is as follows:

[0068]

[0069] In the formula, represents the second emotion association distribution, Z2 is the normalization factor, σ is the standard deviation, preferably 1, |p t-a2| represents the distance between the emotional word label and the second primary emotional word label in the psychological model, where a2 is the second primary emotional word label, and p t represents the emotional word label.

[0070] Therefore, by substituting the distances between the primary emotional label and other labels in the psychological model into the above formula, the value of the emotional association distribution can be obtained. Moreover, by modeling it as a lognormal distribution, compared with modeling it as a normal distribution, the magnitude of the distance in the psychological model can be reflected to a greater extent.

[0071] Furthermore, the psychological model is the Russell Circumplex Model of Affect. The first emotional association distribution includes the first emotional valence association distribution and the first emotional arousal association distribution. The second emotional association distribution includes the second emotional valence association distribution and the second emotional arousal association distribution.

[0072] It can be understood that the Russell Circumplex Model of Affect describes emotions as a two-dimensional circular space, defining emotional states through two main dimensions: valence and arousal. Among them, valence includes positive valence and negative valence, which is used to describe the positive and negative nature of emotions from positive to negative, that is, whether the emotion is positive or negative. Arousal can describe the intensity level of emotions. The adjacent distributions of valence and arousal in the Russell Circumplex Model of Affect can be used to construct a Gaussian model.

[0073] Specifically, |y i -a1| representing the distance between the emotional label and the first primary emotional word label in the Gaussian model includes the valence distance for generating the first emotional valence association distribution and the arousal distance for generating the first emotional arousal association distribution. |p t -a2| representing the distance between the emotional word label and the second primary emotional word label in the Gaussian model includes the valence distance for generating the second emotional valence association distribution and the arousal distance for generating the second emotional arousal association distribution. For example, if the emotional property label is positive and the emotional intensity label is 0.6, the corresponding valence distance is positive 0.6, and if the emotional property label is negative, it is negative 0.6. The arousal distance is determined according to the word generating the label. For example, the arousal distance between the word "easy to use" and the word "very satisfied" is 0.7, which is the positive value of the difference in their emotional intensity labels, and the arousal distance between the word "easy to use" and the word "so-so" is 0.3, which is the negative value of the difference in their emotional intensity labels. Therefore, by substituting the valence distance and the arousal distance into |y i -a1| and |p t -a2| in the above formula, the first emotional association distribution and the second emotional association distribution can be obtained.

[0074] Therefore, the center with the main emotion label following a lognormal distribution is achieved, where the value of its emotion association distribution is the largest. The closer the distance on the psychological model is to the main emotion label, the higher the correlation, and the larger the value of the emotion association distribution.

[0075] Furthermore, the association emotion distribution model has a normalization factor for normalizing the first emotion word set and the second emotion word set.

[0076] Specifically, the calculation process of the normalization factor for the first emotion association distribution is as follows:

[0077]

[0078] In the formula, Z1 is the normalization factor for the first emotion association distribution, σ is the standard deviation, preferably 1, and |y i - a1| represents the distance between the emotion label and the first main emotion word label on the psychological model, where a1 is the first main emotion word label, and y i represents the emotion label.

[0079] Specifically, the calculation process of the normalization factor for the second emotion association distribution is as follows:

[0080]

[0081] In the formula, Z2 is the normalization factor for the second emotion association distribution, σ is the standard deviation, preferably 1, and |p t - a2| represents the distance between the emotion word label and the second main emotion word label on the psychological model, where a2 is the second main emotion word label, and p t represents the emotion word label.

[0082] Therefore, substituting the distance between the main emotion label and other labels on the psychological model into the above formula can obtain the value of the normalization factor.

[0083] Therefore, the value of the first emotion association distribution / second emotion association distribution is within the range of 0 to 1, meeting the standard of text emotion intensity annotation.

[0084] Furthermore, the process of calculating the comprehensive emotion intensity by weighted accumulation of the first emotion association distribution and the second emotion association distribution includes:

[0085] After accumulating multiple of the second emotional valence association distributions, a label valence is generated. After accumulating multiple of the second emotional valence association distributions, a label emotional valence is generated. After accumulating multiple of the second emotional arousal degree association distributions, a label arousal degree is generated. Based on the label emotional valence, the label arousal degree, the first emotional valence association distribution, and the second emotion association distribution, a weighted sum is performed to generate the comprehensive emotional intensity. Specifically:

[0086]

[0087] In the formula, Score i represents the comprehensive emotional intensity of the i-th label, α represents the weight coefficient, b i , a k respectively represent the number of emotion words and the number of emotion labels. represents the label valence, where represents the second emotional valence association distribution. represents the label arousal degree, where represents the second emotional arousal degree association distribution. respectively represent the first emotional valence association distribution and the second emotion association distribution.

[0088] In the above solution, by constructing the emotional valence and emotional arousal degree based on the Russell circular emotion model, it is possible to generate an emotion label value that is more in line with human psychological emotions towards commodities, thereby making the emotion label value more accurate and reducing the dependence of the text emotion analysis large model on the amount of learning data.

[0089] Furthermore, as Figure 4 shown, for the input layer, convolutional pooling layer, and fully connected layer of the text emotion analysis large model, the process of using the first emotion word set and the comprehensive emotional intensity for learning the commodity text emotional intensity of the text emotion analysis large model based on the improved convolutional neural network includes:

[0090] Generate an emotion word vector from the first emotion word set through the word embedding model of the input layer; extract emotion features from the emotion word vector through the convolutional pooling layer; generate a text emotion classification from the emotion features and the comprehensive emotional intensity through the improved loss function of the fully connected layer.

[0091] The word embedding model is preferably a Word2Vec model, which is used to map vocabulary to a high-dimensional vector space and can thus be used for downstream natural language processing.

[0092] It can be understood that the large text sentiment analysis model constructed in this embodiment can be used for the Emotion Distribution Learning task, and good results have been achieved in predicting the emotion distribution through a convolutional neural network in experiments.

[0093] The input layer plays the role of converting the input text into corresponding word vectors. At the same time, it also adjusts the dimension of the input to adapt to the requirements of subsequent convolutional operations. The convolutional operation of the convolutional pooling layer slides in one direction, that is, it performs convolutional operations in the length direction of the text sequence, thereby effectively extracting local features in the text. The pooling operation is responsible for dimensionality reduction and aggregation of the features output by the convolutional layer, so as to screen out the most important features. During the pooling operation, the maximum pooling strategy is adopted, that is, the maximum eigenvalue is selected from the feature vectors generated by each sliding window. The obtained feature vectors are connected and predicted through a fully connected layer. The role of the fully connected layer is to map the output of the pooling layer to the final classification result. Through training and learning, it can effectively convert text features into classification labels of specific product text sentiment.

[0094] Specifically, during the forward propagation process of the model, multi-scale convolutional kernels are applied cyclically. By iteratively processing the input text with convolutional kernels of different sizes, it is possible to capture feature information of different scales and pass it to the subsequent channel attention mechanism for weighted processing.

[0095] The channel attention mechanism is incorporated into the forward propagation method of the model. The channel attention mechanism is applied to the output of the convolutional layer, and through the softmax function of the fully connected layer, the classification task of different features is carried out, so as to achieve enhanced attention to key features and thus improve the performance of the model.

[0096] In a specific implementation, the PyCharm compiler and the PyTorch 1.1 deep learning framework are used, and the model parameters are set as follows: the number of training epochs: 20, the number of samples processed each time during training (Batch_size): 128, the size of the padding operation (Pad_size): 32, the learning rate (Learn_rate): 0.01, the ratio of randomly discarding neurons to prevent overfitting (dropout): 0.5, the size of the convolutional kernel (Filter_size): 2, 3, 4, and the number of filters in each convolutional network layer (Num_filters): 256.

[0097] It can be understood that the neurons of the fully connected layer output the probability distribution of all emotion labels, and the objective function of the model is to improve the loss function.

[0098] Furthermore, the improved loss function is constructed based on the cross - entropy metric according to the comprehensive sentiment intensity of all labels.

[0099] The specific form of the improved loss function is as follows:

[0100]

[0101] In the formula, E(s,d) is the improved loss function, N is the total number of samples, Score j si represents the j - th comprehensive sentiment intensity of the i - th sentiment label of the text, and is the activation function value of each unit in the fully - connected layer.

[0102] Furthermore, as Figure 4 shown, a channel attention mechanism is set after the convolutional pooling layer, and the channel attention mechanism weights the sentiment features.

[0103] The channel attention mechanism focuses on adjusting the weights of each channel in the feature map, enabling the network to focus on the features that are more critical for the task. Since the spatial attention mechanism focuses on adjusting the weights of different spatial positions in an image or feature map, in the text classification task, the spatial attention mechanism is not necessary. Text data is usually represented as a series of words or feature vectors, rather than a two - dimensional spatial structure like an image. Therefore, only the channel attention mechanism is set in this embodiment.

[0104] In the above - mentioned solution, the emotion label values based on the psychological model are embedded into a simple and efficient large - scale text sentiment analysis model, realizing the efficient prediction of text emotions.

[0105] Furthermore, as Figure 3 shown, the process of generating a sentiment label set from the commodity review text through the feature - sentiment word combination tokenization method includes:

[0106] Matching and identifying feature words and sentiment words from the commodity review text through the commodity feature word library and the commodity sentiment word library;

[0107] Selecting the sentiment words that meet the feature - sentiment word combination from the feature words and the sentiment words to generate a temporary sentiment word set;

[0108] Calculating the sentiment intensity label according to the occurrence times of the sentiment property label in the commodity review text.

[0109] Specifically, the sequence of content words after word segmentation and part-of-speech tagging contains a large number of words that are irrelevant to product features and user opinions, so extraction is required. In this paper, based on the product feature word library and the product sentiment word library, the sequence of content words is matched and recognized. If the feature words and sentiment words matched in the sequence of content words conform to the feature-sentiment word combination, the sentiment words are selected and a temporary sentiment word set is generated.

[0110] It can be understood that due to the dependence of the sentiment label of the sentiment word on the context, the same sentiment word may express completely different meanings for a specific comment object (feature word). For example, in "The phone runs fast" and "The phone consumes power fast", "fast" as a sentiment word is used to describe the product feature, but the expressed sentiment tendency is opposite, and the generated sentiment label should also be different. Therefore, when performing sentiment analysis on the feature word of the user comment and its corresponding opinion word as a whole, it is necessary to consider the collocation relationship between the comment object and the opinion word, that is, to construct the feature-sentiment word combination of the feature word and the sentiment word.

[0111] Specifically, the feature-sentiment word combination includes feature word - sentiment word, feature word - sentiment word - sentiment word, feature word - feature word - sentiment word, such as "suitable size", "the screen is big enough and clear enough", "the pixel resolution is high enough" respectively.

[0112] Furthermore, as Figure 3 shown, the process of generating the sentiment label set according to the temporary sentiment word set includes:

[0113] Determine the sentiment property label according to the feature word corresponding to the sentiment word in the temporary sentiment word set and the corresponding score of the product comment text;

[0114] Calculate the sentiment intensity label according to the number of occurrences of the sentiment property label in the product comment text.

[0115] Most of the ratings on e-commerce platforms adopt a five-star rating, that is, users express their evaluations of products and services through the level of stars. In this embodiment, it is stipulated that a 1-star and 2-star rating indicates a negative comment, a 3-star rating may correspond to a negative, neutral or positive comment, and 4-star and 5-star ratings indicate positive comments. This system has the problem of fuzzy definition of sentiment intensity, especially the polarity determination of the 3-star rating is not clear. Therefore, in this embodiment, the sentiment intensity of the opinion word is further divided into 7 levels, and each level corresponds to only one comment polarity, that is, the label intensity level set T = {-3, -2, -1, 0, 1, 2, 3}, and the corresponding sentiment property labels are high negative, medium negative, low negative, neutral, low positive, medium positive, high positive respectively. Among them, the high, medium and low of positive and negative are determined according to the feature word. For example, "power consumption" is high and "running speed" is medium, realizing the one-to-one correspondence between the star rating and the comment polarity and the sentiment property label.

[0116] Use \(T = \{-3, -2, -1, 0, 1, 2, 3\}\) to represent the label strength levels, where each label strength level corresponds to an emotional intensity set. In each emotional intensity set, the emotional intensity label is calculated through the following formula:

[0117]

[0118] In the formula, \(A\) represents the emotional intensity label, and \(q\) ij represents the number of occurrences of the \(i\)-th emotional word in the label strength level \(j\). Therefore, represents the frequency of the \(i\)-th emotional word in the label strength level \(j\). To avoid excessive differences between emotional intensity labels caused by frequencies, the emotional intensity labels can be normalized.

[0119] In the above solution, a method for directly constructing emotional label generation that conforms to the characteristics of commodities without relying on existing word segmentation libraries is realized, which makes the data set more in line with the needs of commodity emotional classification, further reduces the dependence of the text emotional analysis large model on the amount of learning data, and also makes the text emotional analysis large model more in line with the characteristics of commodities.

[0120] It can be understood that in this embodiment, through the associated emotional distribution model based on the psychological model, the emotional relevance expressed by commodity review texts is measured from a psychological perspective, and then a more accurate data set annotation is generated, reducing the dependence of the text emotional analysis large model on the amount of learning data, and also enabling the text emotional analysis large model to learn more accurate commodity emotions through the data set. By constructing the emotional valence and emotional arousal degree based on the Russell circular emotional model, it is possible to generate emotional label values that are more in line with human psychological emotions towards commodities, thereby making the emotional label values more accurate and reducing the dependence of the text emotional analysis large model on the amount of learning data. It realizes embedding the emotional label values based on the psychological model into a text emotional analysis large model with a simple and efficient structure, and realizes the efficient prediction of text emotions. A method for directly constructing emotional label generation that conforms to the characteristics of commodities without relying on existing word segmentation libraries is realized, which makes the data set more in line with the needs of commodity emotional classification, further reduces the dependence of the text emotional analysis large model on the amount of learning data, and also makes the text emotional analysis large model more in line with the characteristics of commodities.

[0121] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention; for those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing method for large model training, characterized in that: include: The product review text is segmented by a characteristic sentiment word combination method to generate a sentiment tag set, and the product review text is segmented by a sentiment dictionary to generate a sentiment word set; The emotion tag set and the emotion word set are sequentially subjected to an associated emotion distribution model based on a psychological model to generate a first emotion association distribution and a second emotion association distribution respectively; Calculate the comprehensive emotion intensity by weighted accumulation of the first emotion association distribution and the second emotion association distribution; The sentiment word set and the comprehensive sentiment intensity are used to learn the sentiment intensity of commodity texts in a text sentiment analysis model based on an improved convolutional neural network.

2. The data processing method for large model training according to claim 1, characterized in that: The process of respectively generating a first emotion association distribution and a second emotion association distribution by sequentially passing the emotion tag set and the emotion word set through the association emotion distribution model based on the psychological model includes: Selecting a first main sentiment word label from the sentiment label set, assigning sentiment words in the sentiment word set to a plurality of sentiment word labels, and selecting a second main sentiment word label from the plurality of sentiment word labels; The first emotion association distribution is calculated by taking the emotion tag set as a variable of the associated emotion distribution model and taking the first main emotion word tag as an expected value of the associated emotion distribution model; Using the sentiment word label as a variable of the associated sentiment distribution model, and using the second main sentiment word label as an expected value of the associated sentiment distribution model to calculate the second sentiment association distribution; Wherein, the associated sentiment distribution model is constructed based on discrete log-normal distribution.

3. The data processing method for large model training according to claim 2, characterized in that: The associated sentiment distribution model has a normalization factor for normalizing the first sentiment word set and the second sentiment word set.

4. The data processing method for large model training according to claim 2, characterized in that: The psychological model is a Russell Circle Emotion Model, the first emotion association distribution includes a first emotion valence association distribution and a first emotion arousal association distribution, and the second emotion association distribution includes a second emotion valence association distribution and a second emotion arousal association distribution.

5. The data processing method for large model training according to claim 4, characterized in that: The process of calculating the comprehensive emotion intensity by weighted accumulation of the first emotion association distribution and the second emotion association distribution includes: Accumulating multiple second emotion valence association distributions to generate a label valence, accumulating multiple second emotion valence association distributions to generate a label emotion valence, and accumulating multiple second emotion arousal association distributions to generate a label arousal; The comprehensive emotion intensity is generated by performing weighted summation based on the label emotion valence, the label arousal, the first emotion valence association distribution, and the second emotion association distribution.

6. The data processing method for large model training according to claim 1, characterized in that: The input layer, convolution pooling layer and full connection layer of the text sentiment analysis large model use the first sentiment word set and the comprehensive sentiment intensity to perform commodity text sentiment intensity learning on the text sentiment analysis large model based on the improved convolutional neural network, including: Generate sentiment word vectors for the first sentiment word set through the word embedding model of the input layer; Extracting emotional features from the emotional word vector through the convolutional pooling layer; The emotional features and the comprehensive emotional strength are passed through the improved loss function of the fully connected layer to generate text emotion classification.

7. The data processing method for large model training according to claim 6, characterized in that: The improved loss function is constructed based on the cross entropy metric according to the comprehensive sentiment strength of all tags.

8. The data processing method for large model training according to claim 7, characterized in that: A channel attention mechanism is set after the convolutional pooling layer, and the channel attention mechanism performs weighted processing on the sentiment features.

9. The data processing method for large model training according to any one of claims 1 to 8, characterized in that: The process of generating a set of sentiment labels from product review texts by combining characteristic sentiment words and segmenting them includes: Match the product review text with the product feature word library and the product sentiment word library to identify the feature words and sentiment words; Selecting emotion words that meet the characteristic emotion word combination from the characteristic words and the emotion words to generate a temporary set of emotion words; The emotion tag set is generated according to the temporary set of emotion words.

10. The data processing method for large model training according to claim 9, characterized in that: The emotion tag set includes emotion property tags and emotion intensity tags. The process of generating the emotion tag set according to the temporary set of emotion words includes: Determine the emotional property label according to the feature words corresponding to the emotional words in the temporary set of emotional words and the corresponding scores of the product review text; The sentiment intensity label is calculated according to the number of occurrences of the sentiment property label in the product review text.

Citation Information

Patent Citations

  • Training data set generation method, system and device based on annotation text and medium

    CN111859857A

  • Emotion analysis method based on improved CNN-LDA

    CN109977413A

  • Mongolian reverse reconstruction emotion distribution learning method based on semantic rules

    CN115146024A

  • Fine-grained false news detection method based on emotion distribution

    CN118410171A

  • Emotion distribution enhanced fine-grained emotion recognition method fusing VAD knowledge

    CN118535741A