Training method and sentiment classification method of sentiment classification model and device thereof
The proposed emotion classification model addresses the inaccuracy of existing methods by masking entity words and optimizing the model based on contextual semantics, improving the precision of sentiment classification for specific entities within a text.
Patent Information
- Application Number
- CN202311055437.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-08-21
AI Technical Summary
Existing fine-grained sentiment analysis techniques have problems such as error propagation, information redundancy and model optimization when dealing with the emotional tendencies of different entity words in text, especially when dealing with the emotional classification of specific entity words.
The emotion classification model is used to mask the entity words in the text, and the context semantic information is used to classify emotions. Through the training of predicting sequences and labeling sequences, the accuracy and efficiency of emotional category classification of entity words are improved.
It improves the classification accuracy and prediction efficiency of entity words, and can output the emotional categories of all entity words in the text at one time, which is suitable for the emotion analysis of specific entity words.
Smart Images

Figure CN117312978B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of natural language processing, deep learning, etc., and particularly relates to a training method and a classification method for an emotion classification model and an apparatus thereof. Background Art
[0002] With the rapid development of Internet technology, more and more users (or netizens) tend to express their opinions and views on specific entity words such as social celebrities, companies, products, etc. on social networks. Mining the attitudes and emotion tendencies of users towards specific entity words from these opinions and views, such as mining the emotion tendency of users towards a certain product, is beneficial to improving the product to better meet the needs of users. Summary of the Invention
[0003] The present disclosure provides a training method for an emotion classification model, a classification method for an emotion classification model, and an apparatus thereof.
[0004] According to one aspect of the present disclosure, there is provided a training method for an emotion classification model, including:
[0005] Obtaining a first sample text labeled with a first labeled tag sequence; wherein each element in the first labeled tag sequence is used to indicate the labeled emotion category of the corresponding first entity word in the first sample text;
[0006] Performing a masking process on each first entity word in the first sample text to obtain a masked first sample text;
[0007] Performing emotion classification on the masked first sample text by using an emotion classification model to obtain a first prediction sequence; wherein the first prediction sequence includes a first probability distribution of each first entity word, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted emotion categories;
[0008] Performing a first training on the emotion classification model according to the first prediction sequence and the first labeled tag sequence.
[0009] According to another aspect of the present disclosure, there is provided an emotion classification method, including:
[0010] Obtaining a text to be classified;
[0011] Performing a masking process on at least one target entity word in the text to be classified to obtain a masked text to be classified;
[0012] Use a sentiment classification model to perform sentiment classification on the masked text to be classified, and obtain a prediction sequence; wherein, the prediction sequence includes the probability distributions of the target entity words, and the probability distributions are used to indicate the prediction probabilities of the corresponding target entity words belonging to multiple predicted sentiment categories;
[0013] According to the prediction probabilities of the multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence, determine the target sentiment category to which each target entity word belongs from the multiple predicted sentiment categories.
[0014] According to another aspect of the present disclosure, there is provided a training device for a sentiment classification model, including:
[0015] A first acquisition module, configured to acquire a first sample text labeled with a first labeled tag sequence; wherein, each element in the first labeled tag sequence is used to indicate the labeled sentiment category of the corresponding first entity word in the first sample text;
[0016] A masking module, configured to perform masking processing on each first entity word in the first sample text to obtain a masked first sample text;
[0017] A classification module, configured to use a sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence; wherein, the first prediction sequence includes the first probability distributions of the first entity words, and the first probability distributions are used to indicate the prediction probabilities of the corresponding first entity words belonging to multiple predicted sentiment categories;
[0018] A first training module, configured to perform first training on the sentiment classification model according to the first prediction sequence and the first labeled tag sequence.
[0019] According to still another aspect of the present disclosure, there is provided a sentiment classification device, including:
[0020] A first acquisition module, configured to acquire a text to be classified;
[0021] A masking module, configured to perform masking processing on at least one target entity word in the text to be classified to obtain a masked text to be classified;
[0022] A classification module, configured to use a sentiment classification model to perform sentiment classification on the masked text to be classified to obtain a prediction sequence; wherein, the prediction sequence includes the probability distributions of the target entity words, and the probability distributions are used to indicate the prediction probabilities of the corresponding target entity words belonging to multiple predicted sentiment categories;
[0023] A determination module, configured to determine, according to the prediction probabilities of multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence, the target sentiment category to which each of the target entity words belongs from the multiple predicted sentiment categories.
[0024] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0025] At least one processor; and
[0026] A memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the sentiment classification model proposed in the above aspect of the present disclosure, or execute the sentiment classification method proposed in the other aspect of the present disclosure above.
[0028] According to yet another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the sentiment classification model proposed in the above aspect of the present disclosure, or execute the sentiment classification method proposed in the other aspect of the present disclosure above.
[0029] According to still another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the training method of the sentiment classification model proposed in the above aspect of the present disclosure, or implements the sentiment classification method proposed in the other aspect of the present disclosure above when executed.
[0030] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0032] Figure 1 is a schematic flowchart of the training method of the sentiment classification model provided in Embodiment 1 of the present disclosure;
[0033] Figure 2 is a schematic flowchart of the training method of the sentiment classification model provided in Embodiment 2 of the present disclosure;
[0034] Figure 3 is a schematic flowchart of the training method of the sentiment classification model provided in Embodiment 3 of the present disclosure;
[0035] Figure 4Schematic flowchart of the training method for the sentiment classification model provided in the fourth embodiment of the present disclosure;
[0036] Figure 5 Schematic flowchart of the training method for the sentiment classification model provided in the fifth embodiment of the present disclosure;
[0037] Figure 6 Schematic flowchart of the training method for the sentiment classification model provided in the sixth embodiment of the present disclosure;
[0038] Figure 7 Schematic flowchart of the acquisition process of the first training set provided in the seventh embodiment of the present disclosure;
[0039] Figure 8 Schematic flowchart of the acquisition process of the first training set provided in the eighth embodiment of the present disclosure;
[0040] Figure 9 Schematic flowchart of the sentiment classification method provided in the ninth embodiment of the present disclosure;
[0041] Figure 10 Schematic flowchart of the sentiment classification method provided in the tenth embodiment of the present disclosure;
[0042] Figure 11 Schematic diagram of the structure of the sentiment classification model provided in the embodiment of the present disclosure;
[0043] Figure 12 Schematic diagram of the structure of the training device for the sentiment classification model provided in the eleventh embodiment of the present disclosure;
[0044] Figure 13 Schematic diagram of the structure of the sentiment classification device provided in the twelfth embodiment of the present disclosure;
[0045] Figure 14 Schematic block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0046] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0047] Traditional sentiment analysis techniques perform coarse-grained sentiment classification on text to obtain the sentiment tendency expressed by the entire text. However, a text may contain multiple entity words, and the sentiment tendencies contained in different entity words may be inconsistent. For example, assuming the text is "Zhang in Group X is quite serious in doing things, while Li often forgets things", it can be seen that the sentiment tendency towards the entity word "Zhang" in this text is positive, while the sentiment tendency towards the entity word "Li" is negative. Therefore, the coarse-grained sentiment analysis technique cannot accurately judge the sentiment tendencies of each entity word in the text.
[0048] Through fine-grained sentiment analysis techniques, the sentiment tendencies of different entity words in the text can be identified, thus obtaining more accurate and comprehensive sentiment analysis conclusions.
[0049] Among them, fine-grained sentiment analysis mainly includes three elements: attribute / entity word, opinion, and sentiment tendency. Its subtasks include opinion extraction, attribute / entity word extraction, sentiment polarity judgment, etc. Currently, the research on fine-grained sentiment analysis techniques mainly adopts the deep learning method, and the commonly used methods are mainly divided into the following several types:
[0050] The first type is the pipeline method, which splits the fine-grained sentiment analysis technique into subtasks such as entity word recognition, opinion recognition, opinion category, and sentiment classification.
[0051] The second type is the joint learning method, which models multiple subtasks of the fine-grained sentiment analysis technique using the same model. Commonly used modeling methods include sequence annotation methods, machine reading comprehension methods, and generative methods, etc.
[0052] However, the above first method has at least the following disadvantages:
[0053] 1. The pipeline-based method splits complex problems. After processing one task and then processing the next task, it is easy to cause error propagation. For example, errors in the entity word recognition module will affect the results of the sentiment recognition module.
[0054] 2. The pipeline method ignores the relevance between tasks. For example, entity words and opinions often appear together. If the opinion is known, then the described entity word can also be judged, but the pipeline method obviously cannot utilize this information.
[0055] 3. The pipeline method is prone to information redundancy. Since it is necessary to perform opinion extraction on all recognized entity words and perform sentiment classification on all extracted opinions, some invalid matching pairs will be generated, reducing the prediction efficiency.
[0056] The second method mentioned above has at least the following disadvantages: The method based on joint learning uses the same model to model multiple subtasks, and it is difficult to select a suitable model to transform the task form; compared with the pipeline method, it is difficult to analyze and optimize individual modules in the joint modeling method.
[0057] In addition, the current fine-grained sentiment analysis technology is mainly applied to the analysis of product reviews, while in the present disclosure, it is the fine-grained sentiment analysis of specific entity words, and the fields are different.
[0058] In view of at least one of the above problems, the present disclosure provides a training method of a sentiment classification model, a sentiment classification method and an apparatus thereof.
[0059] The following describes the training method of the sentiment classification model, the sentiment classification method and the apparatus thereof according to the embodiments of the present disclosure with reference to the accompanying drawings.
[0060] Figure 1 It is a schematic flowchart of the training method of the sentiment classification model provided by Embodiment 1 of the present disclosure.
[0061] In the embodiments of the present disclosure, the training method of the sentiment classification model is configured in a training apparatus of the sentiment classification model for illustration. The training apparatus of the sentiment classification model can be applied to any electronic device so that the electronic device can perform the training function of the sentiment classification model.
[0062] Among them, the electronic device can be any device with computing power. For example, it can be a personal computer, a mobile terminal, a server, etc. The mobile terminal can be a hardware device such as a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. with various operating systems, touch screens and / or display screens.
[0063] As Figure 1 shown, the training method of the sentiment classification model may include the following steps:
[0064] Step S101, obtaining a first sample text labeled with a first labeled tag sequence; wherein, each element in the first labeled tag sequence is used to indicate the labeled sentiment category of the corresponding first entity word in the first sample text.
[0065] Among them, the number of labeled sentiment categories is not limited. For example, the number of labeled sentiment categories can be 3, namely: positive, neutral and negative. For another example, the number of labeled sentiment categories can be 5, namely: positive, neutral, mildly negative, moderately negative and severely negative. That is to say, in the present disclosure, the division granularity of the sentiment category (or sentiment tendency) to which the entity word belongs is not limited, and each sentiment category can be divided according to actual application requirements.
[0066] In the embodiments of the present disclosure, there is no restriction on the way to obtain the first sample text. For example, the first sample text can be text obtained from an existing training set, or the first sample text can be text collected online. For example, through web crawler technology, the first sample text (such as speech information or opinion information in a social network) can be collected online, or the first sample text can also be text provided manually, etc. The embodiments of the present disclosure do not limit this.
[0067] In the embodiments of the present disclosure, the first sample text is labeled with a first labeled tag sequence. Among them, the labeled tag sequence includes the labeled tags of each entity word (denoted as the first entity word in the present disclosure) in the first sample text. Among them, the labeled tag is used to indicate the labeled sentiment category to which the corresponding first entity word belongs.
[0068] As an example, assuming that the number of labeled sentiment categories is 3, namely positive, neutral, and negative, the labeled tag corresponding to positive can be +1, the labeled tag corresponding to neutral can be 0, and the labeled tag corresponding to negative can be -1; or, the labeled tag corresponding to positive can be 2, the labeled tag corresponding to neutral can be 1, and the labeled tag corresponding to negative can be 0, etc. There is no need to list them one by one here.
[0069] Step S102: Perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0070] In the embodiments of the present disclosure, masking processing can be performed on each first entity word in the first sample text to obtain the masked first sample text.
[0071] As an example, fixed masking characters can be used to perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0072] As another example, random masking characters can be used to perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0073] As still another example, masking labels (such as "Mask" labels) can be used to perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0074] Step S103: Use the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence.
[0075] In the embodiments of the present disclosure, an emotion classification model may be used to perform emotion classification on the masked first sample text to obtain a first prediction sequence. The first prediction sequence may include a first probability distribution of each first entity word in the first sample text, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted emotion categories.
[0076] Wherein, the total number of predicted emotion categories is consistent with the total number of labeled emotion categories.
[0077] Taking the number of labeled emotion categories and predicted emotion categories as 3, namely positive, neutral, and negative respectively, as an example for illustration. Assume that the first sample text includes m first entity words, where m is a positive integer. Then the output of the emotion classification model may be {a1, a2, …, a i , …, a m}, where i is a positive integer not greater than m, and a i represents the first probability distribution of the i-th first entity word in the first sample text. a i = {p ipos , p ineg , p ineu}, where p ipos represents the prediction probability that the emotion category (or emotion tendency) of the i-th first entity word is positive, p ineg represents the prediction probability that the emotion category (or emotion tendency) of the i-th first entity word is negative, and p ineu represents the prediction probability that the emotion category (or emotion tendency) of the i-th first entity word is neutral.
[0078] Step S104, perform a first training on the emotion classification model according to the first prediction sequence and the first labeled tag sequence.
[0079] In the embodiments of the present disclosure, the emotion classification model may be trained according to the first prediction sequence and the first labeled tag sequence (denoted as the first training in the present disclosure).
[0080] The training method of the sentiment classification model according to the embodiments of the present disclosure masks each first entity word in the first sample text to obtain the masked first sample text; uses the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence; and performs first training on the sentiment classification model according to the first prediction sequence and the first annotation label sequence annotated by the first sample text. Thus, when using the sentiment classification model to perform sentiment classification on the masked first sample text, the output result of the sentiment classification model does not depend on specific first entity words, but only depends on the context semantic information of the first entity words, which can improve the accuracy of classifying the sentiment categories to which the first entity words belong. Moreover, the sentiment classification model can output the sentiment categories to which all the first entity words in the first sample text belong at one time, which can improve the prediction efficiency.
[0081] It should be noted that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0082] To clearly illustrate how the sentiment classification model in the above embodiments performs sentiment classification on the masked first sample text, the present disclosure also proposes a training method for the sentiment classification model.
[0083] Figure 2 It is a schematic flowchart of the training method for the sentiment classification model provided in the second embodiment of the present disclosure.
[0084] As Figure 2 shown, the training method of the sentiment classification model may include the following steps:
[0085] Step S201, obtain the first sample text annotated with the first annotation label sequence.
[0086] Each element in the first annotation label sequence is used to indicate the annotated sentiment category of the corresponding first entity word in the first sample text.
[0087] For the explanatory description of step S201, reference can be made to the relevant descriptions in any embodiment of the present disclosure, and details are not described herein again.
[0088] Step S202, according to the mask label, mask each first entity word in the first sample text to obtain the masked first sample text.
[0089] In the embodiments of the present disclosure, the mask label can be used to mask each first entity word in the first sample text to obtain the masked first sample text.
[0090] As an example, assume that the first sample text X includes n characters x, such as X = {x1, x2, …, x n}, and there are m first entity words in X, such as X includes E = {e1, e2, …, e m}, e i is the i-th first entity word in X, where e i = {x i1 , x i2 , …, x ik}, that is, e i includes k characters (that is, each first entity word is one or more characters in X), where k is less than n, and m is also less than n. Then, each first entity word e i in X can be replaced with the "[MASK]" symbol to obtain the masked first sample text X′ = {x1, …, [MASK], …, [MASK], …, x n}.
[0091] Step S203, input the masked first sample text into the encoder in the sentiment classification model, so as to use the encoder to encode each masked label based on the context information of each masked label, and obtain the encoded features of each masked label.
[0092] In the embodiments of the present disclosure, the masked first sample text can be input into the encoder in the sentiment classification model, so as to use the encoder to encode each masked label based on the context information (or context semantic information) of each masked label in the masked first sample text, and obtain the encoded features of each encoded label.
[0093] Step S204, use the decoder in the sentiment classification model to predict the sentiment category of each first entity word based on the encoded features of each masked label, so as to obtain a first prediction sequence.
[0094] In the embodiments of the present disclosure, the decoder in the sentiment classification model can be used to predict the sentiment category of each first entity word based on the encoded features of each masked label, and obtain a first prediction sequence, where the first prediction sequence includes the first probability distribution of each first entity word, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted sentiment categories.
[0095] Step S205, perform the first training on the sentiment classification model according to the first prediction sequence and the first labeled label sequence.
[0096] For the explanation of step S205, reference can be made to the relevant descriptions in any embodiment of the present disclosure, and details are not described herein again.
[0097] The training method of the sentiment classification model according to the embodiments of the present disclosure can predict the sentiment categories of each first entity word based on the context semantic information of each first entity word in the first sample text, which can improve the accuracy of the prediction results.
[0098] To clearly illustrate how the above embodiments perform the first training on the sentiment classification model according to the first prediction sequence and the first labeled tag sequence, the present disclosure also proposes a training method for the sentiment classification model.
[0099] Figure 3 It is a schematic flowchart of the training method for the sentiment classification model provided in the third embodiment of the present disclosure.
[0100] As Figure 3 shown, the training method of the sentiment classification model may include the following steps:
[0101] Step S301, obtain the first sample text labeled with the first labeled tag sequence.
[0102] Among them, each element in the first labeled tag sequence is used to indicate the labeled sentiment category of the corresponding first entity word in the first sample text.
[0103] Step S302, perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0104] For the explanatory description of steps S301 to S302, reference can be made to the relevant descriptions in any embodiment of the present disclosure, which will not be elaborated here.
[0105] It should be noted that the first sample text may contain some entity words that are irrelevant to the target task adapted to the sentiment classification model. For example, if the target task is to identify the sentiment category of a specific product, the first sample text may contain the specific product, and may also contain entity words such as place names and personal names. At this time, in order to improve the prediction efficiency of the sentiment classification model, sentiment classification can be performed only on the first entity words related to the target task.
[0106] Therefore, in any embodiment of the present disclosure, a list of entity words associated with the target task can be obtained, and each entity word in the first sample text that is located in the list of entity words can be used as the first entity word. Thus, in the present disclosure, masking processing can be performed only on each first entity word in the first sample text to obtain the masked first sample text.
[0107] Among them, the list of entity words contains each entity word associated with the target task, and the target task is used to indicate the sentiment category of the entity word that the sentiment classification model needs to predict.
[0108] Accordingly, the sentiment classification model only needs to perform sentiment classification on the first entity words associated with the target task, without performing sentiment classification on other entity words irrelevant to the target task, which can further improve the prediction efficiency of the sentiment classification model.
[0109] Step S303: Use the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence.
[0110] The first prediction sequence includes the first probability distributions of the respective first entity words, and the first probability distribution is used to indicate the prediction probabilities of the corresponding first entity words belonging to multiple predicted sentiment categories.
[0111] For the explanation of step S303, reference can be made to the relevant descriptions in any embodiment of the present disclosure, and details are not elaborated herein.
[0112] Step S304: For the i-th first probability distribution in the first prediction sequence, generate a classification loss value for the i-th first entity word in the first sample text according to the prediction probabilities of the multiple predicted sentiment categories indicated by the i-th first probability distribution and the labeled sentiment category indicated by the i-th element in the first labeled label sequence.
[0113] In the embodiments of the present disclosure, for the i-th (i is a positive integer) first probability distribution in the first prediction sequence, the value of the classification loss function of the i-th first entity word in the first sample text can be determined according to the prediction probabilities of the multiple predicted sentiment categories indicated by the i-th first probability distribution and the labeled sentiment category indicated by the i-th element in the first labeled label sequence. In the present disclosure, it is denoted as the classification loss value.
[0114] As an example, taking the classification loss function as the cross-entropy loss function for example, the classification loss value of the i-th first entity word in the first sample text can be:
[0115] where M is the number of predicted sentiment categories, y ic represents the sign function. If the labeled sentiment category of the i-th first entity word is equal to the predicted sentiment category c, then y ic takes 1; otherwise, y ic takes 0, and p ic represents the prediction probability that the i-th first entity word belongs to the predicted sentiment category c.
[0116] Step S305: Determine the target loss value according to the classification loss values of the respective first entity words.
[0117] In the embodiments of the present disclosure, the target loss value can be determined according to the classification loss values of the respective first entity words.
[0118] As an example, the sum, mean, and weighted sum value of the classification loss values of each first entity word can be used as the target loss value.
[0119] Still taking the classification loss function as the cross-entropy loss function as an example, the target loss value L can be, for example:
[0120]
[0121] Where m is the number of first entity words.
[0122] Step S306, perform the first training on the sentiment classification model according to the target loss value.
[0123] In the embodiments of the present disclosure, the sentiment classification model can be trained according to the target loss value. For example, the sentiment classification model can be trained according to the target loss value to minimize the target loss value.
[0124] It should be noted that the above only takes the termination condition of model training as minimizing the target loss value as an example. In actual applications, other termination conditions can also be set. For example, the termination condition can also be that the number of training times reaches a set number, or the termination condition can also be that the training duration reaches a set duration, etc. The present disclosure does not limit this.
[0125] The training method of the sentiment classification model in the embodiments of the present disclosure can train the sentiment classification model according to the classification loss values of all the first entity words, and can improve the prediction accuracy of the sentiment classification model.
[0126] To clearly illustrate how the above embodiments train the sentiment classification model according to the first prediction sequence and the first labeled tag sequence, the present disclosure also proposes a training method for the sentiment classification model.
[0127] Figure 4 It is a schematic flowchart of the training method for the sentiment classification model provided in the fourth embodiment of the present disclosure.
[0128] As Figure 4 shown, the training method of the sentiment classification model can include the following steps:
[0129] Step S401, obtain the first sample text labeled with the first labeled tag sequence.
[0130] Among them, each element in the first labeled tag sequence is used to indicate the labeled sentiment category of the corresponding first entity word in the first sample text.
[0131] Step S402, perform masking processing on each first entity word in the first sample text to obtain the masked first sample text.
[0132] Step S403: Use the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence.
[0133] Among them, the first prediction sequence includes the first probability distribution of each first entity word, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted sentiment categories.
[0134] For the explanations of steps S401 to S403, reference can be made to the relevant descriptions in any embodiment of the present disclosure, which will not be elaborated here.
[0135] Step S404: According to the prediction probabilities of the multiple predicted sentiment categories indicated by the first probability distributions in the first prediction sequence, determine the target sentiment category to which each first entity word belongs from the multiple predicted sentiment categories.
[0136] In the embodiments of the present disclosure, for any one of the first probability distributions in the first prediction sequence, the target sentiment category to which the first entity word corresponding to the first probability distribution belongs can be determined from the multiple predicted sentiment categories according to the prediction probabilities of the multiple predicted sentiment categories indicated by the first probability distribution. For example, the prediction probability of the target sentiment category in the first probability distribution is greater than the prediction probabilities of other predicted sentiment categories.
[0137] Step S405: Generate a prediction label sequence according to the target sentiment category to which each first entity word belongs.
[0138] In the embodiments of the present disclosure, a prediction label sequence can be generated according to the target sentiment category to which each first entity word belongs. Among them, the prediction label sequence includes the prediction labels of each first entity word, and the prediction label is used to indicate the target sentiment category to which the first entity word belongs.
[0139] As an example, assume that the number of target sentiment categories is 3, namely positive, neutral, and negative. Then the prediction label corresponding to positive can be +1, the prediction label corresponding to neutral can be 0, and the prediction label corresponding to negative can be -1; or, the prediction label corresponding to positive can be 2, the prediction label corresponding to neutral can be 1, and the prediction label corresponding to negative can be 0, etc., which will not be listed one by one here.
[0140] It should be noted that the first annotation label sequence can be obtained by sorting the annotation labels of each first entity word according to the positions of the first entity words in the first sample text. Similarly, the prediction label sequence can also be obtained by sorting the prediction labels of each first entity word according to the positions of the first entity words in the first sample text.
[0141] Step S406: Perform the first training on the sentiment classification model according to the difference between the predicted label sequence and the first labeled label sequence.
[0142] In the embodiments of the present disclosure, the first training on the sentiment classification model can be performed according to the difference between the predicted label sequence and the first labeled label sequence.
[0143] For example, the loss value can be generated according to the above difference, where the loss value has a positive correlation with the above difference, that is, the smaller the difference, the smaller the loss value, and vice versa, the larger the difference, the larger the loss value. Therefore, in the present disclosure, the first training on the sentiment classification model can be performed according to the loss value to minimize the loss value.
[0144] It should be noted that the above is only an example with the termination condition of model training being the minimization of the loss value. In actual applications, other termination conditions can also be set. For example, the termination condition can also be that the number of training times reaches a set number, or the termination condition can also be that the training duration reaches a set duration, etc. The present disclosure does not limit this.
[0145] The training method of the sentiment classification model in the embodiments of the present disclosure can realize training the sentiment classification model in different ways, improving the applicability and flexibility of the method.
[0146] In any one of the embodiments of the present disclosure, in order to further improve the prediction accuracy of the sentiment classification model, before training the sentiment classification model, the sentiment classification model can also be pre-trained. The following combines Figure 5 , and details the pre-training process of the sentiment classification model.
[0147] Figure 5 It is a schematic flowchart of the training method of the sentiment classification model provided in the fifth embodiment of the present disclosure.
[0148] As Figure 5 shown, the sentiment classification model can be pre-trained through the following steps:
[0149] Step S501: Obtain the second sample text.
[0150] Among them, there is no limitation on the way to obtain the second sample text. For example, the second sample text can be obtained by a method similar to that of the first sample text.
[0151] Step S502: Mask at least one target text segment in the second sample text to obtain the masked second sample text.
[0152] Among them, the target text segment includes at least one of a specified entity word (pre-specified by relevant personnel), a noun, and a proper noun.
[0153] In an embodiment of the present disclosure, at least one target text segment in the second sample text may be masked to obtain the masked second sample text. The implementation principle is similar to the masking method of the first sample text and will not be elaborated here.
[0154] Step S503: Use the initial sentiment classification model to perform character prediction on the masked second sample text to obtain a predicted text.
[0155] In an embodiment of the present disclosure, the initial sentiment classification model may be used to perform character prediction on the masked second sample text to obtain a predicted text.
[0156] In a possible implementation manner of an embodiment of the present disclosure, the sentiment classification model may only predict the masked (MASKed) characters, which is similar to the task of cloze test, that is, the number of predicted texts output by the sentiment classification model is the same as the number of target text segments.
[0157] In another possible implementation manner of an embodiment of the present disclosure, the sentiment classification model may predict all characters in the entire text in a manner similar to machine translation to obtain a predicted text, that is, the number of predicted texts output by the sentiment classification model is one.
[0158] Step S504: Perform first pre-training on the sentiment classification model according to the predicted text.
[0159] In an embodiment of the present disclosure, the sentiment classification model may be first pre-trained according to the predicted text.
[0160] In a possible implementation manner of an embodiment of the present disclosure, when the number of predicted texts output by the sentiment classification model is the same as the number of target text segments, the sentiment classification model may be first pre-trained according to the difference between each target text segment and each predicted text.
[0161] For example, a loss value may be generated according to the above difference, where the loss value is positively correlated with the difference. Therefore, in the present disclosure, the sentiment classification model may be first pre-trained according to the loss value to minimize the loss value.
[0162] In a possible implementation manner of an embodiment of the present disclosure, when the number of predicted texts output by the sentiment classification model is one, the sentiment classification model may be first pre-trained according to the difference between the predicted text and the second sample text.
[0163] For example, a loss value may be generated according to the above difference, where the loss value is positively correlated with the difference. Therefore, in the present disclosure, the sentiment classification model may be first pre-trained according to the loss value to minimize the loss value.
[0164] It should be noted that the above only takes the minimization of the loss value as the termination condition for model pre-training as an example. In actual applications, other termination conditions can also be set. For example, the termination condition can also be that the number of training times reaches a set number, or the termination condition can also be that the training duration reaches a set duration, etc. The present disclosure does not limit this.
[0165] Thus, according to different methods, the sentiment classification model can be pre-trained, which can improve the flexibility and applicability of this method.
[0166] In the training method of the sentiment classification model according to the embodiments of the present disclosure, before formally training the sentiment classification model, pre-training the sentiment classification model can not only improve the prediction accuracy of the sentiment classification model, but also shorten the training duration of the sentiment classification model.
[0167] In any one of the embodiments of the present disclosure, in order to further improve the prediction accuracy of the sentiment classification model, before training the sentiment classification model, the sentiment classification model can also be pre-trained. The following combines Figure 6 to elaborate in detail on the pre-training process of the sentiment classification model.
[0168] Figure 6 FIG. is a schematic flow chart of the training method of the sentiment classification model provided in Embodiment VI of the present disclosure.
[0169] As Figure 6 shown, the sentiment classification model can be pre-trained through the following steps:
[0170] Step S601, obtain at least one third sample text from at least one dataset irrelevant to the target task adapted to the sentiment classification model.
[0171] Among them, the above dataset is an open-source dataset, such as the in-store review dataset of a catering APP (Application), the fine-grained sentiment analysis dataset of a news or search APP, etc. And, each text in this dataset contains some entity words irrelevant to the target task adapted to the sentiment classification model.
[0172] In the embodiments of the present disclosure, at least one third sample text can be obtained from at least one dataset irrelevant to the target task.
[0173] Step S602, generate a second labeled tag sequence according to the labeled sentiment categories labeled by at least one second entity word in the third sample text.
[0174] In an embodiment of the present disclosure, for any third sample text, a second labeled tag sequence can be generated according to the labeled sentiment categories of each entity word in the third sample text (denoted as the second entity word in the present disclosure, where the second entity word may not be in the entity word list associated with the target task). The second labeled tag sequence includes the labeled tags of each second entity word, and the labeled tag is used to indicate the labeled sentiment category to which the corresponding second entity word belongs.
[0175] Step S603: Use the sentiment classification model to perform sentiment classification on the third sample text to obtain a second prediction sequence.
[0176] Among them, the sentiment classification model can be an initial sentiment classification model, or can also be a sentiment classification model pre-trained for the first time. The embodiments of the present disclosure do not limit this.
[0177] In an embodiment of the present disclosure, the sentiment classification model can be used to perform sentiment classification on the third sample text to obtain a second prediction sequence. The second prediction sequence includes the second probability distribution of each second entity word, and the second probability distribution is used to indicate the prediction probability that the corresponding second entity word belongs to multiple predicted sentiment categories.
[0178] As a possible implementation, a sentiment classification method similar to that of the first sample text can be used to perform sentiment classification on the third sample text to obtain a second prediction sequence.
[0179] For example, each second entity word in the third sample text can be masked to obtain a masked third sample text, and the sentiment classification model is used to perform sentiment classification on the masked third sample text to obtain a second prediction sequence. Its implementation is similar to steps S102 to S103, or its implementation is similar to steps S202 to S204, and will not be elaborated here.
[0180] Thus, when using the sentiment classification model to perform sentiment classification on the masked third sample text, the output result of the sentiment classification model does not depend on the specific second entity word, but only depends on the context semantic information of the second entity word, which can improve the accuracy of classifying the sentiment category to which the second entity word belongs, that is, improve the accuracy of the second prediction sequence output by the model.
[0181] Step S604: Perform second pre-training on the sentiment classification model according to the second labeled tag sequence and the second prediction sequence.
[0182] In an embodiment of the present disclosure, the sentiment classification model can be pre-trained for the second time according to the second annotation tag sequence and the second prediction sequence. The implementation principle is similar to that of step S104, or the implementation principle is similar to that of steps S304 to S306, or the implementation principle is similar to that of steps S404 to S406, which will not be elaborated here.
[0183] In the training method of the sentiment classification model according to the embodiment of the present disclosure, before the sentiment classification model is formally trained, the sentiment classification model is pre-trained with sample texts irrelevant to the target task adapted to the sentiment classification model, which can enable the pre-trained sentiment classification model to have preliminary sentiment classification capabilities. Therefore, when the pre-trained sentiment classification model is formally trained, not only can the prediction accuracy of the sentiment classification model be improved, but also the training duration of the sentiment classification model can be shortened.
[0184] In any embodiment of the present disclosure, the first sample text can be a text obtained from the first training set, where each text in the first training set can be generated according to a text associated with the target task adapted to the sentiment classification model. The following will be combined with Figure 7 , and the generation process of the first training set will be described in detail.
[0185] Figure 7 It is a schematic diagram of the acquisition process of the first training set provided in Embodiment 7 of the present disclosure.
[0186] As Figure 7 shown, the first training set can be obtained by the following steps:
[0187] Step S701: Obtain a plurality of unannotated fourth sample texts associated with the target task adapted to the sentiment classification model.
[0188] In an embodiment of the present disclosure, a plurality of unannotated fourth sample texts associated with the target task can be obtained, where the entity words in the fourth sample texts (denoted as the fourth entity words in the present disclosure) can be in the entity word list associated with the target task.
[0189] In an embodiment of the present disclosure, the fourth sample texts associated with the target task can be collected (such as online collection, offline collection).
[0190] In any embodiment of the present disclosure, in order to improve the effectiveness and accuracy of the text prediction of the sentiment classification model, after the fourth sample texts are obtained, the fourth sample texts can also be preprocessed according to the input requirements of the sentiment classification model, where the preprocessing includes at least one of the following: font conversion, font format conversion.
[0191] As an example, according to the input requirements, perform traditional / simplified conversion, full-width / half-width conversion, and convert capital letters to lowercase on the fourth sample text.
[0192] Thus, it is possible to make the font and font format of the preprocessed fourth sample text match the input requirements of the sentiment classification model, improving the effectiveness and accuracy of the text prediction by the sentiment classification model.
[0193] In any one of the embodiments of the present disclosure, in order to improve the accuracy of the text prediction by the sentiment classification model, multiple fourth sample texts can also be screened, and among them, the third entity words included in the retained fourth sample texts are not repeated.
[0194] Thus, it is possible to make the entity words included in each sample text input into the sentiment classification model non-repetitive, so as to improve the classification accuracy or prediction accuracy of the sentiment classification model for the same entity word.
[0195] In any one of the embodiments of the present disclosure, in order to improve the effectiveness of the text prediction by the sentiment classification model, after obtaining the fourth sample text, it can also be determined whether the number of characters included in the fourth sample text is less than or equal to the set number threshold (such as 512) in the above input requirements. If the number of characters included in the fourth sample text is less than or equal to the set number threshold, subsequent steps can be executed.
[0196] If the number of characters included in the fourth sample text is greater than the set number threshold, the fourth sample text can be segmented so that the number of characters included in the segmented fourth sample text is less than or equal to the set number threshold.
[0197] As an example, taking the third entity word in the fourth sample text as the center, a set number threshold of characters can be intercepted to obtain the segmented fourth sample text.
[0198] Thus, it is possible to make the number of characters included in the sample text input into the sentiment classification model match the input requirements of the sentiment classification model, improving the effectiveness and accuracy of the text prediction by the sentiment classification model.
[0199] Step S702, for any fourth sample text, use the sentiment classification model pre-trained for the second time to perform sentiment classification on any fourth sample text to obtain the third prediction sequence of any fourth sample text.
[0200] In an embodiment of the present disclosure, for any fourth sample text, a sentiment classification model pre-trained for the second time can be used to perform sentiment classification on the fourth sample text to obtain a third prediction sequence of the fourth sample text, where the third prediction sequence includes a third probability distribution of each third entity word in the fourth sample text, and the third probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to multiple predicted sentiment categories.
[0201] As a possible implementation, a sentiment classification method similar to that of the first sample text can be used to perform sentiment classification on the fourth sample text to obtain a third prediction sequence.
[0202] For example, each third entity word (which can be in the entity word list associated with the target task) in the fourth sample text can be masked to obtain the masked fourth sample text, and a sentiment classification model pre-trained for the second time can be used to perform sentiment classification on the masked fourth sample text to obtain a third prediction sequence. Its implementation is similar to steps S102 to S103, or its implementation is similar to steps S202 to S204, which will not be elaborated here.
[0203] Thus, when using a sentiment classification model pre-trained for the second time to perform sentiment classification on the masked fourth sample text, the output result of the sentiment classification model pre-trained for the second time does not depend on specific third entity words, but only on the context semantic information of the third entity words, which can improve the accuracy of classifying the sentiment category to which the third entity word belongs, that is, improve the accuracy of the third prediction sequence output by the model.
[0204] Step S703, label any fourth sample text according to the third prediction sequence to obtain the labeled any fourth sample text.
[0205] In an embodiment of the present disclosure, the fourth sample text can be labeled according to the third prediction sequence to obtain the labeled fourth sample text. For example, each annotation label in the annotation label sequence annotated on the fourth sample text can be used to indicate the sentiment category to which the corresponding third entity word belongs, where the sentiment category to which the third entity word belongs can be the predicted sentiment category corresponding to the maximum predicted probability indicated by the third probability distribution of the third entity word.
[0206] Step S704, generate a first training set according to multiple labeled fourth sample texts.
[0207] In an embodiment of the present disclosure, a first training set can be generated according to multiple labeled fourth sample texts.
[0208] In summary, the sentiment classification model after the second pre-training already has a preliminary sentiment recognition ability. Based on this sentiment recognition model after the second pre-training, the annotation data of each unannotated fourth sample text is output, without the need for manual annotation of all fourth sample texts, which can reduce the annotation cost.
[0209] To clearly illustrate how the first training set is generated according to multiple annotated fourth sample texts in the above embodiments, the present disclosure also proposes a method for obtaining the first training set.
[0210] Figure 8 It is a schematic flowchart of obtaining the first training set provided in Embodiment VIII of the present disclosure.
[0211] As Figure 8 shown, the first training set can be obtained by the following steps:
[0212] Step S801: Obtain multiple unannotated fourth sample texts associated with the target task adapted to the sentiment classification model.
[0213] Step S802: For any fourth sample text, use the sentiment classification model after the second pre-training to perform sentiment classification on any fourth sample text to obtain the third prediction sequence of any fourth sample text.
[0214] Among them, the third prediction sequence includes the third probability distribution of each third entity word in any fourth sample text, and the third probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to multiple predicted sentiment categories.
[0215] Step S803: Annotate any fourth sample text according to the third prediction sequence to obtain the annotated any fourth sample text.
[0216] For the explanatory notes of steps S801 to S803, reference can be made to the relevant descriptions in any embodiment of the present disclosure, and details are not described herein again.
[0217] Step S804: Determine the target sample text with incorrect annotation from multiple annotated fourth sample texts.
[0218] It should be noted that the sentiment classification model after the second pre-training is pre-trained using sample texts irrelevant to the target task. When using this sentiment classification model after the second pre-training to perform sentiment classification on sample texts associated with the target task, there may be a situation where some sample texts are not accurately classified (i.e., the prediction results of the model are inaccurate). Therefore, in the embodiments of the present disclosure, multiple annotated fourth sample texts can also be cleaned. For example, the target sample text with incorrect annotation can be determined from multiple annotated fourth sample texts, so as to use subsequent steps to update the target sample text.
[0219] In any one of the embodiments of the present disclosure, the method for determining the target sample text may be, for example:
[0220] For example, first, multiple labeled fourth sample texts may be divided into K groups, and one group may be selected from the K groups in sequence as the test set.
[0221] After that, for any selected test set, the remaining groups of the K groups except the selected test set may be used as the second training set, and the sentiment classification model may be second-trained using the second training set. Then, based on the sentiment classification model after the second training, the selected test set may be sentiment-classified to obtain the fourth prediction sequences of the labeled fourth sample texts in the selected test set.
[0222] Among them, the fourth prediction sequence may include the fourth probability distribution of each third entity word in the labeled fourth sample text, and the fourth probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to multiple predicted sentiment categories.
[0223] Finally, based on the fourth prediction sequences and the third prediction sequences of the labeled fourth sample texts in the selected test set, the target sample text may be determined from the selected test set, where the fourth prediction sequence and the third prediction sequence of the target sample text are inconsistent.
[0224] Thus, the K-fold cross-validation method can be used to determine the target sample text with incorrect labels from multiple labeled fourth sample texts, improving the accuracy and effectiveness of determining the sample data with incorrect labels.
[0225] Step S805: In response to an update operation on the target sample text, update the annotation data of the target sample text.
[0226] In the embodiments of the present disclosure, in response to an update operation on the target sample text triggered by relevant personnel, the annotation data (i.e., the annotation label sequence) of the target sample text may be updated. That is, the target sample text with incorrect labels may be re-annotated through manual annotation.
[0227] Step S806: Generate a first training set according to the updated target sample text and the remaining texts among the multiple labeled fourth sample texts except the target sample text.
[0228] In the embodiments of the present disclosure, a first training set may be generated according to the updated target sample text and the remaining texts among the multiple labeled fourth sample texts except the target sample text.
[0229] The training method of the sentiment classification model according to the embodiments of the present disclosure can update the annotation data of the target sample text with incorrect annotations, which can improve the training effect of the sentiment classification model, that is, improve the prediction accuracy of the sentiment classification model.
[0230] The above are the embodiments corresponding to the training method of the sentiment classification model. The present disclosure also proposes an application method of the sentiment classification model (i.e., the sentiment classification method).
[0231] Figure 9 It is a schematic flowchart of the sentiment classification method provided in Embodiment IX of the present disclosure.
[0232] As Figure 9 shown, the sentiment classification method may include the following steps:
[0233] Step S901, obtain the text to be classified.
[0234] In the embodiments of the present disclosure, the acquisition method of the text to be classified is not limited. For example, the text to be classified may be the text obtained from the existing test set, or the text to be classified may be the text collected online. For example, the text to be classified can be collected online through web crawler technology, or the text to be classified can also be the text provided manually, etc. The embodiments of the present disclosure do not limit this.
[0235] Step S902, perform masking processing on at least one target entity word in the text to be classified to obtain the masked text to be classified.
[0236] In the embodiments of the present disclosure, at least one target entity word in the text to be classified can be masked to obtain the masked text to be classified. The implementation principle is similar to the masking method of the first sample text and will not be elaborated here.
[0237] Step S903, use the sentiment classification model to perform sentiment classification on the masked text to be classified to obtain a prediction sequence.
[0238] Among them, the sentiment classification model is trained by using the training method of the sentiment classification model in any of the foregoing embodiments. It should be noted that the explanations of the training method of the sentiment classification model in any of the foregoing embodiments also apply to this embodiment, and the implementation principle is similar and will not be elaborated here.
[0239] In the embodiments of the present disclosure, the trained sentiment classification model can be used to perform sentiment classification on the masked text to be classified to obtain a prediction sequence. Among them, the prediction sequence may include the probability distribution of each target entity word, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories.
[0240] Step S904: Determine the target sentiment category to which each target entity word belongs from multiple predicted sentiment categories according to the predicted probabilities of the multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence.
[0241] In an embodiment of the present disclosure, for any probability distribution in the prediction sequence, the target sentiment category to which the target entity word corresponding to the probability distribution belongs can be determined from multiple predicted sentiment categories according to the predicted probabilities of the multiple predicted sentiment categories indicated by the probability distribution. For example, the predicted probability of the target sentiment category in the probability distribution is greater than the predicted probabilities of other predicted sentiment categories.
[0242] The sentiment classification method of the embodiment of the present disclosure masks at least one target entity word in the text to be classified to obtain the masked text to be classified; uses a sentiment classification model to perform sentiment classification on the masked text to be classified to obtain a prediction sequence; and determines the target sentiment category to which each target entity word belongs from multiple predicted sentiment categories according to the predicted probabilities of the multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence. Thus, by using the sentiment classification model to perform sentiment classification on the masked text to be classified, the output result of the sentiment classification model does not depend on specific target entity words, but only on the context semantic information of the target entity words, which can improve the accuracy of classifying the sentiment category to which the target entity word belongs. Moreover, the sentiment classification model can output the sentiment categories to which all the target entity words in the text to be classified belong at one time, which can improve the prediction efficiency.
[0243] To clearly illustrate how the masked text to be classified is sentiment classified to obtain a prediction sequence in the above embodiments, the present disclosure also proposes a sentiment classification method.
[0244] Figure 10 It is a schematic flowchart of the sentiment classification method provided in Embodiment X of the present disclosure.
[0245] As Figure 10 shown, the sentiment classification method may include the following steps:
[0246] Step S1001: Obtain the text to be classified.
[0247] For the explanation of Step S1001, reference can be made to the relevant descriptions in any embodiment of the present disclosure, and details are not described herein again.
[0248] In any embodiment of the present disclosure, the manner of obtaining the text to be classified is, for example:
[0249] First, an initial text to be classified can be obtained. There is no limitation on the manner of obtaining the initial text. For example, the initial text can be collected online, or the initial text can be provided manually, etc.
[0250] After that, according to the input requirements of the sentiment classification model, preprocessing can be performed on the initial text; wherein, the preprocessing includes at least one of the following: font conversion, font format conversion. For example, according to the input requirements, the initial text can be converted between simplified and traditional Chinese, between full-width and half-width, and from uppercase to lowercase, etc.
[0251] Then, it can be determined whether the number of characters contained in the preprocessed initial text is greater than the set number threshold in the input requirements. If so, the preprocessed initial text is segmented to obtain the text to be classified. If not, the preprocessed initial text is used as the text to be classified.
[0252] Thus, it can be ensured that the font and font format of the preprocessed initial text match the input requirements of the sentiment classification model, improving the effectiveness and accuracy of the text prediction by the sentiment classification model. Moreover, it can be ensured that the number of characters contained in the text to be classified input into the sentiment classification model matches the input requirements of the sentiment classification model, improving the effectiveness and accuracy of the text prediction by the sentiment classification model.
[0253] Step S1002, according to the mask label, perform mask processing on each target entity word in the text to be classified to obtain the masked text to be classified.
[0254] In the embodiments of the present disclosure, a mask label can be used to perform mask processing on each target entity word in the text to be classified to obtain the masked text to be classified. Its implementation principle is similar to that of step S202 and will not be elaborated here.
[0255] It should be noted that the text to be classified may contain some entity words that are irrelevant to the target task adapted to the sentiment classification model. For example, when the target task is to identify the sentiment category of a specific product, the text to be classified may contain the specific product, as well as entity words such as place names and personal names. At this time, in order to improve the prediction efficiency of the sentiment classification model, sentiment classification can be performed only on the target entity words related to the target task.
[0256] Therefore, in any embodiment of the present disclosure, a list of entity words associated with the target task can be obtained, and each entity word in the text to be classified that is in the list of entity words is used as a target entity word. Thus, in the present disclosure, mask processing can be performed on each target entity word in the text to be classified to obtain the masked text to be classified.
[0257] Wherein, the list of entity words contains each entity word associated with the target task, and the target task is used to indicate the sentiment category of the entity word that the sentiment classification model needs to predict.
[0258] Accordingly, the sentiment classification model only needs to perform sentiment classification on the target entity words associated with the target task, rather than on other entity words irrelevant to the target task, which can improve the prediction efficiency of the sentiment classification model.
[0259] Step S1003: Input the masked text to be classified into the encoder in the sentiment classification model, so as to use the encoder to encode each masked label based on the context information of each masked label, and obtain the encoded features of each masked label.
[0260] In the embodiments of the present disclosure, the masked text to be classified can be input into the encoder in the trained sentiment classification model, so as to use the encoder to encode each masked label based on the context information (or context semantic information) of each masked label in the masked text to be classified, and obtain the encoded features of each encoded label.
[0261] Step S1004: Use the decoder in the sentiment classification model to predict the sentiment category of each target entity word based on the encoded features of each masked label, so as to obtain a prediction sequence.
[0262] Among them, the prediction sequence includes the probability distribution of each target entity word, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories.
[0263] In the embodiments of the present disclosure, the decoder in the trained sentiment classification model can be used to predict the sentiment category of each target entity word based on the encoded features of each masked label, and obtain a prediction sequence. Among them, the prediction sequence can include the probability distribution of each target entity word, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories.
[0264] Step S1005: Determine the target sentiment category to which each target entity word belongs from multiple predicted sentiment categories according to the prediction probabilities of multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence.
[0265] For the explanation of Step S1005, reference can be made to the relevant description in any embodiment of the present disclosure, and details are not described herein again.
[0266] The sentiment classification method of the embodiments of the present disclosure predicts the sentiment category of each target entity word based on the context semantic information of each target entity word in the text to be classified, which can improve the accuracy of the prediction result.
[0267] In any one of the embodiments of the present disclosure, a fine-grained sentiment classification model provided by the present disclosure for specific entity words associated with a target task can be used to classify the sentiment tendencies of entity words such as personal names and company names. That is, a sentiment classification model independent of entity words is adopted, which does not depend on specific entity words but only on the context semantic information of entity words to predict the sentiment category (i.e., sentiment tendency) of entity words. Therefore, for different target tasks, the list of entity words associated with the target task can be changed to improve the applicability of the method. Moreover, the sentiment classification model can calculate the sentiment categories of all entity words in the text at one time, and the prediction efficiency is relatively high.
[0268] Taking the number of sentiment categories as 3, namely negative, neutral, and positive, as an example, the technical solutions provided by the present disclosure mainly include the following aspects:
[0269] First aspect, task introduction.
[0270] A text may contain multiple entity words, and each entity word may have different sentiment tendencies. Moreover, for the same entity word, if the position of the same entity word in the text is different, due to the different context information of the same entity word at different positions, the same entity word at different positions will also have different sentiment tendencies. Therefore, it is necessary to perform fine-grained sentiment classification on entity words, and the task is defined as follows:
[0271] For the input text X = {x1, x2,..., x n} of the model, where X contains entity words E = {e1, e2,..., e m}, e i is the i-th entity word in X, and e i = {x i1 , x i2 ,..., x ik}, that is, e i includes k characters. Each entity word e i in X can be replaced with the "[MASK]" symbol to obtain the input X' = {x1,..., [MASK],..., [MASK],..., x n} of the model. After the prediction of the model, a prediction sequence A = {a1, a2,..., a i ,..., a m} is obtained, where a i = {p ipos , p ineg , p ineu} represents the probability distribution of the entity word e i , p ipos represents the prediction probability that the sentiment category of the entity word e i is positive, and p ineg represents the entity word ei The predicted probability that the sentiment category is negative, p ineu represents the entity word e i The predicted probability that the sentiment category is neutral.
[0272] Second aspect, model introduction.
[0273] The present disclosure provides a model for fine-grained sentiment analysis for a specific entity word (an entity word associated with a target task) (denoted as a sentiment classification model in the present disclosure), which can not only identify different sentiment tendencies for different entity words, but also identify different sentiment tendencies for the same entity word. The model structure can be as Figure 11 shown.
[0274] Among them, Figure 11 Label1 in is the annotation label corresponding to the entity word "Zhang San", Label2 is the annotation label corresponding to the entity word "Li Si", Logits1 is the probability distribution corresponding to "Zhang San" output by the model, and Logits2 is the probability distribution corresponding to "Li Si" output by the model. [CLS] is the abbreviation of "classification", which usually represents the beginning of a sentence or text in a text classification task. [SEP] is the abbreviation of "separator", which usually represents the end of a sentence or text.
[0275] As Figure 11 shown, the model can use a pre-trained language model (subsequently referred to as a pre-trained model), such as BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoding representation based on Transformer), RoBERTa (A Robustly Optimized BERT Pretraining Approach, a tuned version of the BERT model), ERNIE (Enhanced Language Representation with Informative Entities, enhanced representation through knowledge fusion), etc. as an encoder. First, preprocess the text X input to the model, and perform a masking (MASK) operation on each entity word in the text X. For example, replace the entity word in the text X with the "[MASK]" label to obtain the input part X' of the model. Then, encode the input part X' through the pre-trained model to obtain the hidden layer vector H = {h1, h2,..., h n}, where h i represents the vector representation of the i-th character in X' (denoted as the encoding feature in the present disclosure). This model mainly uses the vector representation H' = {he1 , h e2 , …, h em} (h ej represents the encoded feature of the j-th "[MASK]" label) to predict the sentiment category corresponding to the entity words E = {e1, e2, …, e m}}, and obtain the prediction result A = {a1, a2, …, a i , …, a m}}. This model can use the cross-entropy loss function to train the model.
[0276] The third part, the pre-training stage.
[0277] The first pre-training: Modify the MLM (Masked Language Model) task in the pre-trained model, mask specific entity words, nouns, and proper nouns in the text, and predict the masked character part in the original text through the model.
[0278] The second pre-training: Use the open-source fine-grained sentiment analysis dataset to perform the second pre-training on the model, as follows:
[0279] 1. Collect the open-source fine-grained sentiment analysis dataset;
[0280] 2. Extract the relevant entity words and their corresponding sentiment tendencies in each text in the dataset;
[0281] 3. According to the order of appearance of the entity words, construct annotation labels according to the sentiment tendencies of each entity word in step 2 in turn;
[0282] 4. Use the annotated text to perform the second pre-training on the model.
[0283] The fourth part: The training stage, including the construction of the training set associated with the target task, text cleaning, and training the model using the cleaned text.
[0284] After the two-stage pre-training in the third aspect, the model already has preliminary sentiment classification ability. Therefore, the training set can be constructed by leveraging the model's sentiment classification ability. For example, the training set can be constructed using the following steps:
[0285] 1. Collect unannotated text related to the target task (denoted as the fourth sample text in this disclosure);
[0286] 2. Preprocess the unannotated text, such as converting traditional Chinese to simplified Chinese, full-width to half-width conversion, and converting uppercase letters to lowercase, etc.;
[0287] 3. Divide the unmarked text into sentences according to punctuation marks, and ensure that the same entity word appears only once in the same sentence. If the same entity word appears more than once in the sentence after the sentence division, it will be discarded;
[0288] 4. Use the model to predict the sentiment category of entity words in the sentence;
[0289] 5. Concatenate the sentences in the same text. The text annotation label sequence includes the annotation labels of each entity word in the text (determined according to the sentiment category predicted in step 4), and ensure that the sentence length is less than the set number threshold (for example, 512).
[0290] 6. Construct a training set based on the text annotated with the labeled label sequence.
[0291] Among them, text cleaning: Since the prediction results of the model are not necessarily accurate, the training set needs to be cleaned. K-fold cross validation can be used for text cleaning. The specific steps are as follows:
[0292] 1. Divide the training set into K groups;
[0293] 2. Take one of the groups as the test set, and use the remaining K-1 groups as the training set to train the model, and K models can be obtained;
[0294] 3. Model M generated using the training set selected for the i-th time i Predict the test set selected for the i-th time. If the prediction result is inconsistent with the annotated label sequence, manually annotate the inconsistent text.
[0295] Afterwards, the model can be trained using the cleaned text.
[0296] Part 5: Prediction stage, taking the number threshold as 512 as an example, mainly includes the following steps:
[0297] 1. Definition of entity word list: Prepare entity word list according to business needs, which may include celebrity names, company names, etc.
[0298] 2. Preprocess the input text (the specific operation is the same as that during training);
[0299] 3. For the preprocessed text, 512 characters are intercepted with entity words as the center, and all entity words in the text are replaced with the "[MASK]" symbol as the input part of the model;
[0300] 4. Send the input text to the model for prediction, obtain the probability distribution for each entity word, and take one or more predicted sentiment categories with the highest prediction probability in the probability distribution as the target sentiment category to which the entity word belongs;
[0301] 5. If the entity words in the text contain positive or negative sentiment categories, the text can be output for manual analysis.
[0302] In summary, the sentiment classification model can be used in social scenarios to analyze comments in social networks and identify negative comments about specific celebrities. The sentiment classification model can be used in public opinion analysis scenarios to analyze public opinion on entity words such as newly released products of enterprises, and the positive or negative evaluation of users on the products can be obtained.
[0303] With the above Figures 1 to 8 Corresponding to the training method of the sentiment classification model provided in the embodiment, the present disclosure also provides a training device for the sentiment classification model. Figures 1 to 8 The training method of the sentiment classification model provided in the embodiment corresponds to the embodiment, so the implementation method of the sentiment classification model training method is also applicable to the training device of the sentiment classification model provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0304] Figure 12 This is a schematic diagram of the structure of the training device of the sentiment classification model provided in the eleventh embodiment of the present disclosure.
[0305] like Figure 12 As shown, the training device 1200 for the sentiment classification model may include: a first acquisition module 1201 , a mask module 1202 , a classification module 1203 and a first training module 1204 .
[0306] The first acquisition module 1201 is used to acquire a first sample text annotated with a first annotation label sequence; wherein each element in the first annotation label sequence is used to indicate an annotation sentiment category corresponding to a first entity word in the first sample text.
[0307] The masking module 1202 is used to perform masking processing on each first entity word in the first sample text to obtain a masked first sample text.
[0308] The classification module 1203 is used to use the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence; wherein the first prediction sequence includes a first probability distribution of each first entity word, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted sentiment categories.
[0309] The first training module 1204 is used to perform a first training on the sentiment classification model according to the first prediction sequence and the first annotation label sequence.
[0310] In a possible implementation manner of the embodiments of the present disclosure, the masking module 1202 is configured to: perform masking processing on each first entity word in the first sample text according to the masking label to obtain the masked first sample text; correspondingly, the classification module 1203 is configured to: input the masked first sample text into the encoder in the sentiment classification model, and use the encoder to encode each masking label based on the context information of each masking label to obtain the encoded features of each masking label; use the decoder in the sentiment classification model to predict the sentiment category of each first entity word based on the encoded features of each masking label to obtain the first prediction sequence.
[0311] In a possible implementation manner of the embodiments of the present disclosure, the first training module 1204 is configured to: for the i-th first probability distribution in the first prediction sequence, generate the classification loss value of the i-th first entity word in the first sample text according to the prediction probabilities of the multiple predicted sentiment categories indicated by the i-th first probability distribution and the labeled sentiment category indicated by the i-th element in the first labeled label sequence; where i is a positive integer; determine the target loss value according to the classification loss values of each first entity word; and perform the first training on the sentiment classification model according to the target loss value.
[0312] In a possible implementation manner of the embodiments of the present disclosure, the first training module 1204 is configured to: determine the target sentiment category to which each first entity word belongs from the multiple predicted sentiment categories according to the prediction probabilities of the multiple predicted sentiment categories indicated by each first probability distribution in the first prediction sequence; generate a predicted label sequence according to the target sentiment category to which each first entity word belongs; and perform the first training on the sentiment classification model according to the difference between the predicted label sequence and the first labeled label sequence.
[0313] In a possible implementation manner of the embodiments of the present disclosure, the training device 1200 of the sentiment classification model may further include:
[0314] A second acquisition module, configured to acquire a list of entity words associated with the target task adapted to the sentiment classification model.
[0315] A processing module, configured to use each entity word in the first sample text that is in the list of entity words as the first entity word.
[0316] In a possible implementation manner of the embodiments of the present disclosure, the training device 1200 of the sentiment classification model may further include:
[0317] A third acquisition module, configured to acquire a second sample text.
[0318] The masking module 1202 is further configured to mask at least one target text segment in the second sample text to obtain the masked second sample text; wherein, the target text segment includes at least one of a specified entity word, a noun, and a proper noun.
[0319] The prediction module is configured to perform character prediction on the masked second sample text by using an initial sentiment classification model to obtain a prediction text.
[0320] The first pre-training module is configured to perform first pre-training on the sentiment classification model according to the prediction text.
[0321] In a possible implementation manner of the embodiment of the present disclosure, the number of prediction texts is the same as the number of target text segments, and the first pre-training module is configured to: perform first pre-training on the sentiment classification model according to the differences between each target text segment and each prediction text; or, the number of prediction texts is one, and the first pre-training module is configured to: perform first pre-training on the sentiment classification model according to the differences between the prediction text and the second sample text.
[0322] In a possible implementation manner of the embodiment of the present disclosure, the training device 1200 of the sentiment classification model may further include:
[0323] The fourth acquisition module is configured to acquire at least one third sample text from at least one dataset irrelevant to the target task adapted to the sentiment classification model.
[0324] The generation module is configured to generate a second annotation label sequence according to the annotation sentiment categories labeled by at least one second entity word in the third sample text.
[0325] The classification module 1203 is further configured to perform sentiment classification on the third sample text by using the sentiment classification model to obtain a second prediction sequence; wherein, the second prediction sequence includes the second probability distribution of each second entity word, and the second probability distribution is used to indicate the prediction probability that the corresponding second entity word belongs to multiple predicted sentiment categories.
[0326] The second pre-training module is configured to perform second pre-training on the sentiment classification model according to the second annotation label sequence and the second prediction sequence.
[0327] In a possible implementation manner of the embodiment of the present disclosure, the classification module 1203 is configured to: perform masking processing on each second entity word in the third sample text to obtain the masked third sample text; perform sentiment classification on the masked third sample text by using the sentiment classification model to obtain a second prediction sequence.
[0328] In a possible implementation manner of the embodiment of the present disclosure, the first acquisition module 1201 is configured to: acquire a first sample text from a first training set, where the first training set is acquired by the following steps: acquiring a plurality of unlabeled fourth sample texts associated with a target task adapted to a sentiment classification model; for any one of the fourth sample texts, performing sentiment classification on any one of the fourth sample texts by using the sentiment classification model pre-trained for the second time to obtain a third prediction sequence of any one of the fourth sample texts; where the third prediction sequence includes a third probability distribution of each third entity word in any one of the fourth sample texts, and the third probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to multiple predicted sentiment categories; labeling any one of the fourth sample texts according to the third prediction sequence to obtain the labeled any one of the fourth sample texts; generating a first training set according to the plurality of labeled fourth sample texts.
[0329] In a possible implementation manner of the embodiment of the present disclosure, the first acquisition module 1201 is configured to: for any one of the fourth sample texts, perform masking processing on each third entity word in any one of the fourth sample texts to obtain a masked fourth sample text; use the sentiment classification model pre-trained for the second time to perform sentiment classification on the masked fourth sample text to obtain a third prediction sequence of any one of the fourth sample texts.
[0330] In a possible implementation manner of the embodiment of the present disclosure, the first acquisition module 1201 is configured to: determine a target sample text with incorrect labeling from the plurality of labeled fourth sample texts; in response to an update operation on the target sample text, update the labeling data of the target sample text; generate a first training set according to the updated target sample text and the remaining texts other than the target sample text in the plurality of labeled fourth sample texts.
[0331] In a possible implementation manner of the embodiment of the present disclosure, the first acquisition module 1201 is configured to: divide the plurality of labeled fourth sample texts into K groups, and sequentially select one group from the K groups as a test set; for any selected test set, use the remaining groups other than the any selected test set in the K groups as a second training set, perform second training on the sentiment classification model by using the second training set, and perform sentiment classification on the any selected test set based on the sentiment classification model after the second training to obtain a fourth prediction sequence of each labeled fourth sample text in the any selected test set; determine a target sample text from the any selected test set according to the fourth prediction sequence and the third prediction sequence of each labeled fourth sample text in the any selected test set.
[0332] In a possible implementation manner of the embodiments of the present disclosure, the first acquisition module 1201 is further configured to: preprocess the fourth sample text according to the input requirements of the sentiment classification model, where the preprocessing includes at least one of the following: font conversion, font format conversion; or, screen multiple fourth sample texts, where each third entity word included in the retained fourth sample texts is not repeated.
[0333] In a possible implementation manner of the embodiments of the present disclosure, the first acquisition module 1201 is further configured to: determine that the number of characters included in the fourth sample text is less than or equal to a set number threshold in the input requirements.
[0334] In a possible implementation manner of the embodiments of the present disclosure, the first acquisition module 1201 is further configured to: in the case where the number of characters included in the fourth sample text is greater than the set number threshold, split the fourth sample text so that the number of characters included in the split fourth sample text is less than or equal to the set number threshold.
[0335] The training device of the sentiment classification model in the embodiments of the present disclosure masks each first entity word in the first sample text to obtain the masked first sample text; uses the sentiment classification model to perform sentiment classification on the masked first sample text to obtain a first prediction sequence; and performs first training on the sentiment classification model according to the first prediction sequence and the first annotation label sequence annotated by the first sample text. Thus, by using the sentiment classification model to perform sentiment classification on the masked first sample text, the output result of the sentiment classification model does not depend on specific first entity words, but only depends on the context semantic information of the first entity words, which can improve the accuracy of classifying the sentiment categories to which the first entity words belong. Moreover, the sentiment classification model can output the sentiment categories to which all the first entity words in the first sample text belong at one time, which can improve the prediction efficiency.
[0336] Corresponding to the sentiment classification method provided in the above Figures 9 to 10 embodiment, the present disclosure further provides a sentiment classification device. Since the sentiment classification device provided in the embodiments of the present disclosure corresponds to the Figures 9 to 10 sentiment classification method provided in the above embodiment, the implementation manner of the sentiment classification method is also applicable to the sentiment classification device provided in the embodiments of the present disclosure and will not be described in detail in the embodiments of the present disclosure.
[0337] Figure 13 It is a schematic structural diagram of the sentiment classification device provided in Embodiment Twelve of the present disclosure.
[0338] As Figure 13 shown, the sentiment classification device 1300 may include: a first acquisition module 1301, a masking module 1302, a classification module 1303, and a determination module 1304.
[0339] Among them, the first acquisition module 1301 is configured to acquire the text to be classified.
[0340] The masking module 1302 is configured to perform masking processing on at least one target entity word in the text to be classified, so as to obtain the masked text to be classified.
[0341] The classification module 1303 is configured to perform sentiment classification on the masked text to be classified by using a sentiment classification model, so as to obtain a prediction sequence; wherein, the prediction sequence includes the probability distribution of each target entity word, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories.
[0342] The determination module 1304 is configured to determine the target sentiment category to which each target entity word belongs from multiple predicted sentiment categories according to the prediction probabilities of multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence.
[0343] In a possible implementation manner of the embodiment of the present disclosure, the sentiment classification device 1300 may further include:
[0344] The second acquisition module is configured to acquire a list of entity words associated with a target task adapted to the sentiment classification model, where the list of entity words includes multiple target entity words associated with the target task.
[0345] The processing module is configured to use each entity word in the text to be classified that is located in the list of entity words as a target entity word.
[0346] In a possible implementation manner of the embodiment of the present disclosure, the first acquisition module 1301 is configured to: acquire an initial text to be classified; perform preprocessing on the initial text according to the input requirements of the sentiment classification model; wherein, the preprocessing includes at least one of the following: font conversion, font format conversion; determine whether the number of characters included in the preprocessed initial text is greater than a set number threshold in the input requirements; if the number of characters is less than or equal to the set number threshold, use the preprocessed initial text as the text to be classified; if the number of characters is greater than the set number threshold, perform segmentation on the preprocessed initial text to obtain the text to be classified.
[0347] In a possible implementation manner of the embodiments of the present disclosure, the masking module 1302 is configured to: perform masking processing on each target entity word in the text to be classified according to the masking label to obtain the masked text to be classified; correspondingly, the classification module 1303 is configured to: input the masked text to be classified into the encoder in the sentiment classification model, and use the encoder to encode each masking label based on the context information of each masking label to obtain the encoded features of each masking label; use the decoder in the sentiment classification model to predict the sentiment category of each target entity word based on the encoded features of each masking label to obtain a prediction sequence.
[0348] The sentiment classification device of the embodiments of the present disclosure performs masking processing on at least one target entity word in the text to be classified to obtain the masked text to be classified; uses the sentiment classification model to perform sentiment classification on the masked text to be classified to obtain a prediction sequence; determines the target sentiment category to which each target entity word belongs from multiple predicted sentiment categories according to the prediction probabilities of the multiple predicted sentiment categories indicated by each probability distribution in the prediction sequence. Thus, when using the sentiment classification model to perform sentiment classification on the masked text to be classified, the output result of the sentiment classification model does not depend on specific target entity words, but only depends on the context semantic information of the target entity words, which can improve the accuracy of classifying the sentiment category to which the target entity words belong. Moreover, the sentiment classification model can output the sentiment categories to which all the target entity words in the text to be classified belong at one time, which can improve the prediction efficiency.
[0349] To implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method or the sentiment classification method of the sentiment classification model proposed in any of the above embodiments of the present disclosure.
[0350] To implement the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the training method or the sentiment classification method of the sentiment classification model proposed in any of the above embodiments of the present disclosure.
[0351] To implement the above embodiments, the present disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the training method or the sentiment classification method of the sentiment classification model proposed in any of the above embodiments of the present disclosure.
[0352] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0353] Figure 14 A schematic block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. Among them, the electronic device may include the server and the client in the above embodiments. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0354] As Figure 14 shown, the electronic device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to the computer program stored in the ROM (Read-Only Memory) 1402 or the computer program loaded from the storage unit 1408 into the RAM (Random Access Memory) 1403. In the RAM 1403, various programs and data required for the operation of the electronic device 1400 can also be stored. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other through a bus 1404. The I / O (Input / Output) interface 1405 is also connected to the bus 1404.
[0355] Multiple components in the electronic device 1400 are connected to the I / O interface 1405, including: an input unit 1406, such as a keyboard, a mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, an optical disk, etc.; and a communication unit 1409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1409 allows the electronic device 1400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0356] The computing unit 1401 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 1401 executes the various methods and processes described above, such as the above-mentioned sentiment classification method or the training method of the sentiment classification model. For example, in some embodiments, the above-mentioned sentiment classification method or the training method of the sentiment classification model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into the RAM 1403 and executed by the computing unit 1401, one or more steps of the above-described sentiment classification method or the training method of the sentiment classification model may be executed. Alternatively, in other embodiments, the computing unit 1401 may be configured to execute the above-described sentiment classification method or the training method of the sentiment classification model in any other suitable manner (e.g., by means of firmware).
[0357] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0358] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0359] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0360] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0361] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0362] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server may also be a server of a distributed system or a server combined with a blockchain.
[0363] Among them, it should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0364] Deep learning is a new research direction in the field of machine learning. It learns the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans and be able to recognize data such as text, images, and sounds.
[0365] According to the technical solution of the embodiment of the present disclosure, each first entity word in the first sample text is masked to obtain the masked first sample text; the masked first sample text is subjected to sentiment classification using a sentiment classification model to obtain a first prediction sequence; the sentiment classification model is first trained according to the first prediction sequence and the first annotation label sequence annotated by the first sample text. Thus, when using the sentiment classification model to perform sentiment classification on the masked first sample text, the output result of the sentiment classification model does not depend on the specific first entity word but only on the context semantic information of the first entity word, which can improve the accuracy of classifying the sentiment category to which the first entity word belongs. Moreover, the sentiment classification model can output all the sentiment categories to which the first entity words in the first sample text belong at one time, which can improve the prediction efficiency.
[0366] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution proposed in this disclosure can be achieved, and no limitations are imposed herein.
[0367] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for an emotion classification model, comprising: Obtaining a first sample text labeled with a first labeled tag sequence; wherein, each element in the first labeled tag sequence is used to indicate the labeled emotion category of the corresponding first entity word in the first sample text; Performing masking processing on each first entity word in the first sample text to obtain a masked first sample text; Using an emotion classification model to perform emotion classification on the masked first sample text to obtain a first prediction sequence; wherein, the first prediction sequence includes the first probability distribution of each first entity word, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted emotion categories; Performing first training on the emotion classification model according to the first prediction sequence and the first labeled tag sequence; Wherein, before using the emotion classification model to perform emotion classification on the masked first sample text to obtain a first prediction sequence, the method further includes: Obtaining at least one third sample text from at least one dataset irrelevant to the target task adapted to the emotion classification model; Generating a second labeled tag sequence according to the labeled emotion category labeled by at least one second entity word in the third sample text; Using an emotion classification model to perform emotion classification on the third sample text to obtain a second prediction sequence; wherein, the second prediction sequence includes the second probability distribution of each second entity word, and the second probability distribution is used to indicate the prediction probability that the corresponding second entity word belongs to multiple predicted emotion categories; Performing second pre-training on the emotion classification model according to the second labeled tag sequence and the second prediction sequence.
2. The method according to claim 1, wherein, The performing masking processing on each first entity word in the first sample text to obtain a masked first sample text includes: Performing masking processing on each first entity word in the first sample text according to a masking tag to obtain a masked first sample text; Correspondingly, the using an emotion classification model to perform emotion classification on the masked first sample text to obtain a first prediction sequence includes: Inputting the masked first sample text into an encoder in the emotion classification model to use the encoder to encode each masking tag based on the context information of each masking tag to obtain the encoded feature of each masking tag; Using a decoder in the emotion classification model to predict the emotion category of each first entity word based on the encoded feature of each masking tag to obtain a first prediction sequence.
3. The method according to claim 1, wherein The performing first training on the emotion classification model according to the first prediction sequence and the first labeled tag sequence includes: For the i-th first probability distribution in the first prediction sequence, generating a classification loss value of the i-th first entity word in the first sample text according to the prediction probabilities of multiple predicted emotion categories indicated by the i-th first probability distribution and the labeled emotion category indicated by the i-th element in the first labeled tag sequence; wherein, i is a positive integer; Determining a target loss value according to the classification loss values of each first entity word; Perform the first training on the sentiment classification model according to the target loss value.
4. The method according to claim 1, wherein The performing the first training on the sentiment classification model according to the first prediction sequence and the first labeled tag sequence includes: Determine the target sentiment category to which each of the first entity words belongs from the multiple predicted sentiment categories indicated by the prediction probabilities of the multiple predicted sentiment categories in the first prediction sequence; Generate a predicted tag sequence according to the target sentiment category to which each of the first entity words belongs; Perform the first training on the sentiment classification model according to the difference between the predicted tag sequence and the first labeled tag sequence.
5. The method according to claim 1, wherein Before masking each of the first entity words in the first sample text to obtain the masked first sample text, the method further includes: Obtain a list of entity words associated with a target task adapted to the sentiment classification model; Use each entity word in the first sample text that is in the list of entity words as the first entity word.
6. The method according to claim 1, wherein, Before performing sentiment classification on the masked first sample text using the sentiment classification model to obtain a first prediction sequence, the method further includes: Obtain a second sample text; Mask at least one target text segment in the second sample text to obtain a masked second sample text; wherein the target text segment includes at least one of a specified entity word, a noun, and a proper noun; Perform character prediction on the masked second sample text using an initial sentiment classification model to obtain a predicted text; Perform the first pre-training on the sentiment classification model according to the predicted text.
7. The method according to claim 6, wherein The number of the predicted texts is the same as the number of the target text segments, and the performing the first pre-training on the sentiment classification model according to the predicted text includes: Perform the first pre-training on the sentiment classification model according to the difference between each of the target text segments and each of the predicted texts; Alternatively, the number of the predicted texts is one, and the performing the first pre-training on the sentiment classification model according to the predicted text includes: Perform the first pre-training on the sentiment classification model according to the difference between the predicted text and the second sample text.
8. The method according to claim 1, wherein The performing sentiment classification on the third sample text using the sentiment classification model to obtain a second prediction sequence includes: Mask each of the second entity words in the third sample text to obtain a masked third sample text; Perform sentiment classification on the masked third sample text using the sentiment classification model to obtain a second prediction sequence.
9. The method according to claim 1, wherein, The obtaining the first sample text labeled with the first labeled tag sequence includes: Obtain the first sample text from a first training set, wherein the first training set is obtained by the following steps: Obtain a plurality of unlabeled fourth sample texts associated with a target task adapted to the sentiment classification model; For any fourth sample text, use the sentiment classification model pre-trained for the second time to perform sentiment classification on the any fourth sample text to obtain a third prediction sequence of the any fourth sample text; wherein, the third prediction sequence includes a third probability distribution of each third entity word in the any fourth sample text, and the third probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to multiple predicted sentiment categories; Annotate the any fourth sample text according to the third prediction sequence to obtain the annotated any fourth sample text; Generate the first training set according to multiple annotated fourth sample texts.
10. The method according to claim 9, wherein, The step of using the sentiment classification model pre-trained for the second time to perform sentiment classification on the any fourth sample text to obtain a third prediction sequence of the any fourth sample text includes: For any fourth sample text, perform masking processing on each of the third entity words in the any fourth sample text to obtain a masked fourth sample text; Use the sentiment classification model pre-trained for the second time to perform sentiment classification on the masked fourth sample text to obtain a third prediction sequence of the any fourth sample text.
11. The method according to claim 9, wherein The step of generating the first training set according to multiple annotated fourth sample texts includes: Determine a target sample text with incorrect annotation from the multiple annotated fourth sample texts; In response to an update operation on the target sample text, update the annotation data of the target sample text; Generate the first training set according to the updated target sample text and the remaining texts in the multiple annotated fourth sample texts except the target sample text.
12. The method according to claim 11, wherein, The step of determining a target sample text with incorrect annotation from the multiple annotated fourth sample texts includes: Divide the multiple annotated fourth sample texts to obtain K groups, and sequentially select one group from the K groups as a test set; For any selected test set, use the remaining groups except the any selected test set among the K groups as a second training set, use the second training set to perform second training on the sentiment classification model, and perform sentiment classification on the any selected test set based on the sentiment classification model after the second training to obtain a fourth prediction sequence of each of the annotated fourth sample texts in the any selected test set; Determine the target sample text from the any selected test set according to the fourth prediction sequence and the third prediction sequence of each of the annotated fourth sample texts in the any selected test set.
13. The method according to claim 9, wherein, After obtaining multiple unannotated fourth sample texts associated with the target task adapted to the sentiment classification model, the method further includes: Preprocess the fourth sample text according to the input requirements of the sentiment classification model, wherein the preprocessing includes at least one of the following: font conversion, font format conversion; Or, Screen the multiple fourth sample texts, wherein each third entity word included in the retained fourth sample texts is not repeated.
14. The method according to claim 13, wherein After obtaining a plurality of unlabeled fourth sample texts associated with a target task adapted to the sentiment classification model, the method further includes: Determine that the number of characters included in the fourth sample text is less than or equal to a set number threshold in the input requirement.
15. The method according to claim 14, wherein The method further includes: In the case where the number of characters included in the fourth sample text is greater than the set number threshold, segment the fourth sample text so that the number of characters included in the segmented fourth sample text is less than or equal to the set number threshold.
16. A sentiment classification method, including: Obtain a text to be classified; Perform masking processing on at least one target entity word in the text to be classified to obtain a masked text to be classified; Use the sentiment classification model obtained by the training method of any one of claims 1-15 to perform sentiment classification on the masked text to be classified to obtain a prediction sequence; wherein, the prediction sequence includes the probability distribution of each target entity word, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories; According to the prediction probabilities of the multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence, determine the target sentiment category to which each target entity word belongs from the multiple predicted sentiment categories.
17. The method according to claim 16, wherein, Before performing masking processing on at least one target entity word in the text to be classified to obtain a masked text to be classified, the method further includes: Obtain a list of entity words associated with a target task adapted to the sentiment classification model, wherein the list of entity words includes a plurality of target entity words associated with the target task; Use each entity word in the text to be classified that is located in the list of entity words as the target entity word.
18. The method according to claim 16, wherein The obtaining of the text to be classified includes: Obtain an initial text to be classified; Preprocess the initial text according to the input requirement of the sentiment classification model; wherein, the preprocessing includes at least one of the following: font conversion, font format conversion; Judge whether the number of characters included in the preprocessed initial text is greater than the set number threshold in the input requirement; If the number of characters is less than or equal to the set number threshold, use the preprocessed initial text as the text to be classified; If the number of characters is greater than the set number threshold, segment the preprocessed initial text to obtain the text to be classified.
19. The method according to any one of claims 16 - 18, wherein, The performing of masking processing on at least one target entity word in the text to be classified to obtain a masked text to be classified includes: Perform masking processing on each target entity word in the text to be classified according to a masking label to obtain a masked text to be classified; Correspondingly, the using of the sentiment classification model to perform sentiment classification on the masked text to be classified to obtain a prediction sequence includes: Input the masked text to be classified into an encoder in the sentiment classification model, and use the encoder to encode each masking label based on the context information of each masking label to obtain the encoded features of each masking label; Using the decoder in the sentiment classification model, predict the sentiment categories of each of the target entity words based on the encoded features of each of the masked labels to obtain the prediction sequence.
20. A training device for a sentiment classification model, comprising: A first acquisition module, configured to acquire a first sample text labeled with a first labeled tag sequence; wherein, each element in the first labeled tag sequence is used to indicate the labeled sentiment category of the corresponding first entity word in the first sample text; A masking module, configured to perform masking processing on each first entity word in the first sample text to obtain a masked first sample text; A classification module, configured to perform sentiment classification on the masked first sample text using a sentiment classification model to obtain a first prediction sequence; wherein, the first prediction sequence includes the first probability distribution of each of the first entity words, and the first probability distribution is used to indicate the prediction probability that the corresponding first entity word belongs to multiple predicted sentiment categories; A first training module, configured to perform first training on the sentiment classification model according to the first prediction sequence and the first labeled tag sequence; Wherein, the device further includes: A fourth acquisition module, configured to acquire at least one third sample text from at least one dataset irrelevant to the target task adapted to the sentiment classification model; A generation module, configured to generate a second labeled tag sequence according to the labeled sentiment categories of at least one second entity word in the third sample text; The classification module is further configured to perform sentiment classification on the third sample text using a sentiment classification model to obtain a second prediction sequence; wherein, the second prediction sequence includes the second probability distribution of each of the second entity words, and the second probability distribution is used to indicate the prediction probability that the corresponding second entity word belongs to multiple predicted sentiment categories; A second pre-training module, configured to perform second pre-training on the sentiment classification model according to the second labeled tag sequence and the second prediction sequence.
21. The apparatus according to claim 20, wherein The masking module is configured to: perform masking processing on each of the first entity words in the first sample text according to a masked label to obtain a masked first sample text; Correspondingly, the classification module is configured to: input the masked first sample text into an encoder in the sentiment classification model, so as to use the encoder to encode each of the masked labels based on the context information of each of the masked labels to obtain the encoded features of each of the masked labels; use the decoder in the sentiment classification model to predict the sentiment category of each of the first entity words based on the encoded features of each of the masked labels to obtain a first prediction sequence.
22. The device according to claim 20, wherein, The first training module is configured to: For the i-th first probability distribution in the first prediction sequence, generate a classification loss value of the i-th first entity word in the first sample text according to the prediction probabilities of the multiple predicted sentiment categories indicated by the i-th first probability distribution and the labeled sentiment category indicated by the i-th element in the first labeled tag sequence; wherein, i is a positive integer; Determine a target loss value according to the classification loss values of each of the first entity words; Perform the first training on the sentiment classification model according to the target loss value.
23. The apparatus according to claim 20, wherein, The first training module is configured to: Determine the target sentiment category to which each of the first entity words belongs from the multiple predicted sentiment categories indicated by the first probability distributions in the first prediction sequence; Generate a predicted label sequence according to the target sentiment category to which each of the first entity words belongs; Perform the first training on the sentiment classification model according to the difference between the predicted label sequence and the first annotated label sequence.
24. The apparatus according to claim 20, wherein The apparatus further includes: A second acquisition module, configured to acquire a list of entity words associated with a target task adapted to the sentiment classification model; A processing module, configured to use each entity word in the first sample text that is in the list of entity words as the first entity word.
25. The device according to claim 20, wherein, The apparatus further includes: A third acquisition module, configured to acquire a second sample text; The masking module is further configured to mask at least one target text segment in the second sample text to obtain a masked second sample text; wherein the target text segment includes at least one of a specified entity word, a noun, and a proper noun; A prediction module, configured to perform character prediction on the masked second sample text by using an initial sentiment classification model to obtain a predicted text; A first pre-training module, configured to perform first pre-training on the sentiment classification model according to the predicted text.
26. The apparatus according to claim 25, wherein, The number of the predicted texts is the same as the number of the target text segments, and the first pre-training module is configured to: perform the first pre-training on the sentiment classification model according to the difference between each of the target text segments and each of the predicted texts; Alternatively, the number of the predicted texts is one, and the first pre-training module is configured to: perform the first pre-training on the sentiment classification model according to the difference between the predicted text and the second sample text.
27. The apparatus according to claim 20, wherein The classification module is configured to: Perform masking processing on each of the second entity words in the third sample text to obtain a masked third sample text; Perform sentiment classification on the masked third sample text by using the sentiment classification model to obtain a second prediction sequence.
28. The apparatus according to claim 20, wherein, The first acquisition module is configured to: Acquire a first sample text from a first training set, wherein the first training set is acquired by the following steps: Acquire a plurality of unannotated fourth sample texts associated with a target task adapted to the sentiment classification model; For any one of the fourth sample texts, perform sentiment classification on the any one of the fourth sample texts by using a sentiment classification model that has undergone second pre-training to obtain a third prediction sequence of the any one of the fourth sample texts; wherein the third prediction sequence includes a third probability distribution of each third entity word in the any one of the fourth sample texts, and the third probability distribution is used to indicate the prediction probability that the corresponding third entity word belongs to a plurality of predicted sentiment categories; Annotate the any one of the fourth sample texts according to the third prediction sequence to obtain an annotated any one of the fourth sample texts; Generate the first training set according to a plurality of annotated fourth sample texts.
29. The device according to claim 28, wherein, The first acquisition module is configured to: For any fourth sample text, perform masking processing on each of the third entity words in the any fourth sample text to obtain a masked fourth sample text; Use the second pre-trained sentiment classification model to perform sentiment classification on the masked fourth sample text to obtain a third prediction sequence of the any fourth sample text.
30. The apparatus according to claim 28, wherein, The first acquisition module is configured to: Determine a target sample text with incorrect annotation from the multiple annotated fourth sample texts; In response to an update operation on the target sample text, update the annotation data of the target sample text; Generate the first training set according to the updated target sample text and the remaining texts in the multiple annotated fourth sample texts except the target sample text.
31. The apparatus according to claim 30, wherein, The first acquisition module is configured to: Divide the multiple annotated fourth sample texts into K groups, and sequentially select one group from the K groups as a test set; For any selected test set, use the remaining groups except the any selected test set in the K groups as a second training set, perform second training on the sentiment classification model using the second training set, and perform sentiment classification on the any selected test set based on the sentiment classification model after the second training to obtain a fourth prediction sequence of each of the annotated fourth sample texts in the any selected test set; Determine the target sample text from the any selected test set according to the fourth prediction sequence and the third prediction sequence of each of the annotated fourth sample texts in the any selected test set.
32. The apparatus according to claim 28, wherein The first acquisition module is further configured to: Preprocess the fourth sample text according to the input requirements of the sentiment classification model, where the preprocessing includes at least one of the following: font conversion, font format conversion; Or, Screen multiple fourth sample texts, where the third entity words included in the remaining fourth sample texts are not repeated.
33. The apparatus according to claim 32, wherein, The first acquisition module is further configured to: Determine that the number of characters included in the fourth sample text is less than or equal to a set number threshold in the input requirements.
34. The apparatus according to claim 33, wherein, The first acquisition module is further configured to: In the case where the number of characters included in the fourth sample text is greater than the set number threshold, split the fourth sample text so that the number of characters included in the split fourth sample text is less than or equal to the set number threshold.
35. A sentiment classification device, comprising: A first acquisition module, configured to acquire a text to be classified; A masking module, configured to perform masking processing on at least one target entity word in the text to be classified to obtain a masked text to be classified; A classification module, which is used to perform sentiment classification on the masked text to be classified by using the sentiment classification model obtained by the training method of the sentiment classification model according to any one of claims 1-15, so as to obtain a prediction sequence; wherein, the prediction sequence includes the probability distribution of each of the target entity words, and the probability distribution is used to indicate the prediction probability that the corresponding target entity word belongs to multiple predicted sentiment categories; A determination module, which is used to determine the target sentiment category to which each of the target entity words belongs from the multiple predicted sentiment categories according to the prediction probabilities of the multiple predicted sentiment categories indicated by the probability distributions in the prediction sequence.
36. The apparatus according to claim 35, wherein, The apparatus further includes: A second acquisition module, which is used to acquire a list of entity words associated with a target task adapted to the sentiment classification model, wherein the list of entity words includes multiple target entity words associated with the target task; A processing module, which is used to use each entity word in the text to be classified that is located in the list of entity words as the target entity word.
37. The apparatus according to claim 35, wherein, The first acquisition module is used for: Acquiring an initial text to be classified; Preprocessing the initial text according to the input requirements of the sentiment classification model; wherein, the preprocessing includes at least one of the following: font conversion, font format conversion; Judging whether the number of characters included in the preprocessed initial text is greater than a set number threshold in the input requirements; If the number of characters is less than or equal to the set number threshold, using the preprocessed initial text as the text to be classified; If the number of characters is greater than the set number threshold, splitting the preprocessed initial text to obtain the text to be classified.
38. The apparatus according to any one of claims 35-37, wherein The masking module is used for: performing masking processing on each of the target entity words in the text to be classified according to a masking label, so as to obtain a masked text to be classified; Correspondingly, the classification module is used for: inputting the masked text to be classified into an encoder in the sentiment classification model, so as to use the encoder to encode each masking label based on the context information of each masking label to obtain an encoded feature of each masking label; using a decoder in the sentiment classification model to predict the sentiment category of each target entity word based on the encoded feature of each masking label, so as to obtain the prediction sequence.
39. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the training method of the sentiment classification model according to any one of claims 1-15, or execute the sentiment classification method according to any one of claims 16-19.
40. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the training method of the sentiment classification model according to any one of claims 1-15, or execute the sentiment classification method according to any one of claims 16-19.
41. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method for training an emotion classification model according to any one of claims 1-15, or implements the steps of the emotion classification method according to any one of claims 16-19.
Citation Information
Patent Citations
Emotion analysis model training method, emotion analysis method, equipment and storage medium
CN115719063A