A method and apparatus for training a classification model and classifying text.
By optimizing the classification label filling position of the pre-trained language model and using masking and consistency training to train the semantic relevance of the text, the problem of low accuracy of the classification model in texts with inconsistent expression and character count is solved, thus improving the robustness and accuracy of the model.
Patent Information
- Application Number
- CN202211570814.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Existing classification models have low accuracy when faced with texts that have different expressions and inconsistent numbers of characters, resulting in insufficient robustness and effectiveness.
By acquiring a pre-trained language model and combining it with the original training text, masked training text, and consistent training text, we can optimize the learning of the classification label filling position, including semantic relevance training of masked input text and consistent input text, and optimize the pre-trained model to improve the model's adaptability to different expression methods and character counts.
This improved the classification accuracy of the model for texts with different expressions and character counts, enhanced the model's robustness and effectiveness, and ensured accurate classification capabilities in diverse texts.
Smart Images

Figure CN116304015B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer science, and in particular to a method and apparatus for training a classification model and classifying text. Background Technology
[0002] Currently, an increasing number of service platforms have emerged online. These platforms provide services to users; for example, shopping platforms offer shopping services to shoppers, gaming platforms offer gaming services to gamers, and chat platforms offer chat services to chatters. To better serve users and enhance the competitiveness of these platforms, user preference categories can be collected. Services can then be tailored to these preferences, aiming to meet users' core needs as much as possible while providing services, thereby increasing user stickiness to the platform. Summary of the Invention
[0003] This application discloses a training classification model, a method for classifying text, and an apparatus.
[0004] Firstly, a method for training a classification model is presented, comprising: obtaining a pre-trained language model; obtaining original training text and original classification labels of the original training text, wherein the original training text contains at least original input text and template text, the template text including prompt text and classification label fill-in positions; the original classification labels of the original training text are obtained at least based on the classification labels of the original input text; obtaining masked training text and masked classification labels of the masked training text based on the original training text and the original classification labels; the masked training text contains at least masked input text and template text, the masked input text being obtained by masking at least one character in the original input text, and the masked classification labels being obtained at least based on the original classification labels. The labels are obtained and / or at least based on at least one masked character in the original input text; consistent training text and consistent classification labels of the consistent training text are obtained based on the original training text and the original classification labels; the consistent training text contains at least template text and consistent input text that is semantically related to the original input text; the characters in the consistent input text are not all the same as the characters in the original input text, and the consistent classification labels are obtained at least based on the original classification labels; the pre-trained language model is optimized for the classification prediction task of the classification label filling position using at least the original training text, the original classification labels, the masked training text, the masked classification labels, the consistent training text, and the consistent classification labels to obtain the classification model.
[0005] Secondly, a method for text classification is presented, comprising: acquiring a text to be processed, which includes at least a text to be classified and a template text, the template text including prompt text and a classification label filling position; predicting the classification label of the text to be processed for filling in the classification label filling position in the template text based on a trained classification model; and acquiring the classification label of the text to be classified based on the classification label of the text to be processed; wherein the trained classification model is obtained by optimizing a pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, the original classification label, the masked training text, the masked classification label, the consistent training text, and the consistent classification label; the original training text includes at least the original input text. The training text includes both text and template text, with the template text including prompt text and category label fields; the original category label of the original training text is obtained at least based on the category label of the original input text; the masked training text includes at least masked input text and template text, with the masked input text obtained by masking at least one character in the original input text, and the masked category label obtained at least based on the original category label of the original training text and / or at least based on at least one masked character in the original input text; the consistency training text includes template text and consistency input text semantically related to the original input text; the characters in the consistency input text are not all the same as the characters in the original input text, and the consistency category label is obtained at least based on the original category label.
[0006] Thirdly, an apparatus for training a classification model is shown, comprising: a first acquisition module for acquiring a pre-trained language model; acquiring original training text and original classification labels of the original training text, wherein the original training text contains at least original input text and template text, the template text including prompt text and classification label fill-in positions; the original classification labels of the original training text are obtained at least based on the classification labels of the original input text; a second acquisition module for acquiring masked training text and masked classification labels of the masked training text based on the original training text and original classification labels; the masked training text contains at least masked input text and template text, the masked input text being obtained by masking at least one character in the original input text, and the masked classification labels being obtained at least based on the original classification labels of the original training text. The class labels are obtained and / or at least based on at least one masked character in the original input text; the third acquisition module is used to acquire consistent training text and consistent classification labels of consistent training text based on the original training text and the original classification labels; the consistent training text contains at least template text and consistent input text that is semantically related to the original input text; the characters in the consistent input text are not all the same as the characters in the original input text, and the consistent classification labels are obtained at least based on the original classification labels; the optimization module is used to optimize the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, the original classification labels, the masked training text, the masked classification labels, the consistent training text, and the consistent classification labels, to obtain the classification model.
[0007] Fourthly, an apparatus for text classification is shown, comprising: a fourth acquisition module for acquiring a text to be processed, the text to be processed containing at least a text to be classified and a template text, the template text including prompt text and a classification label filling position; a prediction module for predicting, based on a trained classification model, the classification label of the text to be processed to be filled in the classification label filling position in the template text of the text to be processed; and a fifth acquisition module for acquiring the classification label of the text to be classified according to the classification label of the text to be processed; wherein the trained classification model is obtained by optimizing a pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, the original classification label, the masked training text, the masked classification label, the consistent training text, and the consistent classification label; ... masked training text, the masked classification label, the consistent training text, and the consistent classification label; the original training text, the original classification label, the original classification label, the original classification label, the original classification label, the original training text, the original classification label, the original classification label, the original classification label, the original training text, the original classification label, the original classification label, the original classification label, the original training text, the original classification label, the original classification label, the original classification label, the original training text, the original classification label, the original classification label, the original classification label, the original classification label, the The training text contains at least the original input text and template text, the template text including prompt text and category label fields; the original category label of the original training text is obtained at least based on the category label of the original input text; the masked training text contains at least the masked input text and template text, the masked input text is obtained by masking at least one character in the original input text, and the masked category label is obtained at least based on the original category label of the original training text and / or at least based on at least one masked character in the original input text; the consistency training text contains template text and consistency input text semantically related to the original input text; the characters in the consistency input text are not all the same as the characters in the original input text, and the consistency category label is obtained at least based on the original category label.
[0008] Fifthly, this application discloses an electronic device comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the methods shown in any of the foregoing aspects.
[0009] Sixthly, this application discloses a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods shown in any of the foregoing aspects.
[0010] In a seventh aspect, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the methods shown in any of the foregoing aspects.
[0011] On one hand, in the scenario of optimizing a pre-trained language model for a classification prediction task involving filling in classification labels, there is consistent training text and consistent classification labels for the consistent training text. The consistent training text contains at least template text and consistent input text semantically related to the original input text. The characters in the consistent input text are not entirely the same as those in the original input text, and the consistent classification labels are obtained at least based on the original classification labels. Furthermore, it also contains the original training text and its original classification labels.
[0012] This approach allows for consistent regularization optimization of the pre-trained language model based on the original training text, original classification labels, consistent training text, and consistent classification labels. This ensures that minor modifications to the input text do not alter the language model's prediction of the input text's classification label. It can improve the accuracy of the trained classification model in classifying texts with more diverse expressions / concepts, enhance the robustness and effectiveness of the trained classification model, and ultimately improve its accuracy in text classification.
[0013] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained at least based on the original classification label in the original training text.
[0014] This approach allows for the optimization and learning of a pre-trained language model based on the masked training text and its masked classification labels. This enables the pre-trained language model to learn from training texts with fewer characters, allowing it to focus more intently on learning relevant information from these smaller texts. Consequently, the trained classification model demonstrates stronger ability to extract relevant information from texts with fewer characters, improving its accuracy in classifying texts with fewer characters. This enhances the robustness and effectiveness of the trained classification model, ultimately leading to higher accuracy in text classification.
[0015] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained based on at least one masked character in the original input text.
[0016] This allows the pre-trained language model to optimize the learning of semantic relationships between text contexts based on the masked training text and the masked classification labels of the masked training text. This enables the classification model to learn the semantic and contextual relationships between words in the text, thereby improving the accuracy of the trained classification model in predicting masked words based on words other than the masked words in text containing "input text and template text (the template text contains at least prompt text and masked words, etc.)". Furthermore, it can improve the accuracy of the trained classification model in predicting the classification labels of the input text used to fill in the classification label positions based on the input text and prompt text in text containing "input text and template text (the template text contains at least prompt text and classification label filling positions, etc.)". Ultimately, this improves the accuracy of the trained classification model in text classification. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of a method for training a classification model according to this application.
[0018] Figure 2 This is a flowchart illustrating the steps of a text classification method according to this application.
[0019] Figure 3 This is a structural block diagram of a device for training a classification model according to this application.
[0020] Figure 4 This is a structural block diagram of a text classification device according to this application.
[0021] Figure 5 This is a structural block diagram of a device according to this application. Detailed Implementation
[0022] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] The inventors discovered that in many cases, when users encounter problems, they communicate with service platform staff via text to resolve their issues (meet their needs). Therefore, the inventors collected historical communication texts between a large number of users and service platform staff, and conducted extensive statistical analysis on these texts. They found that the content of these historical communication texts often reveals user preference categories (what types of services users like and dislike, etc.).
[0024] In light of this, to obtain user preference classifications, the inventors conceived of using historical communication texts between users and service platform staff. Specifically, they devised a method where a model can be used to categorize user preferences based on these historical communication texts. For example, a classification model can be used to categorize historical communication texts to obtain classification labels (e.g., sentiment classification of historical communication texts to obtain sentiment labels for the user's feelings towards the services provided by the service platform, which can include positive and negative labels). These labels can reflect what types of services the user likes and dislikes. Thus, user preference classifications can be derived from these labels, achieving the goal of obtaining user preference classifications based on classification labels from historical communication texts.
[0025] To classify historical communication texts and obtain classification labels, a classification model needs to be trained beforehand. Then, when text classification is required, the text (e.g., historical communication texts between users and service platform staff) can be input into the classification model so that it can process the text, obtain classification labels, and output the labels. However, after conducting statistical analysis on the results of classifying a large number of different texts using the classification model, the inventors found that the results of text classification using the model are often inaccurate.
[0026] Thus, the need arose to improve the accuracy of text classification using classification models. To achieve this goal, the inventors conducted statistical analysis on the training text used to train the classification model and the online text requiring classification processed by the model after its deployment. The analysis revealed that:
[0027] On the one hand, the way the training text is expressed sometimes differs from that of the online text. After in-depth analysis, the inventors discovered the reasons why the way the training text is expressed sometimes differs from that of the online text.
[0028] For example, training texts are often selected (e.g., manually selected). Each selected text has its own unique expression / style. However, since the number of training texts is limited (an unlimited number is usually not possible), sometimes the selected training texts may only exhibit a few specific expressions / styles. That is, the selected training texts do not represent all objectively existing expressions / styles. Online texts, on the other hand, are often unselected and are input into classification models to predict class labels. The expressions / styles of online texts are often uncontrollable; any given online text may represent any of the objectively existing expressions / styles. In other words, the expressions / styles of online texts may not be the specific few mentioned above (i.e., not the selected training texts), but may exhibit various expressions / styles. This makes it possible for the expressions / styles of online texts to differ from those of the training texts. When the online text is expressed differently from the training text, the classification model may not be accurate, resulting in low accuracy.
[0029] In one example, the way a text is presented / expressed can include its structural components (subject, predicate, object, attributive, adverbial, and complement, etc.), the order of these components (subject, predicate, object, attributive, adverbial, and complement, etc.), and the vocabulary that contributes significantly to classification (e.g., adjectives). For instance, in a sentiment classification scenario, words describing positive sentiment in training text that contribute significantly to sentiment classification often include: good, excellent, or like. However, words describing positive sentiment in online text that contribute significantly to sentiment classification might include: great, outstanding, and awesome. "Good, excellent, or like" is different from "great, outstanding, and awesome."
[0030] Because the training text may differ from the online text, and since the online text was not considered during the training of the classification model, the model may have low processing power for the online text, resulting in low robustness and effectiveness, and consequently, low accuracy in text classification.
[0031] On the other hand, the number of characters in online text is sometimes less than the number of characters in training text. After in-depth analysis, the inventors discovered the reason why the number of characters in online text is less than that in training text. For example, training text is often selected (e.g., manually selected). Each selected text has at least a certain number of characters; that is, the selected training text does not include text with fewer than a certain number of characters. However, online text is often not selected and is often input into classification models.
[0032] To predict category labels, the number of characters in online text is often uncontrollable. The number of characters in any given online text can be arbitrary; for example, it may be greater than a certain number, equal to a certain number, or less than a certain number.
[0033] There's a certain number of characters in online text, and sometimes that number won't exceed a certain threshold. This means that the number of characters in the online text might be less than the number of characters in the training text. When the number of characters in the online text is less than the number of characters in the training text, the classification model will obtain fewer characters, resulting in less information. Consequently, the classification model's results may be inaccurate, leading to low accuracy when classifying text with zero characters.
[0034] For example, a classification model trained on texts containing more than a certain number of characters has a strong processing capability for texts with more than a certain number of characters. It can extract sufficient features (sufficient information) for classification from texts containing more than a certain number of characters, thus enabling it to classify texts containing more than a certain number of characters.
[0035] The system can accurately classify texts containing more than a certain number of characters, thus achieving high accuracy in classifying such texts. However, its processing capability is weak for texts containing fewer than a certain number of characters.
[0036] The features extracted from texts containing fewer than a certain number of characters are often insufficient (insufficient information), resulting in low accuracy for classifications with fewer than a certain number of characters.
[0037] That is, because there may be situations where the number of characters in the online text is less than the number of characters in the training text,
[0038] Furthermore, because text containing characters that are "less than the number of characters in the online text" was not used during the previous training of the classification model, the classification model is prone to misinterpreting the number of characters included.
[0039] Online texts with fewer than the number of characters in the training text have low processing power, resulting in low robustness and effectiveness of the classification model, which in turn leads to low accuracy in text classification.
[0040] In summary, the low accuracy of text classification is due to the low robustness and low robustness of the classification model. Therefore, to improve the accuracy of text classification using the model, at least one method (5) can improve the robustness and robustness of the classification model. To improve the robustness and robustness of the classification model, in order to...
[0041] In one of these methods, at least the accuracy of the trained classification model in classifying texts with more expressions / concepts can be improved, as well as the accuracy of the trained classification model in classifying texts with fewer characters. This is to improve the accuracy of the trained classification model in classifying texts with more expressions / concepts and to improve the accuracy of the trained classification model in classifying texts with fewer characters.
[0042] The model's accuracy in classifying texts with fewer characters, in at least one manner, see [reference needed]. Figure 1 This paper illustrates a method for training a classification model, which is applied in an electronic device. The method includes:
[0043] In step S101, a pre-trained language model is obtained. The original training text and its original classification labels are obtained. The original training text contains at least the original input text and template text, with the template text including prompt text and classification label fields. The original classification labels of the original training text are obtained at least based on the classification labels of the original input text.
[0044] Pre-trained language models include at least: pre-trained BERT (Bidirectional Encoder Representation from Transformers), GPT (Gererate Pre-Training Model), PLM (Pre-trained Language Model), RoBERTa (Robustly Optimized BERT Approach), and ALBERT, etc. Of course, other types of pre-trained language models may also be included, and this application does not limit this. A pre-trained language model is a neural network model that learns semantic information from a large-scale corpus in an unsupervised or supervised manner. It is a complex learning model with multiple layers of neural networks. Pre-trained language models can more accurately capture semantic information (possessing a knowledge space) in text, improving the accuracy of the model in downstream tasks.
[0045] In one embodiment of this application, obtaining the original training text and its original classification labels can be achieved through the following process:
[0046] 1011. Obtain labeled training text and labeled classification labels of labeled training text, and / or, obtain unlabeled training text and use a pre-trained language model to predict pseudo-classification labels of unlabeled training text.
[0047] In one embodiment of this application, obtaining the labeled training text and its labeled classification labels can be achieved through the following process:
[0048] 11) Get the labeled input text, get the labeled category tags of the labeled input text, and get the template text.
[0049] The labeled input text can be one or more, and each labeled input text has a labeled category label. The labeled input text can be collected from the internet, etc. The labeled category labels of the labeled input text can be manually labeled, etc. The template text is used to guide and mine the embedded knowledge of the pre-trained language model to more accurately solve the classification prediction task for the category label filling position, based on the classification prediction task and requirements for the category label filling position. The classification prediction task for the category label filling position includes: determining the category label to be filled in the category label filling position from multiple category labels based on the prompt text. The template text includes prompt text and the category label filling position. The prompt text can include natural language vocabulary, such as Chinese or English words, etc. The prompt text can be manually designed in advance according to the application scenario, etc.
[0050] For example, in a category prediction task targeting the category label input position, which is a sentiment classification task, it is necessary to predict the sentiment classification of the input text. Thus, the template text could be: "The sentiment expressed in this sentence is []". Here, "The sentiment expressed in this sentence is" is the prompt text, and the position of "[]" within the template text "The sentiment expressed in this sentence is []" is the category label input position. "[]" is used to mark the category label input position.
[0051] The multiple category labels include at least two preset category labels. The number of category labels and the specific category labels can be determined according to the actual situation. For example, in a binary emotional scenario, the binary emotions include positive emotions and negative emotions. The category labels corresponding to positive emotions can include "like" or "good", etc., and the category labels corresponding to negative emotions can include "dislike" or "bad", etc.
[0052] In this application, the classification label input field corresponds to the classification label and is used to guide the pre-trained language model to predict the classification label to be output at the classification label input field during the optimization process. The input data of the pre-trained language model can be training data containing input text and template text, etc., and the output is: the classification label of the input text predicted in each classification label according to the prompt text in the template text, which is used to fill in the classification label input field. The output classification label serves as the supervised target to optimize the pre-trained language model for the classification prediction task of the classification label input field. For example, it can fine-tune the parameters of the pre-trained language model to optimize the parameters of the pre-trained language model until the parameters in the pre-trained language model converge, thereby obtaining the classification model.
[0053] For example, suppose the original input text is "This job is done very well," and we need to perform sentiment classification on the original input text "This job is done very well." The template text is: The sentiment expressed by this sentence is []. Here, "The sentiment expressed by this sentence is" is the prompt text, and the position of "[]" in the template text "The sentiment expressed by this sentence is []" is the classification label input position. "[]" is used to mark the classification label input position. Therefore, we add the template text "The sentiment expressed by this sentence is []" after the original input text "This job is done very well." This gives us the original training text "This job is done very well, the sentiment expressed by this sentence is []". Then, the original training text "This work is done very well, the emotion expressed in this sentence is []" is input into the pre-trained language model so that the pre-trained language model can predict the classification label of the original input text to be filled in the classification label filling position "[]" in each classification label according to the prompt text in the template text. Based on the predicted classification label and the original classification label labeled in the original input text "This work is done very well", the pre-trained language model is optimized for the classification prediction task of the classification label filling position, so as to optimize the network parameters in the pre-trained language model.
[0054] In this embodiment, template text including prompt text and category label fill-in positions can be added to the original input text to obtain the original training text. The category labels of the original input text are used as the original category labels (supervision targets) of the original training text. The pre-trained language model is optimized for the category prediction task of category label fill-in positions to obtain a classification model, so that the classification model performs better (prediction accuracy) in the category prediction task of category label fill-in positions.
[0055] 12) Generate labeled training text with at least labeled input text and template text based on the preset positional relationship between the input text and the template text.
[0056] The preset positional relationship between the input text and the template text can be pre-set, and can include: the input text is placed before the template text, or the input text is placed after the template text, etc.
[0057] In one example, the input text and template text can be combined according to the preset positional relationship between the input text and the template text to obtain labeled training text.
[0058] 13) Obtain at least the labeled category labels of the labeled training text based on the labeled category labels of the labeled input text.
[0059] In one example, the labeled category labels of the labeled input text can be used as the labeled category labels of the labeled training text, etc.
[0060] Accordingly, in another embodiment, when obtaining unlabeled training text, unlabeled input text and template text can be obtained. Then, based on a preset positional relationship between the input text and the template text, unlabeled training text containing at least the unlabeled input text and the template text is generated (the specific generation method can be found in the description of step 12 above, and will not be detailed here). The unlabeled input text may not be labeled with a category tag, thus, the unlabeled training text may not be labeled with a category tag. There can be multiple unlabeled training texts, for example, the number of unlabeled training texts may be greater than the number of labeled training texts.
[0061] In another embodiment of this application, when using a pre-trained language model to predict pseudo-classification labels for unlabeled training text, the unlabeled training text can be input into the pre-trained language model so that the pre-trained language model processes the unlabeled training text to obtain the predicted probabilities of filling in each classification label in the classification label filling position of the template text in the unlabeled training text. The classification label corresponding to the highest probability is used as the pseudo-classification label of the unlabeled training text. Each classification label includes at least two preset classification labels. The number of classification labels and the specific types of each classification label can be determined according to the actual situation. For example, in a binary classification scenario of emotion, the binary emotion includes positive emotion and negative emotion. The classification labels corresponding to positive emotion can include "like" or "good," etc., and the classification labels corresponding to negative emotion can include "dislike" or "bad," etc.
[0062] To improve the accuracy of the trained classification model in classifying text, in another embodiment, the following processing can be performed in the stage of "predicting pseudo-classification labels of unlabeled training text using a pre-trained language model": For example, the pre-trained language model often includes multiple network layers. In two adjacent network layers, the feature matrix output by the previous network layer can be input into the next network layer so that the next network layer can process the feature matrix, etc.
[0063] For at least one pair of adjacent network layers in a pre-trained language model, after the preceding network layer outputs the feature matrix corresponding to the unlabeled training text, at least some of the feature elements in the feature matrix are hidden to obtain a hidden feature matrix (for example, setting at least some of the feature elements in the feature matrix to specific numbers, such as the number 0, so that at least some feature elements are hidden). The hidden feature matrix is then input into the following network layer of the two adjacent network layers for the following network layer to process the hidden feature matrix. This process continues until the last layer of the pre-trained language model outputs the pseudo-classification label of the unlabeled training text.
[0064] In this embodiment, by hiding at least some of the feature elements in the feature matrix, the pre-trained language model obtains fewer feature elements. This allows the pre-trained language model to focus more on learning the relevant information of the feature elements other than at least some of the feature elements in the feature matrix. As a result, the trained classification model has a stronger ability to extract relevant information from the feature matrix with fewer elements, and the trained classification model can achieve higher accuracy in classifying texts with feature matrices containing fewer elements. This improves the robustness and effectiveness of the trained classification model.
[0065] 1012. Obtain the original training text and its original classification label based on the labeled training text and its labeled classification label, and / or obtain the original training text and its original classification label based on the unlabeled training text and its pseudo-classification label.
[0066] The original training text can be labeled training text, the labeled classification labels of the labeled training text can be used as the original classification labels of the original training text, and / or, unlabeled training text can be used as the original training text, and the pseudo-classification labels of the unlabeled training text predicted by the pre-trained language model can be used as the original classification labels of the original training text. Thus, the original training text can include labeled training text and / or unlabeled training text. When the original training text includes labeled training text, the original classification label of the labeled training text can be: the labeled classification label of the labeled training text. When the original training text includes unlabeled training text, the original classification label of the unlabeled training text can be: the pseudo-classification label of the unlabeled training text predicted by the pre-trained language model. When the original training text includes both labeled and unlabeled training text, the original classification label of the labeled training text, which is a part of the original training text, can be: the labeled classification label of the labeled training text. Furthermore, the original classification labels of the unlabeled training texts, which constitute another part of the original training texts, can be pseudo-classification labels of the unlabeled training texts predicted using a pre-trained language model. In cases where the number of labeled training texts is small, this embodiment can collect more unlabeled training texts and automatically expand the collection with more classification labels to optimize the pre-trained language model based on these unlabeled texts and the pre-trained language model. This reduces the cost (e.g., manual labor) and workload of labeling training texts, and lowers the difficulty of obtaining a larger number of labeled training texts.
[0067] In step S102, masked training text and its masked classification label are obtained based on the original training text and its original classification label. The masked training text contains at least masked input text and template text. The masked input text is obtained by masking at least one character in the original training text. The masked classification label is obtained at least based on the original classification label of the original training text and / or at least based on at least one masked character in the original input text.
[0068] In one embodiment of this application, the mask training text includes a first mask training text, which contains at least a first mask input text and template text. The first mask input text is obtained by masking at least one character in the original training text. The first mask classification label of the first mask training text is obtained at least based on the original classification label of the original training text. For example, at least one character in the original input text can be masked in the original training text, and then the masked original training text can be used as the first mask training text. The mask can include replacing at least one character in the original input text with a preset character (e.g., a mask token). The preset character can be set according to the actual situation, such as including [XXX], (YYY), or "ZZZ", etc., which is not limited in this application. The first mask training text contains at least the unmasked content of the first mask input text, the preset character, and the template text. In addition, the original classification label of the original training text can be used as the first mask classification label of the first mask training text.
[0069] In another embodiment of this application, the mask training text includes a second mask training text, which contains at least a second mask input text and template text. The second mask input text is obtained by masking at least one character in the original input text. The second mask classification label of the second mask training text is obtained based on at least one masked character in the original input text. For example, at least one character in the original input text can be masked in the original training text, and then the masked original training text can be used as the second mask training text. The mask can include replacing at least one character in the original input text with a preset character (e.g., a mask token). The preset character can be set according to the actual situation, such as including [XXX], (YYY), or "ZZZ", etc., which is not limited in this application. The second mask training text contains at least the unmasked content of the second mask input text, the preset character, and the template text. In addition, at least one masked character in the original input text can be used as the second mask classification label of the second mask training text.
[0070] In another embodiment of this application, the mask training text includes a first mask training text and a second mask training text. The first mask training text contains at least a first mask input text and a template text. The first mask input text is obtained by masking at least one character in the original training text. The first mask classification label of the first mask training text is obtained at least based on the original classification label of the original training text. The second mask training text contains at least a second mask input text and a template text. The second mask input text is obtained by masking at least one character in the original training text. The second mask classification label of the second mask training text is obtained at least based on the masked at least one character in the original input text.
[0071] In step S103, consistent training text and its consistent classification label are obtained based on the original training text and its original classification label. The consistent training text contains at least template text and consistent input text semantically related to the original input text. The characters in the consistent input text are not entirely the same as those in the original input text, and the consistent classification label is obtained at least based on the original classification label of the original training text.
[0072] In one embodiment of this application, the original input text in the original training text is text in a first language. Thus, this step can be implemented through the following process:
[0073] 1031. Translate the original input text into the second language input text at least once in the original training text to obtain the second language training text corresponding to the original training text.
[0074] For example, the original training text can be input into a first translator, causing the first translator to translate the original input text into at least a second language input text, thus obtaining the second language training text corresponding to the original training text. In one example, the second language training text has a second language input text and a template text. The languages of both the second language input text and the template text in the second language training text can be the second language, or the language of the second language input text in the second language training text can be the second language, while the language of the template text in the second language training text can be the first language. This application does not restrict which language the first language or the second language is, as long as the first language and the second language are different languages. In one example, the second language training text has a second language input text and a template text. The positional relationship between the second language input text and the template text in the second language training text is the same as the positional relationship between the original input text and the template text in the original training text.
[0075] 1032. In the training text of the second language, at least the input text of the second language is translated into the input text of the first language to obtain the training text of the first language corresponding to the training text of the second language.
[0076] For example, training text in a second language can be input into a second translator, so that the second translator translates the second language input text into the first language input text within the second language training text, thus obtaining the first language training text corresponding to the second language training text. The first and second translators can be existing translators, etc., and can be the same translator or different translators. The positional relationship between the first language input text and the template text in the first language training text is the same as the positional relationship between the second language input text and the template text in the second language training text. After "translating the second language input text into the first language input text within the second language training text to obtain the first language training text corresponding to the second language training text," it can be determined whether the characters in the first language input text within the first language training text are completely identical to the characters in the original input text of the original training text. If the characters in the first language input text within the first language training text are not completely identical to the characters in the original input text of the original training text, step 1033 is executed. Alternatively, if the characters in the first language input text of the training text in the first language are exactly the same as the characters in the original input text of the original training text, the second language training text is input into another translator, so that the other translator translates the second language input text into the first language input text in the second language training text, and obtains another first language training text corresponding to the second language training text, until the characters in the first language input text of the first language training text are not exactly the same as the characters in the original input text of the original training text, and then step 1033 is executed.
[0077] 1033. Obtain consistent training text based on the training text in the first language, and obtain consistent classification labels for the consistent training text based at least on the original classification labels of the original training text.
[0078] For example, the training text in the first language can be used as the consistency training text, and the original classification label of the original training text can be used as the consistency classification label of the consistency training text. The forward and reverse translation processes—"translating the original input text into the second language input text at least once in the original training text to obtain the second language training text corresponding to the original training text" and "translating the second language input text into the first language input text in the second language training text to obtain the first language training text corresponding to the second language training text"—do not significantly alter the semantics of the text or its classification. That is, the original classification label of the original training text can be the same as the classification label obtained by "translating the original training text in the first language into the second language training text, and then translating it back into the second language training text to obtain the first language training text."
[0079] In text classification scenarios, the words that contribute most to the classification are often adjectives (adjectives often reflect the main idea of the text, which in turn reflects its classification). Therefore, we can identify the adjectives in the original input text within the original training text, then obtain their synonyms or near-synonyms (using a dictionary lookup), and replace these near-synonyms with the original input text to obtain the near-synonymous training text. The near-synonymous training text has near-synonymous input text and template text. The words in the original input text other than the adjective are the same as those in the near-synonymous input text, excluding their near-synonyms or near-synonyms. We can then obtain consistent training text based on the near-synonymous training text, and at least obtain consistent classification labels for the consistent training text based on the original classification labels of the original training text. For example, we can use near-synonymous training text as consistent training text, and the original classification labels of the original training text as consistent classification labels for the consistent training text.
[0080] In text classification scenarios, the words that contribute most to the classification are often adjectives (adjectives often reflect the main idea of the text, which in turn reflects its classification). Therefore, adjectives in the text often determine the text's classification. Since adjectives in the original input text of the training text are often synonyms or near-synonyms of those adjectives, the influence of adjectives on the classification of the original input text is often the same as the influence of near-synonyms on the classification of near-synonyms in the near-synonyms training text. Thus, near-synonyms can serve as consistency training texts, and their classification labels can be the original classification labels of the original training texts. In other words, the consistency classification label of a near-synonyms training text can be the same as the original classification label of the original training text.
[0081] The understanding of consistent input text related to the semantics of the original input text includes, but is not limited to: the semantics of the original input text being the same as or very similar to the semantics of the consistent input text, and the classification labels of the original input text being the same as the classification labels of the consistent input text. For example, in the scenario of sentiment classification, the classification labels of the original input text and the consistent input text are both positive sentiment or both are negative sentiment, etc.
[0082] The languages mentioned in this application may include English, Russian, German, French, Spanish, Japanese, and Chinese, etc. Of course, other languages may also be included depending on the actual situation, which will not be detailed here.
[0083] This application does not limit the execution order between steps S102 and S103.
[0084] In step S104, at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels are used to optimize the pre-trained language model for the classification prediction task of the classification label filling position, so as to obtain the classification model.
[0085] In one embodiment, at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels are used to perform at least one round of optimization learning on the pre-trained language model for the classification prediction task of classification label filling positions, so as to obtain a classification model.
[0086] Any one of the optimization learning processes includes:
[0087] 1041. Obtain the KL divergence (Kullback-Leibler divergence, relative entropy) loss value of the language model based at least on the original training text, original classification labels, consistent training text, and consistent classification labels.
[0088] A language model can be used to process both the original training text and the consistent training text separately. The KL divergence loss value of the language model can then be obtained based on the processing results of the original training text and the consistent training text, the original classification labels, and the consistent classification labels. For example, the original training text can be input into a language model (e.g., the language model obtained after the previous round of optimization learning) to process the original training text and obtain the predicted original probabilities for filling in each classification label in the template text within the original training text. Similarly, the consistent training text can be input into the language model (e.g., the language model obtained after the previous round of optimization learning) to process the consistent training text and obtain the predicted consistency probabilities for filling in each classification label in the template text within the consistent training text. The KL divergence loss value of the language model can then be calculated based on the original probabilities of each classification label, the consistency probabilities of each classification label, the original classification labels, the consistent classification labels, and the KL divergence loss function.
[0089] 1042. Obtain the original cross-entropy loss value of the language model based on the original training text and the original classification labels.
[0090] A language model can be used to process the original training text, and then the original cross-entropy loss value of the language model can be obtained based on the processing results and the original classification labels. For example, the original training text can be input into a language model (such as the language model obtained after the previous round of optimization learning) to process the original training text and obtain the predicted original probabilities for filling in each classification label in the template text within the original training text. Then, the original cross-entropy loss value of the language model can be calculated based on the original probabilities of each classification label, the original classification labels, and the cross-entropy loss function.
[0091] 1043. Obtain the masked cross-entropy loss value of the language model based on the masked training text and the masked classification labels.
[0092] In one embodiment, the masked training text includes a first masked training text, which contains at least a first masked input text and a template text. The first masked input text is obtained by masking at least one character in the original training text. The first masked classification label of the first masked training text is obtained at least based on the original classification label of the original training text. Thus, when obtaining the masked cross-entropy loss value of the language model based on the masked training text and the masked classification label, the first masked cross-entropy loss value of the language model can be obtained based on the first masked training text and the first masked classification label. For example, the language model can be used to process the first masked training text, and then the first masked cross-entropy loss value of the language model can be obtained based on the processing result of the first masked training text and the first masked classification label. For example, the first masked training text can be input into a language model (e.g., a language model obtained after the previous round of optimization learning) so that the language model processes the first masked training text to obtain the predicted first mask probabilities for filling in each classification label in the classification label filling position of the template text in the first masked training text. Then, the first mask cross-entropy loss value of the language model can be calculated based on the first mask probability of each classification label, the first mask classification label, and the cross-entropy loss function.
[0093] In another embodiment, the masked training text includes a second masked training text, which contains at least a second masked input text and a template text. The second masked input text is obtained by masking at least one character in the original input text from the original training text. The second masked classification label of the second masked training text is obtained based on at least one masked character in the original input text. Thus, when obtaining the masked cross-entropy loss value of the language model based on the masked training text and the masked classification label, the second masked cross-entropy loss value of the language model can be obtained based on the second masked training text and the second masked classification label. For example, the language model can be used to process the second masked training text, and then the second masked cross-entropy loss value of the language model can be obtained based on the processing result of the second masked training text and the second masked classification label. For example, the second masked training text can be input into a language model (e.g., a language model obtained after the previous round of optimization learning) so that the language model processes the second masked training text to obtain predicted second mask probabilities for filling in each classification label in the classification label filling position of the template text in the second masked training text. Then, the second mask cross-entropy loss value of the language model can be calculated based on the second mask probability of each classification label, the second mask classification label, and the cross-entropy loss function.
[0094] In another embodiment, the masked training text includes a first masked training text and a second masked training text. The first masked training text contains at least a first masked input text and a template text. The first masked input text is obtained by masking at least one character in the original training text. The first masked classification label of the first masked training text is obtained at least based on the original classification label of the original training text. The second masked training text contains at least a second masked input text and a template text. The second masked input text is obtained by masking at least one character in the original training text. The second masked classification label of the second masked training text is obtained at least based on the masked at least one character in the original input text. Thus, when obtaining the masked cross-entropy loss value of the language model based on the masked training text and the masked classification label, the first masked cross-entropy loss value of the language model can be obtained based on the first masked training text and the first masked classification label (see the above description for details), and the second masked cross-entropy loss value of the language model can be obtained based on the second masked training text and the second masked classification label (see the above description for details). In this way, the first masked cross-entropy loss value and the second masked cross-entropy loss value can be obtained.
[0095] 1044. At least based on the KL divergence loss value, the original cross-entropy loss value, and the masked cross-entropy loss value, optimize the language model for the classification prediction task of the classification label filling position.
[0096] In one embodiment of this application, when the masked training text includes a first masked training text and a second masked training text, the masked cross-entropy loss value includes a first masked cross-entropy loss value and a second masked cross-entropy loss value. In this case, this step can be implemented through the following process, including:
[0097] 21) Obtain the comprehensive loss value based at least on the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value.
[0098] In one example, the weighted sum of the KL divergence loss, the original cross-entropy loss, the first mask cross-entropy loss, and the second mask cross-entropy loss can be used to obtain the comprehensive loss value.
[0099] For example, the KL divergence loss, the original cross-entropy loss, the first mask cross-entropy loss, and the second mask cross-entropy loss can be weighted and summed according to the following formula to obtain the comprehensive loss value:
[0100] Loss = L KL +λ1*L1+λ2*L2+λ3*L3.
[0101] In the above formula, Loss is the overall loss value, and L KL Let L1 be the KL divergence loss value, L2 be the original cross-entropy loss value, L3 be the first mask cross-entropy loss value, and L4 be the second mask cross-entropy loss value. Let λ1 be the first preset coefficient, λ2 be the second preset coefficient, and λ3 be the third preset coefficient. In one example, the sum of the first preset coefficient λ1, the second preset coefficient λ2, and the third preset coefficient λ3 can be equal to a specific value, such as 1, 1.5, or 2, etc. This application does not limit this. The first preset coefficient λ1 can be greater than or equal to 0 and less than or equal to a specific value, the second preset coefficient λ2 can be greater than or equal to 0 and less than or equal to a specific value, and the third preset coefficient λ3 can be greater than or equal to 0 and less than or equal to a specific value, etc.
[0102] 22) At least based on the comprehensive loss value, optimize the pre-trained language model for the classification prediction task of the classification label filling position.
[0103] The network parameters in the language model can be adjusted, at least based on the comprehensive loss value, for the classification prediction task targeting the label fill-in position. After at least one round of optimization learning on the pre-trained language model for the classification prediction task targeting the label fill-in position, and with the convergence of the network parameters in the language model, the latest language model is used as the classification model, and then the classification model can be deployed online.
[0104] On one hand, in the scenario of optimizing a pre-trained language model for a classification prediction task involving filling in classification labels, there is consistent training text and consistent classification labels for the consistent training text. The consistent training text contains at least template text and consistent input text semantically related to the original input text. The characters in the consistent input text are not entirely the same as those in the original input text, and the consistent classification labels are obtained at least based on the original classification labels. Furthermore, it also contains the original training text and its original classification labels.
[0105] This approach allows for consistent regularization optimization of the pre-trained language model based on the original training text, original classification labels, consistent training text, and consistent classification labels. This ensures that minor modifications to the input text do not alter the language model's prediction of the input text's classification label. It can improve the accuracy of the trained classification model in classifying texts with more diverse expressions / concepts, enhance the robustness and effectiveness of the trained classification model, and ultimately improve its accuracy in text classification.
[0106] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained at least based on the original classification label in the original training text.
[0107] This approach allows for the optimization and learning of a pre-trained language model based on the masked training text and its masked classification labels. This enables the pre-trained language model to learn from training texts with fewer characters, allowing it to focus more intently on learning relevant information from these smaller texts. Consequently, the trained classification model demonstrates stronger ability to extract relevant information from texts with fewer characters, improving its accuracy in classifying texts with fewer characters. This enhances the robustness and effectiveness of the trained classification model, ultimately leading to higher accuracy in text classification.
[0108] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained based on at least one masked character in the original input text.
[0109] This allows the pre-trained language model to optimize the learning of semantic relationships between text contexts based on the masked training text and the masked classification labels of the masked training text. This enables the classification model to learn the semantic and contextual relationships between words in the text, thereby improving the accuracy of the trained classification model in predicting masked words based on words other than the masked words in text containing "input text and template text (the template text contains at least prompt text and masked words, etc.)". Furthermore, it can improve the accuracy of the trained classification model in predicting the classification labels of the input text used to fill in the classification label positions based on the input text and prompt text in text containing "input text and template text (the template text contains at least prompt text and classification label filling positions, etc.)". Ultimately, this improves the accuracy of the trained classification model in text classification.
[0110] In another embodiment, in Figure 1The method shown also includes: obtaining interference training text and interference classification labels of interference training text based on the original training text and the original classification labels; the interference training text contains at least interference text, original input text and template text; the interference classification labels are obtained at least based on the original classification labels.
[0111] For example, a random text can be generated and used as noise text. This noise text, along with the original input text and template text, is combined to obtain the noise training text. The noise text can be placed before or after the original input text. The noise text can include at least one character, and the number of characters in the noise text can be less than the number of characters in the original input text. The semantics of the noise text can be related to or unrelated to the semantics of the original input text. For example, sometimes the user input text contains two sentences, one expressing the user's sentiment classification and the other not. For the classification model, the result is the user input text containing two sentences, from which the user's sentiment classification can be derived. In the scenario where the user's sentiment classification is obtained based on these two sentences, one sentence mainly contributes to the classification. Thus, the classification model primarily obtains the user's sentiment classification based on one sentence. However, the other sentence can interfere with the classification model's ability to obtain the user's sentiment classification based on the first sentence. Therefore, by interfering with the training text and interfering with the classification labels to participate in the training of the pre-trained language model, the trained classification model becomes more robust to interference. For example, it can minimize the interference of the other sentence on the classification model's classification results, so as to focus on classifying the user's sentiment based on the first sentence and improve the accuracy of classification.
[0112] Accordingly, in step S104, the original training text, original classification labels, masked training text, masked classification labels, consistent training text, consistent classification labels, interference training text, and interference classification labels can be used to optimize the pre-trained language model for the classification prediction task of the classification label filling position, so as to obtain the classification model.
[0113] Thus, the method in any round of optimization learning mentioned in step S104 also includes: obtaining the interference cross-entropy loss value of the language model based on the interference training text and interference classification labels.
[0114] For example, interfering training text is input into a language model (e.g., the language model obtained after the previous round of optimization learning) so that the language model processes the interfering training text and obtains the predicted interference probabilities for filling in the classification label fields in the template text within the interfering training text. Then, the interference cross-entropy loss value of the language model can be calculated based on the interference probabilities of each classification label, the interfering classification labels, and the cross-entropy loss function.
[0115] Accordingly, in step 1044, when optimizing the language model for the classification prediction task of the classification label fill-in position based at least on the KL divergence loss value, the original cross-entropy loss value, and the masked cross-entropy loss value, the language model can be optimized for the classification prediction task of the classification label fill-in position based on the KL divergence loss value, the original cross-entropy loss value, the masked cross-entropy loss value, and the interference cross-entropy loss value.
[0116] For example, a comprehensive loss value can be obtained based on at least the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, the second mask cross-entropy loss value, and the interference cross-entropy loss value; then, based on at least the comprehensive loss value, the pre-trained language model can be optimized for the classification prediction task of the classification label filling position.
[0117] In one example, the weighted sum of the KL divergence loss, the original cross-entropy loss, the first mask cross-entropy loss, the second mask cross-entropy loss, and the interference cross-entropy loss can be used to obtain the comprehensive loss value.
[0118] For example, the KL divergence loss, the original cross-entropy loss, the first mask cross-entropy loss, the second mask cross-entropy loss, and the interference cross-entropy loss can be weighted and summed according to the following formula to obtain the comprehensive loss value:
[0119] Loss = L KL +λ1*L1+λ2*L2+λ3*L3+λ4*L4.
[0120] In the above formula, Loss is the overall loss value, and L KLLet L1 be the KL divergence loss value, L2 be the original cross-entropy loss value, L3 be the first mask cross-entropy loss value, L4 be the interference cross-entropy loss value, and λ1 be the first preset coefficient, λ2 be the second preset coefficient, λ3 be the third preset coefficient, and λ4 be the fourth preset coefficient. In one example, the sum of the first preset coefficient λ1, the second preset coefficient λ2, the third preset coefficient λ3, and the fourth preset coefficient λ4 can be equal to a specific value, such as 1, 1.5, or 2, etc., which is not limited in this application. The first preset coefficient λ1 can be greater than or equal to 0 and less than or equal to a specific value, the second preset coefficient λ2 can be greater than or equal to 0 and less than or equal to a specific value, the third preset coefficient λ3 can be greater than or equal to 0 and less than or equal to a specific value, and the fourth preset coefficient λ4 can be greater than or equal to 0 and less than or equal to a specific value, etc.
[0121] In the scenario of optimizing a pre-trained language model for classification prediction tasks involving filling in classification labels, there is interfering training text. This interfering training text contains at least three elements: interfering text, the original input text, and template text. The interfering classification labels of the interfering training text are derived from at least the original classification labels. This allows for optimization of the pre-trained language model based on the interfering training text and its interfering classification labels. This enables the pre-trained language model to resist interference from interfering text, ensuring that even when interfering text interferes with the input text, the model's prediction of the input text's classification label remains unchanged. This improves the accuracy of the trained classification model in classifying texts with interfering text, enhancing its robustness and overall performance, ultimately leading to higher accuracy in text classification.
[0122] Furthermore, after optimizing the pre-trained language model for the classification prediction task targeting the classification label fill-in position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels, the resulting classification model can be deployed online, for example, for online text classification. See details... Figure 2 This application illustrates a method for text classification, which is applied in an electronic device and includes the following steps:
[0123] In step S201, the text to be processed is obtained. The text to be processed contains at least the text to be classified and the template text. The template text includes prompt text and classification label fields.
[0124] The text to be categorized refers to text that needs to be classified. This text can include user-inputted comments and search queries. For example, in one scenario, a user inputs a comment about a product. The electronic device can then categorize the comment and, based on the categorization (e.g., positive or negative sentiment), recommend other products to the user. Thus, the input text received by the electronic device can be the text to be categorized. Once the electronic device receives the text to be categorized, it can obtain template text, which is pre-stored in the electronic device. For a detailed description of template text, please refer to [link to relevant documentation]. Figure 1 The illustrated embodiment will not be described in detail here. Then, text to be processed, containing at least the text to be classified and the template text, can be generated based on the preset positional relationship between the text to be classified and the template text. For example, the preset positional relationship between the text to be classified and the template text can be pre-set, and may include: the text to be classified is located before the template text, or the text to be classified is located after the template text, etc. The preset positional relationship between the text to be classified and the template text... Figure 1 In the illustrated embodiments, the preset positional relationship between the input text and the template text can be the same. Thus, based on the preset positional relationship between the text to be classified and the template text, the text to be classified and the template text can be combined to obtain the text to be processed.
[0125] In step S202, based on the trained classification model, the classification label of the text to be processed is predicted and used to fill in the classification label filling position in the template text of the text to be processed.
[0126] The text to be processed can be input into the trained classification model, so that the trained classification model can perform a classification prediction task for the classification label filling position in the template text of the text to be processed, and output the classification label with the highest probability (the output classification label is one of at least two preset classification labels). The output classification label can be understood as the classification label of the text to be processed.
[0127] In step S203, the classification labels of the text to be classified are obtained based on the classification labels of the text to be processed.
[0128] The classification labels output by the trained classification model can be understood as the classification labels of the text to be processed. The purpose of this application is to obtain the classification labels of the text to be classified. Specifically, the prompt text in the template text of the text to be processed is used to guide the trained classification model in performing the task of "predicting the classification label to be output in the classification label field of the template text in the text to be processed" during the processing of the text. The prompt text in the template text of the text to be processed does not play a decisive role in determining which classification label the text to be classified is; therefore, the classification label of the text to be processed can be used as the classification label of the text to be classified.
[0129] The trained classification model is obtained by optimizing a pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels. The original training text contains at least the original input text and template text, with the template text including prompt text and classification label filling positions. The original classification labels of the original training text are obtained based on at least the classification labels of the original input text. The masked training text contains at least the masked input text and template text. The masked input text is obtained by masking at least one character in the original input text, and the masked classification labels are obtained based on at least the original classification labels of the original training text and / or at least based on at least one masked character in the original input text. The consistent training text contains template text and consistent input text semantically related to the original input text. The characters in the consistent input text are not all the same as those in the original input text, and the consistent classification labels are obtained based on at least the original classification labels. For specific training methods, please refer to [link to training documentation]. Figure 1 The embodiments shown are not described in detail here.
[0130] The trained classification model is obtained by optimizing the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels. This results in a higher accuracy of the trained classification model for classifying texts with more expressions / concepts, a higher accuracy of classifying texts with fewer characters, and a higher accuracy of predicting masked words based on words other than the masked words in texts containing "input text and template text (the template text contains at least prompt text and masked words)". This makes the trained classification model more robust and effective, thus improving the accuracy of text classification.
[0131] In another embodiment, the trained classification model is obtained by optimizing a pre-trained language model for a classification prediction task targeting the classification label filling position, using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, consistent classification labels, interference training text, and interference classification labels. This results in a higher accuracy of the trained classification model for classifying texts with more expressions / concepts, a higher accuracy of classifying texts with fewer characters, a higher accuracy of predicting masked words based on words other than the masked words in texts containing "input text and template text (the template text contains at least prompt text and masked words, etc.)," and a higher accuracy of classifying texts with interference text. This makes the trained classification model more robust and effective, thus improving the accuracy of text classification.
[0132] The scenarios for text classification in this application may include sentiment classification of text, classification of relationships between entities in text, and classification of user intent reflected in text, etc. Of course, other scenarios may also be included, and this application does not limit them.
[0133] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0134] Reference Figure 3The diagram illustrates a structural block diagram of an apparatus for training a classification model, comprising: a first acquisition module 11, used to acquire a pre-trained language model; acquire original training text and original classification labels of the original training text, wherein the original training text contains at least original input text and template text, and the template text includes prompt text and classification label fields; the original classification labels of the original training text are obtained at least based on the classification labels of the original input text; a second acquisition module 12, used to acquire masked training text and masked classification labels of the masked training text based on the original training text and original classification labels; the masked training text contains at least masked input text and template text, wherein the masked input text is obtained by masking at least one character in the original input text, and the masked classification labels are obtained at least based on the original classification labels of the original training text. The class label is obtained and / or at least based on at least one masked character in the original input text; the third acquisition module 13 is used to obtain consistent training text and consistent classification labels of consistent training text based on the original training text and the original classification label; the consistent training text has at least template text and consistent input text that is semantically related to the original input text; the characters in the consistent input text are not all the same as the characters in the original input text, and the consistent classification label is obtained at least based on the original classification label; the optimization module 14 is used to optimize the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, the original classification label, the masked training text, the masked classification label, the consistent training text, and the consistent classification label, to obtain the classification model.
[0135] The training module includes: at least one optimization unit, which is used to perform at least one round of optimization learning on the pre-trained language model for the classification prediction task of classification label filling positions, using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels, to obtain a classification model; any one of the optimization units includes: a first acquisition unit, used to acquire the KL divergence loss value of the language model based on the original training text, original classification labels, consistent training text, and consistent classification labels; a second acquisition unit, used to acquire the original cross-entropy loss value of the language model based on the original training text and original classification labels; a third acquisition unit, used to acquire the masked cross-entropy loss value of the language model based on the masked training text and masked classification labels; and an optimization unit, used to perform optimization learning on the language model for the classification prediction task of classification label filling positions, using at least the KL divergence loss value, the original cross-entropy loss value, and the masked cross-entropy loss value.
[0136] The masked training text includes a first masked training text and a second masked training text; the first masked training text contains at least a first masked input text and a template text, the first masked input text is obtained by masking at least one character in the original input text in the original training text, and the first masked classification label of the first masked training text is obtained at least based on the original classification label; the second masked training text contains at least a second masked input text and a template text, the second masked input text is obtained by masking at least one character in the original input text in the original training text, and the second masked classification label of the second masked training text is obtained at least based on at least one masked character in the original input text; the third acquisition unit includes: a first acquisition subunit, used to acquire the first masked cross-entropy loss value of the language model based on the first masked training text and the first masked classification label; and a second acquisition subunit, used to acquire the second masked cross-entropy loss value of the language model based on the second masked training text and the second masked classification label.
[0137] The optimization unit includes: a third acquisition subunit, used to acquire a comprehensive loss value based at least on the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value; and an optimization subunit, used to optimize the pre-trained language model for the classification prediction task of the classification label filling position based at least on the comprehensive loss value.
[0138] Specifically, the third acquisition subunit is used to: at least weight and sum the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value to obtain the comprehensive loss value.
[0139] The first acquisition module includes: a fourth acquisition unit for acquiring labeled training text and its labeled classification labels; a fifth acquisition unit for acquiring original training text and its original classification labels based on the labeled training text and its labeled classification labels; and / or a sixth acquisition unit for acquiring unlabeled training text; a prediction unit for predicting pseudo-classification labels of unlabeled training text using a pre-trained language model; and a seventh acquisition unit for acquiring original training text and its original classification labels based on the unlabeled training text and the pseudo-classification labels.
[0140] The fourth acquisition unit includes: a fourth acquisition subunit, used to acquire labeled input text, acquire labeled category labels of the labeled input text, and acquire template text; a first generation subunit, used to generate labeled training text with at least labeled input text and template text based on a preset positional relationship between the input text and the template text; and a fifth acquisition subunit, used to acquire labeled category labels of the labeled training text based at least on the labeled category labels of the labeled input text.
[0141] The fifth acquisition unit includes: a sixth acquisition subunit, used to acquire unlabeled input text and template text; and a second generation subunit, used to generate unlabeled training text with at least unlabeled input text and template text based on a preset positional relationship between the input text and the template text.
[0142] The prediction unit includes: a first input subunit for inputting unlabeled training text into a pre-trained language model; a hiding subunit for hiding at least some feature elements in at least one pair of adjacent network layers in the pre-trained language model after the first network layer outputs the feature matrix corresponding to the unlabeled training text, to obtain a hidden feature matrix; a second input subunit for inputting the hidden feature matrix into the next network layer of the two adjacent network layers, so that the next network layer of the two adjacent network layers can process the hidden feature matrix, and so on, until the last layer of the pre-trained language model outputs the pseudo-classification label of the unlabeled training text.
[0143] The original input text in the original training text is text in the first language; the third acquisition module includes: a first translation unit, used to translate the original input text into at least the input text in the second language in the original training text, to obtain the training text in the second language corresponding to the original training text; a second translation unit, used to translate the input text in the second language into at least the input text in the first language in the training text in the second language, to obtain the training text in the first language corresponding to the training text in the second language; the characters in the input text in the first language are not all the same as the characters in the original input text; and an eighth acquisition unit, used to acquire consistent training text based on the training text in the first language, and to acquire consistent classification labels based at least on the original classification labels.
[0144] On one hand, in the scenario of optimizing a pre-trained language model for a classification prediction task involving filling in classification labels, there is consistent training text and consistent classification labels for the consistent training text. The consistent training text contains at least template text and consistent input text semantically related to the original input text. The characters in the consistent input text are not entirely the same as those in the original input text, and the consistent classification labels are obtained at least based on the original classification labels. Furthermore, it also contains the original training text and its original classification labels.
[0145] This approach allows for consistent regularization optimization of the pre-trained language model based on the original training text, original classification labels, consistent training text, and consistent classification labels. This ensures that minor modifications to the input text do not alter the language model's prediction of the input text's classification label. It can improve the accuracy of the trained classification model in classifying texts with more diverse expressions / concepts, enhance the robustness and effectiveness of the trained classification model, and ultimately improve its accuracy in text classification.
[0146] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained at least based on the original classification label in the original training text.
[0147] This approach allows for the optimization and learning of a pre-trained language model based on the masked training text and its masked classification labels. This enables the pre-trained language model to learn from training texts with fewer characters, allowing it to focus more intently on learning relevant information from these smaller texts. Consequently, the trained classification model demonstrates stronger ability to extract relevant information from texts with fewer characters, improving its accuracy in classifying texts with fewer characters. This enhances the robustness and effectiveness of the trained classification model, ultimately leading to higher accuracy in text classification.
[0148] On the other hand, in the scenario of optimizing the pre-trained language model for the classification prediction task of filling in the classification label, there is a masked training text, which contains at least a masked input text and a template text. The masked input text is obtained by masking at least one character in the original input text in the original training text, and the masked classification label of the masked training text is obtained based on at least one masked character in the original input text.
[0149] This allows the pre-trained language model to optimize the learning of semantic relationships between text contexts based on the masked training text and the masked classification labels of the masked training text. This enables the classification model to learn the semantic and contextual relationships between words in the text, thereby improving the accuracy of the trained classification model in predicting masked words based on words other than the masked words in text containing "input text and template text (the template text contains at least prompt text and masked words, etc.)". Furthermore, it can improve the accuracy of the trained classification model in predicting the classification labels of the input text used to fill in the classification label positions based on the input text and prompt text in text containing "input text and template text (the template text contains at least prompt text and classification label filling positions, etc.)". Ultimately, this improves the accuracy of the trained classification model in text classification.
[0150] Furthermore, the device also includes: a fourth acquisition module, used to acquire interference training text and interference classification labels of interference training text based on the original training text and the original classification labels; the interference training text contains at least interference text, original input text, and template text; the interference classification labels are obtained at least based on the original classification labels; the optimization module is specifically used to: use the original training text, original classification labels, masked training text, masked classification labels, consistent training text, consistent classification labels, interference training text, and interference classification labels to optimize the pre-trained language model for the classification prediction task of classification label filling positions, thereby obtaining a classification model.
[0151] Furthermore, in any round of optimization learning, the training module also includes: a ninth acquisition unit, used to acquire the interference cross-entropy loss value of the language model based on the interference training text and interference classification labels; correspondingly, the optimization unit is specifically used to: optimize the language model for the classification prediction task of the classification label filling position based on the KL divergence loss value, the original cross-entropy loss value, the mask cross-entropy loss value and the interference cross-entropy loss value.
[0152] In the scenario of optimizing a pre-trained language model for a classification prediction task involving filling in classification labels, there is interfering training text. This interfering training text contains at least three elements: interfering text, the original input text, and template text. The interfering classification labels of the interfering training text are derived from at least the original classification labels. This allows for optimization of the pre-trained language model based on the interfering training text and its interfering classification labels. This enables the pre-trained language model to resist interference from interfering text, ensuring that even when interfering text interferes with the input text, the model's prediction of the input text's classification label remains unchanged. This improves the accuracy of the trained classification model in classifying texts with interfering text, enhancing its robustness and overall performance, ultimately leading to higher accuracy in text classification.
[0153] Reference Figure 4 The diagram shows a structural block diagram of a text classification apparatus according to this application, comprising: a fourth acquisition module 21, for acquiring a text to be processed, the text to be processed containing at least a text to be classified and a template text, the template text including prompt text and a classification label filling position; a prediction module 22, for predicting, based on a trained classification model, the classification label of the text to be processed to be filled in the classification label filling position in the template text of the text to be processed; and a fifth acquisition module 23, for acquiring the classification label of the text to be classified according to the classification label of the text to be processed.
[0154] The fourth acquisition module includes: a tenth acquisition unit, used to acquire template text upon receiving input text to be classified; and a generation unit, used to generate text to be processed that contains at least the text to be classified and the template text, based on a preset positional relationship between the text to be classified and the template text.
[0155] The trained classification model is obtained by optimizing the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels. The original training text contains at least the original input text and template text, and the template text includes prompt text and classification label filling positions. The original classification labels of the original training text are obtained at least based on the classification labels of the original input text. The masked training text contains at least the masked input text and template text. The masked input text is obtained by masking at least one character in the original input text. The masked classification labels are obtained at least based on the original classification labels of the original training text and / or at least based on at least one masked character in the original input text. The consistent training text contains template text and consistent input text that is semantically related to the original input text. The characters in the consistent input text are not all the same as the characters in the original input text. The consistent classification labels are obtained at least based on the original classification labels.
[0156] The trained classification model is obtained by optimizing the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels. This results in a higher accuracy of the trained classification model for classifying texts with more expressions / concepts, a higher accuracy of classifying texts with fewer characters, and a higher accuracy of predicting masked words based on words other than the masked words in texts containing "input text and template text (the template text contains at least prompt text and masked words)". This makes the trained classification model more robust and effective, thus improving the accuracy of text classification.
[0157] In another embodiment, the trained classification model is obtained by optimizing a pre-trained language model for a classification prediction task targeting the classification label filling position, using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, consistent classification labels, interference training text, and interference classification labels. This results in a higher accuracy of the trained classification model for classifying texts with more expressions / concepts, a higher accuracy of classifying texts with fewer characters, a higher accuracy of predicting masked words based on words other than the masked words in texts containing "input text and template text (the template text contains at least prompt text and masked words, etc.)," and a higher accuracy of classifying texts with interference text. This makes the trained classification model more robust and effective, thus improving the accuracy of text classification.
[0158] This application also provides a non-volatile readable storage medium storing one or more programs. When these programs are used in a device, they enable the device to execute this application.
[0159] Instructions for each method step in the embodiment.
[0160] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more methods as described in the above embodiments. In this application, the electronic device includes a server, a gateway, sub-devices, etc., and the sub-devices are devices such as Internet of Things (IoT) devices.
[0161] The embodiments disclosed herein can be implemented using any suitable hardware, firmware, software, or any combination thereof.
[0162] The device that performs the desired configuration may include servers (clusters), terminal devices such as IoT devices, and other electronic devices.
[0163] Figure 5 An exemplary apparatus 1300 that can be used to implement the various embodiments of this application is schematically illustrated. For one embodiment, Figure 5 An exemplary device 1300 is shown, which has one or more processors 1302 coupled together.
[0164] The processor 1302 includes at least one of a control module (chipset) 1304, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304. The processor 1302 may include one or more single-core or multi-core processors.
[0165] It may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1300 can function as a server device such as a gateway in the embodiments of this application.
[0166] 0 In some embodiments, the device 1300 includes one or more computer-readable media having instructions 1314 (e.g.,
[0167] For example, a memory 1306 or an NVM / storage device 1308) and one or more processors 1302 that are combined with the one or more computer-readable media and configured to execute instructions 1314 to implement the module and thus perform the actions in this disclosure.
[0168] In one embodiment, the control module 1304 may include any suitable interface controller to send commands to (one or more) interfaces.
[0169] At least one of the processors 1302 and / or any suitable device or component communicating with the control module 1304 provides any suitable interface. The control module 1304 may include a memory controller module to provide an interface to the memory 1306.
[0170] The memory controller module can be a hardware module, a software module, and / or a firmware module. Memory 1306 can be used, for example, to load and store data and / or instructions 1314 for device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1306 may include double data rate quad synchronous dynamic random access memory (DDR4 SDRAM).
[0171] 0. In one embodiment, the control module 1304 may include one or more input / output controllers to send signals to the NVM / store.
[0172] Storage device 1308 and one or more input / output devices 1310 provide interfaces. For example, NVM / storage device 1308 may be used to store data and / or instructions 1314. NVM / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable one or more non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives). NVM / storage device 1308 may include storage resources that are physically part of a device on which device 1300 is mounted, or that can be accessed by the device without being part of the device. For example, NVM / storage device 1308 may be accessed via a network via one or more input / output devices 1310. One or more input / output devices 1310 may provide interfaces for device 1300 to communicate with any other suitable device, and input / output devices 1310 may include communication components, input components, sensor components, etc. Network interface 1312 may provide an interface for device 1300 to communicate through one or more networks. Device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.
[0173] In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1304. In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers of the control module 1304 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die to form a system-on-a-chip (SoC).
[0174] In various embodiments, device 1300 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1300 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0175] This application provides an electronic device, including: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform one or more methods as described in this application.
[0176] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0177] The embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The embodiments of this application are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable information processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable information processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable information processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 a process or multiple processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable information processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0178] Although preferred embodiments of the embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of this application. Finally, it should be noted that in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element. The foregoing has provided a detailed description of the training classification model, text classification method, and apparatus provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for training a classification model, characterized in that, The method includes: Obtain the pre-trained language model; obtain the original training text and its original classification labels. The original training text contains at least the original input text and template text, and the template text includes prompt text and classification label fields. The original classification labels of the original training text are obtained based on the classification labels of the original input text. Obtain masked training text and masked classification labels of masked training text based on original training text and original classification labels; the masked training text contains at least masked input text and template text, the masked input text is obtained by masking at least one character in the original input text, and the masked classification labels are obtained at least based on the original classification labels and / or at least based on at least one masked character in the original input text. Obtain consistent training text and consistent classification labels from the original training text and original classification labels; the consistent training text must contain at least template text and consistent input text that is semantically related to the original input text; the characters in the consistent input text are not all the same as the characters in the original input text, and the consistent classification labels are obtained at least based on the original classification labels; At least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels are used to optimize the pre-trained language model for the classification prediction task of classification label filling positions, so as to obtain a classification model.
2. The method according to claim 1, characterized in that, The process involves optimizing a pre-trained language model for a classification prediction task targeting the classification label fill-in position, using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels, to obtain a classification model, including: At least one round of optimization learning is performed on the pre-trained language model for the classification prediction task of the classification label filling position using the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels to obtain a classification model. Any one of these rounds of optimization learning includes: The KL divergence loss value of the language model is obtained based on the original training text, the original classification labels, the consistent training text, and the consistent classification labels. Obtain the original cross-entropy loss value of the language model based on the original training text and the original classification labels; Obtain the masked cross-entropy loss value of the language model based on the masked training text and the masked classification labels; Based on at least the KL divergence loss, the original cross-entropy loss, and the masked cross-entropy loss, the language model is optimized for the classification prediction task of the classification label filling position.
3. The method according to claim 2, characterized in that, The mask training text includes a first mask training text and a second mask training text; The first mask training text contains at least a first mask input text and a template text. The first mask input text is obtained by masking at least one character in the original input text in the original training text. The first mask classification label of the first mask training text is obtained at least based on the original classification label. The second mask training text contains at least a second mask input text and a template text. The second mask input text is obtained by masking at least one character in the original input text in the original training text. The second mask classification label of the second mask training text is obtained based on at least one masked character in the original input text. The step of obtaining the masked cross-entropy loss value of the language model based on the masked training text and the masked classification labels includes: The first mask cross-entropy loss value of the language model is obtained based on the first mask training text and the first mask classification label; Additionally, the second mask cross-entropy loss value of the language model is obtained based on the training text and the classification label of the second mask.
4. The method according to claim 3, characterized in that, The optimization learning of the language model for the classification prediction task of the classification label fill-in position, based at least on the KL divergence loss value, the original cross-entropy loss value, and the masked cross-entropy loss value, includes: The comprehensive loss value is obtained based on at least the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value. Based at least on the aforementioned comprehensive loss value, the pre-trained language model is optimized for the classification prediction task of the classification label filling position.
5. The method according to claim 4, characterized in that, The process of obtaining the comprehensive loss value based at least on the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value includes: The combined loss value is obtained by weighted summation of at least the KL divergence loss value, the original cross-entropy loss value, the first mask cross-entropy loss value, and the second mask cross-entropy loss value.
6. The method according to claim 2, characterized in that, The method further includes: Obtain interfering training texts and interfering classification labels for the interfering training texts based on the original training texts and original classification labels; the interfering training texts must contain at least interfering texts, the original input text, and template text; the interfering classification labels are obtained at least based on the original classification labels; Accordingly, the process of optimizing the pre-trained language model for the classification prediction task of classifying label fill-in positions using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels to obtain a classification model includes: Using the original training text, original classification labels, masked training text, masked classification labels, consistent training text, consistent classification labels, interference training text, and interference classification labels, the pre-trained language model is optimized for the classification prediction task of the classification label filling position to obtain the classification model. Furthermore, in any round of optimization learning, the method further includes: Obtain the interference cross-entropy loss value of the language model based on the interference training text and interference classification labels; Accordingly, the optimization learning of the language model for the classification prediction task of the classification label fill-in position, based at least on the KL divergence loss value, the original cross-entropy loss value, and the masked cross-entropy loss value, includes: Based on the KL divergence loss, the original cross-entropy loss, the masked cross-entropy loss, and the interference cross-entropy loss, the language model is optimized for the classification prediction task of the classification label filling position.
7. The method according to claim 1, characterized in that, The process of obtaining the original training text and its original classification labels includes: Obtain the labeled training text and its labeled classification labels; based on the labeled training text and its labeled classification labels, obtain the original training text and its original classification labels. And / or, Obtain unlabeled training text and use a pre-trained language model to predict pseudo-classification labels for the unlabeled training text; obtain the original training text and its original classification label based on the unlabeled training text and the pseudo-classification labels.
8. The method according to claim 7, characterized in that, The process of obtaining labeled training text and its labeled classification labels includes: Get the labeled input text, get the labeled category tags of the labeled input text, and get the template text; Generate labeled training text that contains at least the labeled input text and the template text based on the preset positional relationship between the input text and the template text; At least obtain the labeled category labels of the labeled training text based on the labeled category labels of the labeled input text.
9. The method according to claim 7, characterized in that, The method of using a pre-trained language model to predict pseudo-classification labels for unlabeled training text includes: Unlabeled training text is input into a pre-trained language model. For at least one pair of adjacent network layers in the pre-trained language model, after the first network layer of the two adjacent network layers outputs the feature matrix corresponding to the unlabeled training text, at least some feature elements in the feature matrix are hidden to obtain a hidden feature matrix. The hidden feature matrix is then input into the next network layer of the two adjacent network layers for the next network layer to process the hidden feature matrix. This process continues until the last layer of the pre-trained language model outputs the pseudo-classification label of the unlabeled training text.
10. The method according to claim 1, characterized in that, The original input text in the original training text is text in the first language; The step of obtaining consistent training text and consistent classification labels for consistent training text based on original training text and original classification labels includes: In the original training text, at least the original input text is translated into the second language input text to obtain the second language training text corresponding to the original training text; In the training text of the second language, the input text of the second language is translated into the input text of the first language at least once to obtain the training text of the first language corresponding to the training text of the second language; the characters in the input text of the first language are not all the same as the characters in the original input text. Obtain consistent training text based on the training text in the first language, and obtain consistent classification labels based at least on the original classification labels.
11. A method for classifying text, characterized in that, The method includes: Obtain the text to be processed, which contains at least the text to be categorized and the template text. The template text includes prompt text and category label fields. Based on the trained classification model, predict the classification label of the text to be processed to be filled in the classification label filling position in the template text of the text to be processed; Obtain the category labels of the text to be categorized based on the category labels of the text to be processed; The trained classification model is obtained by optimizing the pre-trained language model for the classification prediction task of the classification label filling position using at least the original training text, original classification labels, masked training text, masked classification labels, consistent training text, and consistent classification labels. The original training text contains at least the original input text and template text, and the template text includes prompt text and classification label filling positions. The original classification labels of the original training text are obtained at least based on the classification labels of the original input text. The masked training text contains at least the masked input text and template text. The masked input text is obtained by masking at least one character in the original input text. The masked classification labels are obtained at least based on the original classification labels of the original training text and / or at least based on at least one masked character in the original input text. The consistent training text contains at least the template text and consistent input text that is semantically related to the original input text. The characters in the consistent input text are not all the same as the characters in the original input text. The consistent classification labels are obtained at least based on the original classification labels.
12. The method according to claim 11, characterized in that, The process of obtaining the text to be processed includes: Upon receiving the input text to be categorized, obtain the template text; Generate a text to be processed that contains at least the text to be classified and the template text, based on the preset positional relationship between the text to be classified and the template text.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The steps of the method as claimed in any one of claims 1 to 12 are implemented when the processor executes the program.
14. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as claimed in any one of claims 1 to 12.
Citation Information
Patent Citations
Text classification method and server
CN113961705A
Extracting chemical structures from digitized images
EP3876236A1