Text classification method and device, electronic equipment and storage medium

Through the two-stage training method and feature vector construction, the problem of low text classification accuracy in low resource scenarios is solved, and efficient accuracy and robustness in malicious comment recognition is achieved, which is suitable for malicious comment recognition in the Internet community.

CN120470125APending Publication Date: 2025-08-12AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510657068.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In low-resource scenarios, pre-trained language models need more training data and training costs when migrating to new scenarios, resulting in low accuracy of text classification, especially in specific scenarios that malicious comment recognition is difficult to effectively carry out.

Method used

The two-stage training method is adopted, first the first stage training is carried out in a general scenario, the model is trained using the general sample text, and then the second stage training is carried out in the target scenario. By fine-tuning the model to adapt to low-resource scenarios, data expansion is used using the prompt word technology of pre-trained language models and large language models, combining characters, pronunciations and glyph feature vectors for feature construction and model fine-tuning.

Benefits of technology

It improves the accuracy of text classification in the target scenario, especially the malicious comment recognition ability under low resource conditions, improves the recognition robustness and generalization ability of the model in specific fields, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470125A_ABST
    Figure CN120470125A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a text classification method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a to-be-classified text; determining whether the to-be-classified text belongs to a target classification or not in a target scene through a preset classification model; wherein the preset classification model is obtained based on two stages of training; the training of the two stages comprises first-stage training performed based on a first sample text in a general scene and then second-stage training performed based on a second sample text in the target scene. The accuracy of text classification in a target scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a text classification method, device, electronic device, and storage medium. Background Art

[0002] With the widespread adoption and development of the internet, content on various online platforms has exploded. Users commenting on content of interest has significantly enhanced the friendly community atmosphere and user experience. However, malicious comments can also undermine the community ecosystem. Identifying and blocking malicious comments helps maintain the community ecosystem and user experience.

[0003] With the development of pre-trained language model technology, its understanding of text content has become relatively profound. However, migrating pre-trained language models to new scenarios still requires a large amount of training data and training costs. However, whether a comment is classified as malicious is closely related to the community context. In certain scenarios, the scale of annotated training sets is relatively small, so text classification accuracy in low-resource scenarios faces certain challenges. Summary of the Invention

[0004] The embodiments of the present application provide a text classification method, device, electronic device, and storage medium, which can improve the accuracy of text classification in target scenarios.

[0005] In a first aspect, an embodiment of the present application provides a text classification method, comprising:

[0006] Get the text to be classified;

[0007] Determine whether the text to be classified belongs to the target classification in the target scenario through a preset classification model;

[0008] Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in the target scenario.

[0009] In a second aspect, an embodiment of the present application further provides a text classification device, comprising:

[0010] Acquisition module, used to obtain the text to be classified;

[0011] A classification module is used to determine whether the text to be classified belongs to the target classification in the target scenario through a preset classification model;

[0012] Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in the target scenario.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, comprising:

[0014] one or more processors;

[0015] a storage device for storing one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the text classification method as described in any one of the embodiments of the present application.

[0017] In a fourth aspect, an embodiment of the present application further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the text classification method as described in any one of the embodiments of the present application.

[0018] In a fifth aspect, an embodiment of the present application further provides a computer program product, characterized in that the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the text classification method as described in any one of the embodiments of the present application.

[0019] In the technical solution of the embodiment of the present application, a text to be classified can be obtained; through a preset classification model, it is determined whether the text to be classified belongs to the target classification in the target scenario; wherein, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage training based on a first sample text in a general scenario, and then a second stage training based on a second sample text in the target scenario.

[0020] During model training, a first-stage training phase using a first sample text from a general scenario yields a model with target classification and recognition capabilities. A second-stage training phase using a second sample text from a target scenario fine-tunes the first-stage trained model to produce a pre-set classification model with improved target classification and recognition capabilities for low-resource target scenarios. Consequently, the pre-set classification model, resulting from two phases of training, can improve text classification accuracy in target scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other features, advantages, and aspects of the various embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0022] Figure 1 A flowchart of a text classification method provided in an embodiment of the present application;

[0023] Figure 2 A flowchart of the first stage training process in a text classification method provided in an embodiment of the present application;

[0024] Figure 3 A schematic block diagram of a first feature vector construction process in a text classification method provided in an embodiment of the present application;

[0025] Figure 4 A schematic block diagram of a partial structure of a preset classification model in a text classification method provided in an embodiment of the present application;

[0026] Figure 5 A flowchart of the second stage training process in a text classification method provided in an embodiment of the present application;

[0027] Figure 6 A schematic block diagram of the input and output of a preset language model in a text classification method provided in an embodiment of the present application;

[0028] Figure 7 A schematic diagram of the structure of a text classification device provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the structure of a model building module in a text classification device provided in an embodiment of the present application;

[0030] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0032] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0033] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0038] Figure 1 This is a flowchart of a text classification method provided in an embodiment of the present application. This embodiment of the present application is applicable to target classification and identification in low-resource target scenarios, such as identifying malicious comments in specific community scenarios. The method can be performed by a text classification device, which can be implemented in software and / or hardware and configured in an electronic device, such as a computer.

[0039] like Figure 1 As shown, the text classification method provided in this embodiment may include:

[0040] S110, obtaining the text to be classified;

[0041] S120: Determine whether the text to be classified belongs to the target category in the target scenario through a preset classification model;

[0042] Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in a target scenario.

[0043] In this embodiment, the target scenarios may include, for example, various Internet communities. An Internet community may refer to an online platform that brings together users with common interests, needs, or goals and provides interactive services based on the Internet platform. In an Internet community, it may revolve around a specific topic or issue, and users may interact by posting content, participating in discussions, sharing resources, and so on. Target classification may include text categories with preset attributes, such as malicious text categories. Malicious text categories may include, for example, text categories involving abuse, discrimination, fraud, rumors, and other illegal and unethical content.

[0044] In this embodiment, the preset classification model may include a text classification model obtained by performing two stages of training on the basis of a pre-trained language model.

[0045] A pre-trained language model refers to a machine learning model with a large number of parameters and complex structure. It is often used in a conversational manner, can better understand user input, and can be used for various natural language processing tasks, such as language comprehension, text generation, and machine translation. Pre-training of a language model typically involves training with large amounts of unlabeled text data to learn the general laws and representations of language. Pre-trained language models can achieve performance on various natural language processing tasks through transfer learning.

[0046] In this embodiment, the two-stage training based on the pre-trained language model may include: first-stage training based on a first sample text in a general scenario, and then second-stage training based on a second sample text in a target scenario.

[0047] The first sample text may include samples belonging to the target classification and samples not belonging to the target classification in a general scenario; the second sample text may include samples belonging to the target classification and samples not belonging to the target classification in a target scenario. The sample size of the first sample text is usually larger than the sample size of the second sample text, and it can be considered that the target scenario is a low-resource scenario with a small sample size. It is worth noting that in this embodiment, the first sample text and the second sample text should be obtained in accordance with the requirements of relevant laws, regulations and relevant provisions.

[0048] The pre-trained language model can be retrained using the first sample text in a general scenario (i.e., first-stage training), so that the retrained language model can have the ability to recognize target-classified text. Then, the retrained language model can be fine-tuned using the second sample text in the target scenario (i.e., second-stage training), and a preset classification model with the ability to recognize target-classified text in the target scenario can be obtained. Since the second-stage training only fine-tunes the model parameters, the ability to recognize target-classified text in the target scenario can be achieved, which can ensure the accuracy of text classification in low-resource scenarios.

[0049] It is understandable that the pre-trained language model, the retrained language model, and the preset classification model have the same model structure, but the model parameters need to be gradually adjusted. For ease of description, this application may also refer to the pre-trained language model and the retrained language model as preset classification models, that is, the preset classification model adjusts the model parameters through the first stage training and the second stage training to have the classification ability to determine whether the text in the target scenario belongs to the target classification.

[0050] Accordingly, the text to be classified can be input into a pre-trained classification model, which can output whether the text to be classified belongs to the target category in the target scenario. For example, if the output is 0, it can indicate that the text to be classified does not belong to the target category; if the output is 1, it can indicate that the text to be classified belongs to the target category.

[0051] In this embodiment, corresponding text processing operations can also be performed based on whether the text to be classified belongs to the target classification. For example, if the target scenario is an internet community scenario, the text to be classified is a comment, and the target classification is malicious text, if the preset classification model determines that the comment is malicious text, then the comment can be blocked, thereby helping to maintain the community ecology and user experience.

[0052] In the technical solution of the embodiment of the present application, a text to be classified can be obtained; through a preset classification model, it is determined whether the text to be classified belongs to the target classification in the target scenario; wherein, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage training based on a first sample text in a general scenario, and then a second stage training based on a second sample text in the target scenario.

[0053] During model training, a first-stage training phase using a first sample text from a general scenario yields a model with target classification and recognition capabilities. A second-stage training phase using a second sample text from a target scenario fine-tunes the first-stage trained model to produce a pre-set classification model with improved target classification and recognition capabilities for low-resource target scenarios. Consequently, the pre-set classification model, resulting from two phases of training, can improve text classification accuracy in target scenarios.

[0054] The embodiments of the present application can be combined with the various optional solutions in the text classification method provided in the above embodiments. The text classification method provided in this embodiment describes the first stage training process in detail.

[0055] Figure 2 This is a flow chart of the first stage training process in a text classification method provided in an embodiment of the present application. Figure 2 As shown, in the text classification method provided in this embodiment, the first stage training process may include:

[0056] S210: Obtain a third sample text in a general scenario and a first classification truth value corresponding to the third sample text.

[0057] In this embodiment, the third sample text can be considered to belong to the first sample text. The third sample text and its first classification truth value should be collected in accordance with relevant laws, regulations, and other requirements. The first classification truth value can be considered to be a true value label indicating whether the third sample text belongs to the target classification, for example, it can be 0 or 1. As mentioned above, 0 can indicate that the third sample text does not belong to the target classification; 1 can indicate that the third sample text does belong to the target classification.

[0058] S220. Perform text enhancement on the third sample text by a preset enhancement method to obtain a first sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement.

[0059] In this embodiment, homophone replacement and homograph replacement can be used to enhance the text sample text belonging to the target category in the third sample text. Furthermore, research and analysis have revealed that the target category text may contain variant characters, such as variant numbers, letters, or symbols. Based on the transformation algorithm determined by the research, character replacement can be performed on the sample text belonging to the target category in the third sample text to further expand the third sample text.

[0060] The third sample text, as well as the third sample text after homophone replacement, homograph replacement, or character replacement, can be collectively referred to as the first sample text. The first classification truth value corresponding to the third sample text can be assigned to the corresponding third sample text after homophone replacement, homograph replacement, or character replacement, thereby obtaining the first classification truth value of each first sample text. Text enhancement of the third sample text through multiple replacement methods helps improve the model's recognition ability for target classification texts of various variants.

[0061] In addition, in this embodiment, after obtaining the first sample text, before performing feature construction on the first sample text, the first sample text can also be segmented and processed according to the input length acceptable to the preset classification model, so that the first sample text meets the input dimension requirements of the preset classification model.

[0062] S230. Perform feature construction on the first sample text through a feature construction module of a preset classification model to obtain a first feature vector; wherein the first feature vector is determined based on a character feature vector, a pronunciation feature vector, and a glyph feature vector.

[0063] In this embodiment, the preset classification model may include a feature construction module, wherein the feature construction module may extract a character feature vector, a pronunciation feature vector, and a glyph feature vector of the first sample text, and may obtain a first feature vector based on the character feature vector, the pronunciation feature vector, and the glyph feature vector.

[0064] By extracting multi-dimensional feature information of the meaning, pronunciation and shape of the first sample text, the text features can be fully mined, which helps to improve the accuracy of the model in detecting the target classification text.

[0065] In some optional implementations, feature construction is performed on the first sample text to obtain a first feature vector, which may include: based on the first sample text, expanding the dictionary corresponding to the preset classification model to obtain a character feature vector corresponding to the first sample text; obtaining the first pronunciation of the first sample text, performing vector conversion on the characters in the first pronunciation, and obtaining the pronunciation feature vector corresponding to the first sample text; obtaining the first decomposed glyph of the first sample text, performing vector conversion on the first decomposed glyph, and obtaining the glyph feature vector corresponding to the first sample text; and fusing the character feature vector, the pronunciation feature vector, and the glyph feature vector to obtain the first feature vector.

[0066] Among them, the dictionary corresponding to the preset classification model can be expanded by using the first sample text using an expansion method of an existing model dictionary. For example, the characters that were not successfully recognized in the first sample text can be converted from characters to identifier (identity, ID) encoding, and then converted from ID encoding to character vector representation. Based on the character vector representation of each character in the first sample text, a character feature vector corresponding to the first sample text can be obtained. For example, the character vector representation of each character can be sequentially spliced to obtain a character feature vector.

[0067] By using the first sample text to expand the dictionary corresponding to the preset classification model, the preset classification model can be equipped with the ability to recognize each word in the first sample text to adapt to the usage scenario of target classification text recognition.

[0068] Among them, an existing character-to-pronunciation conversion tool can be used to obtain the pronunciation of each character in the first sample text. The pronunciation can be split into multiple characters, and then the character-to-vector conversion is performed using an existing method (such as a convolutional neural network) to obtain a pronunciation vector representation of each character. Based on the pronunciation vector representation of each character in the first sample text, a pronunciation feature vector corresponding to the first sample text can be obtained. For example, the pronunciation vector representations of each character can be sequentially spliced to obtain a pronunciation feature vector.

[0069] Among them, existing character decomposition tools can be used to decompose characters into structural units or stroke dimensions. For example, Chinese characters can be decomposed into structural units to obtain components such as radicals and radicals. Among them, the decomposition results can be converted into vectors through existing methods (such as convolutional neural networks) to obtain the glyph vector representation of each character. According to the glyph vector representation of each character in the first sample text, the glyph feature vector corresponding to the first sample text can be obtained. For example, the glyph vector representation of each character can be sequentially spliced to obtain a phonetic feature vector.

[0070] For example, Figure 3 This is a schematic diagram of the process of constructing the first feature vector in a text classification method provided in an embodiment of the present application. Figure 3 , a fully connected layer can be used to perform linear layer fusion on the character feature vector, the phonetic feature vector and the glyph feature vector to obtain a complete vector representation of the first sample text (i.e., the first feature vector).

[0071] In these optional implementations, the character feature vector, the phonetic feature vector and the glyph feature vector of the first sample text can be determined by obtaining the character vector representation, the phonetic vector representation and the glyph vector representation of each character.

[0072] S240. Process the first feature vector through the multi-head attention module, multi-head adapter module and feedforward module of the preset classification model, and output a first predicted classification.

[0073] For example, Figure 4 This is a schematic block diagram of a partial structure of a preset classification model in a text classification method provided in an embodiment of the present application. The main model of the preset classification model may include a multi-head attention module 410, a multi-head adapter module 420 and a feedforward module 430. Figure 4 , the multi-head attention module 410 may include a multi-head attention layer, a residual layer and a normalization layer; the feedforward module 430 may include a feedforward layer, a residual layer and a normalization layer. Figure 4 In , the residual layer and the normalization layer in the multi-head attention module 410 and the feedforward module 430 can be represented by the same layer.

[0074] A multi-head adapter module can be set between each multi-head attention module and feedforward module in the preset classification model. The multi-head adapter module can have the ability to convert the input vector into an output vector of the same dimension, which enables the multi-head adapter to better learn the essential characteristics of the vector and improve the vector representation effect.

[0075] The processing of the multi-head adapter module 430 may include: sequentially processing the input vector through the down-projection layer, the nonlinear activation function layer, and the up-projection layer in the multi-head adapter module 430. Figure 4 Regarding the expanded portion of the multi-head adapter module, the multi-head adapter module 430 may include a lower projection layer, a nonlinear activation function layer, and an upper projection layer. The activation function corresponding to the nonlinear activation function layer may include, for example, a rectified linear unit (ReLU) function. The lower projection layer projects the input vector into a low-dimensional feature space, the nonlinear activation function layer performs a nonlinear transformation of the low-dimensional features, and the upper projection layer projects the low-dimensional features into the original-dimensional feature space to obtain a transformed vector.

[0076] By deploying a multi-head adapter module consisting of a down-projection layer, a nonlinear activation function layer, and an up-projection layer, the essential low-dimensional characteristics of the vector can be learned, which can better represent the contextual information of the vector and help improve the vector representation effect.

[0077] S250. Determine a first loss value based on the first predicted classification and the first classification true value, and perform a first-stage training on the preset classification model based on the first loss value.

[0078] In this embodiment, the loss value of the first predicted classification and the first classification true value corresponding to the same first sample text can be calculated using the existing classification loss function to obtain the first loss value. For example, when the first predicted classification and the first classification true value are binary classifications, the following cross entropy loss function can be used to calculate the first loss value L1:

[0079]

[0080] Among them, y can represent the first classification truth value; Can represent the first predicted classification; where y and The value can be 0 or 1, where 0 means it does not belong to the target category; 1 means it belongs to the target category.

[0081] During the first stage of training, the first loss value can be back-propagated to adjust the parameters of each network layer of the preset classification model to fully learn the characteristic information of the target classification text in general scenarios.

[0082] The technical solution of the embodiment of the present application provides a detailed description of the first stage training process. By extracting the character feature vector, phonetic feature vector and glyph feature vector corresponding to the first sample text, and combining the multi-dimensional feature information to determine the first feature vector, not only can the text that obviously belongs to the target classification be detected, but also the variant text of the target classification and the text of the implicit target classification can be identified, which can improve the recognition accuracy of the target classification. In addition, by adding a multi-head adapter module to the model structure, the feature vector can better represent the context, which can improve the representation effect of the feature vector and also help improve the recognition accuracy of the target classification.

[0083] In addition, the text classification method provided in the embodiment of the present application and the text classification method provided in the above embodiment belong to the same public concept. Technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0084] The text classification method provided in the embodiment of the present application can be combined with the various optional solutions in the above embodiments. The text classification method provided in this embodiment describes the second stage training process in detail.

[0085] Figure 5 This is a flow chart of the second stage training process in a text classification method provided in an embodiment of the present application. Figure 5 As shown, the text classification method provided in this embodiment, the second stage training process may include:

[0086] S510. Obtain a fourth sample text in a target scenario and a second classification truth value corresponding to the fourth sample text; wherein, the fifth sample text in the fourth sample text and the second classification truth value corresponding to the fifth sample text are generated based on a preset language model.

[0087] In this embodiment, the fourth sample text in the target scenario may include the fifth sample text and other sample texts except the fifth sample text. The fifth sample text may be generated by a preset language model; the other sample texts may be collected in the target scenario in accordance with the requirements of relevant laws, regulations and related provisions. The second classification truth value corresponding to the fifth sample text may also be generated by a preset language model; the second classification truth values of other sample texts may include pre-labeled classification truth values. The second classification truth value may be considered as a true value label of whether the fourth sample text belongs to the target classification, for example, it may be 0 or 1. Referring to the above, 0 may indicate that the fourth sample text does not belong to the target classification; 1 may indicate that the fourth sample text belongs to the target classification.

[0088] Among them, the preset language model may include, for example, an existing large language model (LLM). Among them, the prompt word technology of the large language model can be used to generate the fifth sample text and its second classification truth value. Among them, the prompt word can be understood as a word or phrase that points to or prompts a specific language point of the large language model, which can help the large language model recall and apply the learned language knowledge, and at the same time guide the large language model to perform correct language expression. Using the large language model prompt word technology to expand the data set can achieve data enhancement of the data set in low-resource target scenarios, which is conducive to improving the model training effect.

[0089] S520: Perform text enhancement on the fourth sample text by a preset enhancement method to obtain a second sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement.

[0090] Here, the fourth sample text may be enhanced with reference to the process of enhancing the third sample text to obtain the second sample text.

[0091] S530. Perform feature construction on the second sample text through a feature construction module of a preset classification model to obtain a second feature vector; wherein the second feature vector is determined based on the character feature vector, the pronunciation feature vector, and the glyph feature vector.

[0092] Here, the feature construction process for the first sample text may be referred to, and the feature construction for the second sample text may be performed to obtain a second feature vector.

[0093] S540: Process the second feature vector through the multi-head attention module, multi-head adapter module and feedforward module of the preset classification model, and output a second predicted classification.

[0094] Among them, the output process of the first prediction classification can be referred to, and the second feature vector can be processed by the multi-head attention module, multi-head adapter module and feedforward module of the preset classification model to output the second prediction classification.

[0095] S550: Determine a second loss value based on the second predicted classification and the second classification true value, and perform a second stage of training on the multi-head adapter module based on the second loss value.

[0096] The second loss value can be determined based on the second predicted classification and the second classification's true value, with reference to the process for determining the first loss value. During the second phase of training, the second loss value can be backpropagated, and the parameters of the other parts of the prediction and classification model are fixed during backpropagation, while only the parameters of the multi-head adapter module are fine-tuned. Because the multi-head adapter module has a small number of parameters, it can accurately recognize the target classification text in the target scenario even when the second sample text is small.

[0097] In some optional implementations, the process of generating the fifth sample text and the second classification truth value corresponding to the fifth sample text may include:

[0098] Construct a first prompt word based on the first content data in the target scenario; generate a first comment text based on the first prompt word through a preset language model; construct a second prompt word based on the first content data and the first comment text; generate a third classification truth value of the first comment text based on the second prompt word through a preset language model; determine the probability that the first comment text belongs to the target classification in the target scenario through the preset classification model obtained through the first stage training; in response to the probability belonging to the preset range, determine the corresponding first comment text as the fifth sample text, and determine the third classification truth value corresponding to the fifth sample text as the second classification truth value.

[0099] The preset classification model can be applied to determine whether a comment text in an Internet community scenario is malicious text. For example, Figure 6 This is a schematic diagram of the input and output of a preset language model in a text classification method provided in an embodiment of the present application. Figure 6 ,The preset language model can be generated through two-stage prompt words, and the first comment text and the third classification truth value are obtained respectively.

[0100] like Figure 6 In the first stage, prompt word generation may include: constructing a first prompt word according to the posting content (i.e., first content data) in the Internet community scenario; Figure 6In the example, the post content can be represented by "{passage}", and the first prompt word can include Figure 6 The content in the middle A box; the first prompt word can be input into the preset language model to make it output the first comment text; Figure 6 In the example, the first comment text can be represented by "{comment}".

[0101] like Figure 6 In the second stage, the prompt word generation may include: constructing a second prompt word according to the first content data and the first comment text; wherein the second prompt word may include, for example Figure 6 The content in the middle B box; the second prompt word can be input into the preset language model to make it output the third classification truth value corresponding to the first comment text; Figure 6 In the example, the output of the third classification truth value may include "yes" or "no".

[0102] Since the classification model can usually determine the probability of belonging to each category before outputting the final classification, in this embodiment, the first comment text can be input into the preset classification model obtained by the first stage training, and the probability of the first comment text belonging to the target category in the target scenario determined by the model can be obtained. Among them, the preset range can be pre-set according to the experience value or experimental value, such as [45%, 55%], etc. In the case that the probability of the first comment text belonging to the target category in the target scenario determined by the preset classification model falls within the preset range, it can be considered that the first comment text belongs to a sample that is more difficult to classify by the preset classification model. At this time, these first comment texts that are more difficult to classify can be determined as the fifth sample text, and the third classification true value corresponding to the fifth sample text can be determined as the second classification true value.

[0103] In these optional implementations, the first comment text and the third classification truth value can be obtained through two-stage prompt word generation based on a preset language model. By mining the fifth sample text that is more difficult for the preset classification model to classify from the first comment text, and performing the second-stage training of the preset classification model using the fifth sample text and its corresponding second classification truth value, the preset classification model's ability to recognize the target classification text in the target scenario can be improved, thereby improving the model's classification accuracy. While reducing the cost of Internet community comment management, it can also enhance the community ecology and user experience.

[0104] The technical solution of the embodiment of the present application provides a detailed description of the second-stage training process. Through the effective use of the prompt word technology of the large language model, it is possible to expand the data of samples in the target scenario with low resources, which is beneficial to improving the fine-tuning effect of the preset classification model in the second-stage training. At the same time, by combining the character feature vector, the pronunciation feature vector and the glyph feature vector of the second sample text, the second feature vector is determined, which can improve the recognition accuracy of the target classification. In addition, by only adjusting the parameters of the multi-head adapter module in the second-stage training, it is possible to achieve better training effects with a small number of samples while ensuring that the preset classification model adapts to the specific field of the target scenario. Combining the above methods, it is possible to achieve while ensuring the recognition accuracy of the model, and also to improve the robustness and generalization ability of the model when facing the challenges of low-resource scenarios.

[0105] In addition, the text classification method provided in the embodiment of the present application and the text classification method provided in the above embodiment belong to the same public concept. Technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0106] Figure 7 This is a schematic diagram of the structure of a text classification device provided in an embodiment of the present application. The text classification device provided in this embodiment is suitable for target classification and identification in low-resource target scenarios, such as the identification of malicious comments in a specific community scenario.

[0107] like Figure 7 As shown, the text classification device provided in the embodiment of the present application may include:

[0108] An acquisition module 710 is used to acquire the text to be classified;

[0109] The classification module 720 is used to determine whether the text to be classified belongs to the target category in the target scenario through a preset classification model;

[0110] Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in a target scenario.

[0111] For example, Figure 8 This is a structural diagram of a model building module in a text classification device provided in an embodiment of the present application. Figure 8 In some optional implementations, the text classification device further includes a model construction module; wherein the model construction module may include:

[0112] The first stage training unit is used to perform the following first stage training process:

[0113] Obtaining a third sample text in a general scenario and a first classification truth value corresponding to the third sample text;

[0114] Performing text enhancement on the third sample text by a preset enhancement method to obtain the first sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement;

[0115] Performing feature construction on the first sample text using a feature construction module of a preset classification model to obtain a first feature vector; wherein the first feature vector is determined based on the character feature vector, the pronunciation feature vector, and the glyph feature vector;

[0116] Processing the first feature vector through a multi-head attention module, a multi-head adapter module, and a feedforward module of a preset classification model to output a first predicted classification;

[0117] A first loss value is determined based on the first predicted classification and the first classification true value, and a first stage of training is performed on the preset classification model based on the first loss value.

[0118] like Figure 8 As shown, the first stage training unit may include a first data processor and a first model trainer. The first data processor may include a first data preprocessor and a first feature constructor. The third sample text may be enhanced by the first data preprocessor. The first feature constructor may be considered to belong to a feature construction module and may determine a first feature vector. The first model trainer may output a first predicted classification, determine a first loss value, and train parameters in network modules such as a feature construction module, a multi-head attention module, a multi-head adaptability module, and a feedforward module in a preset classification model according to the first loss value to achieve the first stage training.

[0119] In some optional implementations, the first feature constructor may be used to:

[0120] Based on the first sample text, expanding the dictionary corresponding to the preset classification model to obtain a character feature vector corresponding to the first sample text;

[0121] Obtaining a first pronunciation of a first sample text, performing vector conversion on characters in the first pronunciation, and obtaining a pronunciation feature vector corresponding to the first sample text;

[0122] Obtaining a first decomposed glyph of a first sample text, performing vector conversion on the first decomposed glyph to obtain a glyph feature vector corresponding to the first sample text;

[0123] The character feature vector, the pronunciation feature vector and the glyph feature vector are fused to obtain a first feature vector.

[0124] In some optional implementations, the first model trainer may be used to:

[0125] The input vector is processed sequentially through the down-projection layer, nonlinear activation function layer, and up-projection layer in the multi-head adapter module.

[0126] In some optional implementations, the model building module may further include:

[0127] The second stage training unit is used to perform the following second stage training process:

[0128] Obtaining a fourth sample text in the target scenario and a second classification truth value corresponding to the fourth sample text; wherein the fifth sample text in the fourth sample text and the second classification truth value corresponding to the fifth sample text are generated based on a preset language model;

[0129] Performing text enhancement on the fourth sample text by a preset enhancement method to obtain a second sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement;

[0130] Performing feature construction on the second sample text using a feature construction module of a preset classification model to obtain a second feature vector; wherein the second feature vector is determined based on the character feature vector, the pronunciation feature vector, and the glyph feature vector;

[0131] Processing the second feature vector through a multi-head attention module, a multi-head adapter module, and a feedforward module of a preset classification model to output a second predicted classification;

[0132] A second loss value is determined according to the second predicted classification and the second classification true value, and the multi-head adapter module is trained in the second stage according to the second loss value.

[0133] See again Figure 8 The preset classification model obtained by the first stage training unit can be fine-tuned by the second stage training unit. The second stage training unit may include a data expander, a second data processor, and a second model trainer. The data expander may generate a fifth sample text and a second classification truth value corresponding to the fifth sample text through a preset language model. The second data processor may include a second data preprocessor and a second feature constructor ( Figure 8(not shown). The fourth sample text can be enhanced by a second data preprocessor. The second feature constructor can be considered to also belong to a feature construction module of the preset classification model and can determine a second feature vector. The second model trainer can output a second predicted classification, determine a second loss value, and fine-tune the parameters in the multi-head adaptability of the preset classification model based on the second loss value to achieve second-stage training.

[0134] In some optional implementations, the data expander may generate the fifth sample text and the second classification truth value corresponding to the fifth sample text through the following process:

[0135] Constructing a first prompt word according to the first content data in the target scenario;

[0136] Generate a first comment text according to the first prompt word using a preset language model;

[0137] constructing a second prompt word according to the first content data and the first comment text;

[0138] Generate a third classification truth value of the first comment text according to the second prompt word through a preset language model;

[0139] The preset classification model obtained through the first stage of training determines the probability that the first comment text belongs to the target category in the target scenario;

[0140] In response to the probability belonging to the preset range, the corresponding first comment text is determined as the fifth sample text, and the third classification true value corresponding to the fifth sample text is determined as the second classification true value.

[0141] The text classification device provided in the embodiments of this application can execute the text classification method provided in any of the embodiments of this application and has functional modules corresponding to the execution of the method. For technical details not fully described in this embodiment, please refer to the above-mentioned embodiments of the text classification method, and the same technical features in this embodiment have the same beneficial effects as in the above-mentioned embodiments.

[0142] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of this application.

[0143] Reference below Figure 9 , which shows an electronic device (eg Figure 9The terminal device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0144] like Figure 9 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0145] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display, a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Figure 9 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0146] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the text classification method of the embodiment of the present application are performed.

[0147] The electronic device provided in the embodiment of the present application and the text classification method provided in the above embodiment belong to the same public concept. Technical details not fully described in this embodiment can be found in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0148] An embodiment of the present application provides a storage medium of computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they can be used to execute the text classification method provided in the above embodiment.

[0149] It should be noted that the storage medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium containing or storing executable instructions that can be used by or in combination with an instruction execution system, device or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable executable instructions. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit executable instructions for use by or in conjunction with an instruction execution system, apparatus, or device. The executable instructions contained on the storage medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0150] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0151] The storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0152] The storage medium carries one or more executable instructions. When the one or more executable instructions are executed by the electronic device, the electronic device:

[0153] Obtain the text to be classified; determine whether the text to be classified belongs to the target classification in the target scenario through a preset classification model; wherein the preset classification model is obtained based on two-stage training; the two-stage training includes a first-stage training based on a first sample text in a general scenario, and a second-stage training based on a second sample text in the target scenario.

[0154] The executable instructions for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The executable instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0155] An embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, can implement the text classification method provided in any embodiment of the present application.

[0156] The computer program product includes a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the text classification method. The program code can be written in one or more programming languages or a combination thereof, the programming language including an object-oriented programming language such as Java, Smalltalk, C++, and also including a conventional procedural programming language such as "C" language or similar programming language. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, via the Internet using an Internet service provider).

[0157] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0158] The units involved in the embodiments described in this application may be implemented in software or hardware. In some cases, the names of the units and modules do not constitute limitations on the units and modules themselves.

[0159] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.

[0160] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0161] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

[0162] In addition, although adopting specific order to describe each operation, this should not be interpreted as requiring these operations to be executed in the specific order shown or in sequential order.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the application.Some features described in the context of separate embodiment can also be implemented in a single embodiment in combination.On the contrary, the various features described in the context of a single embodiment also can be implemented in multiple embodiments individually or in the mode of any suitable subcombination.

[0163] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A text classification method, characterized in that: include: Get the text to be classified; Determine whether the text to be classified belongs to the target classification in the target scenario through a preset classification model; Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in the target scenario.

2. The method according to claim 1, characterized in that The first phase of training includes: Obtaining a third sample text in the general scenario and a first classification truth value corresponding to the third sample text; Performing text enhancement on the third sample text by a preset enhancement method to obtain the first sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement; Performing feature construction on the first sample text using the feature construction module of the preset classification model to obtain a first feature vector; wherein the first feature vector is determined based on a character feature vector, a pronunciation feature vector, and a glyph feature vector; Processing the first feature vector through the multi-head attention module, the multi-head adapter module, and the feedforward module of the preset classification model to output a first predicted classification; A first loss value is determined based on the first predicted classification and the first classification true value, and the first stage training is performed on the preset classification model based on the first loss value.

3. The method according to claim 2, characterized in that The step of constructing features of the first sample text to obtain a first feature vector includes: Based on the first sample text, expanding the dictionary corresponding to the preset classification model to obtain a character feature vector corresponding to the first sample text; Obtaining a first pronunciation of the first sample text, performing vector conversion on characters in the first pronunciation, and obtaining a pronunciation feature vector corresponding to the first sample text; Obtaining a first decomposed glyph of the first sample text, performing vector conversion on the first decomposed glyph to obtain a glyph feature vector corresponding to the first sample text; The character feature vector, the pronunciation feature vector and the glyph feature vector are fused to obtain the first feature vector.

4. The method according to claim 2, characterized in that The processing process of the multi-head adapter module includes: The input vector is processed sequentially through the down-projection layer, the nonlinear activation function layer, and the up-projection layer in the multi-head adapter module.

5. The method according to claim 1, wherein The second phase of training includes: Obtaining a fourth sample text in the target scenario and a second classification truth value corresponding to the fourth sample text; wherein a fifth sample text in the fourth sample text and the second classification truth value corresponding to the fifth sample text are generated based on the preset language model; Performing text enhancement on the fourth sample text by a preset enhancement method to obtain the second sample text; wherein the preset enhancement method includes: homophone replacement, homograph replacement, and character replacement; Performing feature construction on the second sample text using the feature construction module of the preset classification model to obtain a second feature vector; wherein the second feature vector is determined based on the character feature vector, the pronunciation feature vector, and the glyph feature vector; Processing the second feature vector through the multi-head attention module, the multi-head adapter module, and the feedforward module of the preset classification model to output a second predicted classification; A second loss value is determined based on the second predicted classification and the second classification true value, and the second stage training is performed on the multi-head adapter module based on the second loss value.

6. The method according to claim 5, characterized in that The process of generating the fifth sample text and the second classification true value corresponding to the fifth sample text includes: Constructing a first prompt word according to the first content data in the target scenario; Generate a first comment text according to the first prompt word using the preset language model; constructing a second prompt word according to the first content data and the first comment text; Generating a third classification truth value of the first comment text according to the second prompt word using the preset language model; Determining the probability that the first comment text belongs to the target classification in the target scenario using the preset classification model obtained through the first stage training; In response to the probability belonging to a preset range, the corresponding first comment text is determined as the fifth sample text, and the third classification true value corresponding to the fifth sample text is determined as the second classification true value.

7. A text classification device, characterized in that: include: Acquisition module, used to obtain the text to be classified; A classification module is used to determine whether the text to be classified belongs to the target classification in the target scenario through a preset classification model; Among them, the preset classification model is obtained based on two stages of training; the two stages of training include a first stage of training based on a first sample text in a general scenario, and a second stage of training based on a second sample text in the target scenario.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the text classification method according to any one of claims 1 to 6.

9. A storage medium comprising computer-executable instructions, wherein the computer-executable instructions are used to perform the text classification method according to any one of claims 1 to 6 when executed by a computer processor.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the text classification method according to any one of claims 1 to 6.