Text classification model training method and device, equipment and medium
Patent Information
- Application Number
- CN202410434475.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-04-11
AI Technical Summary
[0003]目前常见的文本分类模型的训练方法多为将同一标签对应多个文本数据作为一个训练集合进行模型训练,使得模型能够学习当前标签对应的多个文本数据的共通性,从而得到训练完成的文本分类模型,例如,在金融领域中,通过对多个已知类别的保险说明文本进行模型训练,得到训练完成的文本分类模型,并使用训练完成的文本分类模型对未知类别的保险说明文本进行分类,但是这种训练方法在遇到文本内容较为相似,但标签略有不同的保险说明文本时,由于模型的泛化能力,从而容易造成这些保险说明文本具有相似的标签,使得文本分类模型的精准度不高
[0053]本发明实施例通过构建分词训练文本的正样本文本及负样本文本,并利用所述分词训练文本、所述正样本文本及所述负样本文本对预设文本分类模型进行训练,从而使得训练完成的文本分类模型在提取金融类文本数据特征时具有更好的差异表现,进一步地,利用预设文本分类模型分别预测所述分词训练文本与所述正样本文本及所述负样本文本的文本相似度,根据文本相似度与预设目标相似度之间的损失值调整所述预设文本分类模型的参数,直至所述第一损失值小于第一预设阈值,得到初步训练完成的文本分类模型,使得所述文本分类模型具有更强的文本表征能力,从而提升文本分类的精准度;更进一步地,对训练完成的文本分类模型中多层感知机模块的层数进行调整,并对调整后的改进多层感知机模块进行训练,得到训练完成的多层感知机模块,提高了多层感知机模块进行文本分类的精准度,使得金融类文本分类的结果更为精准。因此,本发明提供的一种文本分类模型的训练方法、装置、设备及存储介质,能够提高金融类文本分类模型的精准度,从而使得金融工作者在查找相关金融类文本时,更加便捷有效。
Smart Images

Figure CN118312778B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and more particularly to a training method, apparatus, electronic device, and readable storage medium for a text classification model. Background Technology
[0002] Text classification refers to identifying the category of any text according to predefined rules. For example, in the financial field, various insurance documents can be classified according to different given tags, making document retrieval more convenient.
[0003] Currently, common training methods for text classification models often involve training the model using multiple text data sets corresponding to the same label as a training set. This allows the model to learn the commonalities among the multiple text data sets corresponding to the current label, thus obtaining a trained text classification model. For example, in the financial field, a trained text classification model is obtained by training multiple insurance description texts of known categories, and then using the trained text classification model to classify insurance description texts of unknown categories. However, when encountering insurance description texts with similar content but slightly different labels, this training method can easily result in these insurance description texts having similar labels due to the model's generalization ability, leading to low accuracy of the text classification model. Summary of the Invention
[0004] This invention provides a training method, apparatus, electronic device, and readable storage medium for a text classification model, which effectively improves the convenience and accuracy of financial text classification, with the aim of improving the precision of the text classification model.
[0005] To achieve the above objectives, the present invention provides a method for training a text classification model, the method comprising:
[0006] Obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules;
[0007] The text similarity between the word segmentation training text and the positive sample text and the negative sample text is predicted using a preset text classification model.
[0008] The first loss value between the similarity and the preset target similarity is calculated using a preset first loss function, and the parameters of the preset text classification model are adjusted according to the first loss value until the first loss value is less than a first preset threshold, thus obtaining a text classification model that has been initially trained.
[0009] Based on the number of text tags, the number of multilayer perceptron modules in the initially trained text classification model is adjusted to obtain an improved multilayer perceptron module. The feature vector of the word segmentation training text is then extracted using the initially trained text classification model to obtain the word segmentation text feature vector.
[0010] The improved multilayer perceptron module is used to calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector;
[0011] The second loss value between the predicted probability and the preset target probability is calculated using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, thus obtaining a trained text classification model.
[0012] Optionally, the step of using a preset text classification model to predict the text similarity between the segmented training text and the positive sample text and the negative sample text respectively includes:
[0013] Add corresponding classification identifiers to the words of the word segmentation training text, the positive sample text, and the negative sample text respectively to obtain word segmentation training text, positive sample text, and negative sample text with classification identifiers;
[0014] The word segmentation training text, positive sample text, and negative sample text with classification identifiers are encoded using the encoding layer in the preset text classification model to obtain the word segmentation text classification identifier encoding vector, the positive sample classification identifier encoding vector, and the negative sample classification identifier encoding vector.
[0015] The similarity between the segmented text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector is calculated using the multilayer perceptron module in the preset text classification model, respectively, to obtain the text similarity between the segmented training text and the positive sample text and the negative sample text.
[0016] Optionally, the step of using the multilayer perceptron module in the preset text classification model to calculate the similarity between the segmented text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector, respectively, to obtain the text similarity between the segmented training text and the positive sample text and the negative sample text, includes:
[0017] The word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector are vector-connected using the preset connection function in the multilayer perceptron module to obtain a positive sample comparison connection vector; and the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector are vector-connected using the connection function to obtain a negative sample comparison connection vector.
[0018] The first preset parameter in the multilayer perceptron module is used to multiply the positive sample comparison connection vector and the negative sample comparison connection vector respectively to obtain the positive comparison multiplication vector and the negative comparison multiplication vector;
[0019] Using the preset activation function in the multilayer perceptron module, the positive sample similarity between the word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector in the positive contrast multiplication vector, and the negative sample similarity between the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector in the negative contrast multiplication vector are calculated respectively.
[0020] The positive sample similarity and the negative sample similarity are normalized using the second preset parameters in the multilayer perceptron module to obtain the text similarity between the word segmentation training text and the positive sample text and the negative sample text.
[0021] Optionally, constructing the positive and negative sample texts of the word segmentation training text according to preset rules includes:
[0022] Select one segmented training text from the segmented training texts arbitrarily and without repetition as the target segmented training text;
[0023] The segmented training texts that have the same text labels as the target segmented training text are used as the first positive sample texts of the target segmented training text.
[0024] Randomly select a text label from the text labels corresponding to the target word segmentation training text as the text label to be trained, and take the word segmentation training text with the text label to be trained as the second positive sample text of the target word segmentation training text.
[0025] Data augmentation processing is performed on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text;
[0026] The word segmentation training text that has an inclusion relationship with the text tags in the target word segmentation training text and is trained for different text tags is used as the first negative sample text of the target word segmentation training text.
[0027] The word segmentation training text that is completely identical to the text label except for the text label to be trained is used as the second negative sample text of the target word segmentation training text;
[0028] The word segmentation training text that has the highest encoding vector similarity to the target word segmentation training text and has a different text label is used as the third negative sample text of the target word segmentation training text.
[0029] Optionally, the step of performing data augmentation processing on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text includes:
[0030] The target word segmentation training text is back-translated using a preset translation dictionary to obtain the back-translated target word segmentation training text.
[0031] Extract text keywords from the target word segmentation training text;
[0032] The text keywords are matched with phrases in a preset thesaurus to obtain successfully matched thesaurus phrases;
[0033] The text keywords in the target word segmentation training text are replaced with the successfully matched synonym phrases to obtain the replaced target word segmentation training text.
[0034] By integrating the back-translated target segmentation training text and the replaced target segmentation training text, a third positive sample text of the target segmentation training text is obtained.
[0035] Optionally, the step of extracting the feature vector of the word segmentation training text using the pre-trained text classification model to obtain the word segmentation text feature vector includes:
[0036] The word segmentation training text is encoded using the encoding layer in the preliminarily trained text classification model to obtain the word segmentation text encoding vector;
[0037] The word segmentation text encoding vector is subjected to position index encoding to obtain the word segmentation text position encoding vector;
[0038] The word segmentation text position encoding vector is combined with the word segmentation text encoding vector to obtain the word segmentation text feature vector.
[0039] Optionally, the step of calculating the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector using the improved multilayer perceptron module includes:
[0040] The third, fourth and fifth preset parameters in the improved multilayer perceptron module are used to perform linear transformations on the word segmentation text feature vectors to obtain query vectors, key vectors and numerical vectors.
[0041] The similarity matrix is obtained by multiplying the query vector with the transpose of the key vector.
[0042] The similarity matrix is normalized to obtain a normalized matrix;
[0043] The activation matrix is obtained by performing activation calculation on the normalized matrix;
[0044] The dot product of the activation matrix and the numerical vector yields the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector.
[0045] To address the aforementioned problems, the present invention also provides a training apparatus for a text classification model, the apparatus comprising:
[0046] The positive and negative sample construction module is used to obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules;
[0047] The model preliminary training module is used to predict the text similarity between the word segmentation training text and the positive sample text and the negative sample text using a preset text classification model, calculate the first loss value between the similarity and the preset target similarity using a preset first loss function, and adjust the parameters of the preset text classification model according to the first loss value until the first loss value is less than a first preset threshold, so as to obtain the text classification model that has been preliminarily trained.
[0048] The complete model training module is used to adjust the number of multilayer perceptron modules in the initially trained text classification model according to the number of text tags, to obtain an improved multilayer perceptron module, and to extract the feature vector of the segmented training text using the initially trained text classification model, to obtain the segmented text feature vector. The improved multilayer perceptron module is used to calculate the predicted probability of the text tag of the segmented training text corresponding to the segmented text feature vector, and to calculate the second loss value between the predicted probability and the preset target probability using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, to obtain the trained text classification model.
[0049] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:
[0050] Memory, storing at least one computer program; and
[0051] The processor executes the computer program stored in the memory to implement the training method of the text classification model described above.
[0052] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the above-described text classification model training method.
[0053] This invention constructs positive and negative sample texts for word segmentation training and trains a preset text classification model using these texts. This results in a better difference performance when extracting features from financial text data. Furthermore, the preset text classification model predicts the text similarity between the word segmentation training text and the positive and negative sample texts. The parameters of the preset text classification model are adjusted based on the loss value between the text similarity and a preset target similarity, until the first loss value is less than a first preset threshold. This yields a pre-trained text classification model with stronger text representation capabilities, thus improving the accuracy of text classification. Even further, the number of layers in the multilayer perceptron module of the trained text classification model is adjusted, and the adjusted improved multilayer perceptron module is trained to obtain a trained multilayer perceptron module. This improves the accuracy of the multilayer perceptron module in text classification, resulting in more accurate classification of financial texts. Therefore, the training method, apparatus, device, and storage medium for a text classification model provided by this invention can improve the accuracy of financial text classification models, thereby making it more convenient and effective for financial workers to find relevant financial texts. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating a training method for a text classification model provided in an embodiment of the present invention.
[0055] Figure 2 and Figure 3 A detailed implementation flowchart of one step in the training method of a text classification model provided in an embodiment of the present invention;
[0056] Figure 4 A schematic diagram of a training device for a text classification model provided in an embodiment of the present invention;
[0057] Figure 5 A schematic diagram of the internal structure of an electronic device for implementing a training method for a text classification model according to an embodiment of the present invention;
[0058] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0059] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0060] This invention provides a method for training a text classification model. The execution entity of the text classification model training method includes, but is not limited to, at least one of electronic devices, such as a server or a terminal, that can be configured to execute the method provided in this application. In other words, the text classification model training method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server can include an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0061] Reference Figure 1 The flowchart shown is a schematic diagram of a text classification model training method provided in an embodiment of the present invention. In this embodiment, the text classification model training method includes:
[0062] S1. Obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules.
[0063] In this embodiment of the invention, the word segmentation training text can be financial text data that has undergone word segmentation processing. This financial text data can be explanatory text data of financial products, such as insurance explanatory text data and wealth management product explanatory text data. The text tags can be labels assigned to the word segmentation training text by relevant personnel based on its text content. For example, in the financial industry, insurance explanatory text might be tagged with tags such as pension insurance, work injury insurance, and medical insurance.
[0064] In this embodiment of the invention, the preset rules include rules for constructing positive sample text and rules for constructing negative samples. The rules for constructing positive sample text include the following: Word segmentation training texts with completely identical text labels are considered as each other's positive sample texts; when training a text classification model for texts with the same content, word segmentation training texts in the word segmentation training text pair are considered as each other's positive sample texts; and augmented word segmentation training texts after data augmentation are considered as positive sample texts along with their corresponding word segmentation training texts. The rules for constructing negative samples include the following: When faced with word segmentation training text pairs that have most of the same text labels and a small number of different labels, when training a text classification model for different text labels, the word segmentation training texts in the word segmentation training text pairs are regarded as each other's negative sample texts; when faced with word segmentation training text pairs whose text labels have an inclusion relationship, when training a text classification model for different text labels, the word segmentation training texts in the word segmentation training text pairs are regarded as each other's negative sample texts; the word segmentation training texts with the highest encoding vector similarity to the word segmentation training text to be paired, but with different text labels, are regarded as the negative sample texts of the word segmentation training text to be paired.
[0065] In one optional embodiment of the present invention, text data can be downloaded from the network and segmented into words to obtain segmented training text. Then, the segmented training text can be labeled manually or by a pre-trained text labeling model to obtain text labels.
[0066] In this embodiment of the invention, positive and negative sample texts of the word segmentation training text are constructed according to preset rules, providing comparative samples for the training of the text classification model, thereby enabling the trained text classification model to obtain better difference representation and improve the accuracy of text classification.
[0067] Furthermore, as an optional embodiment of the present invention, the step of constructing the positive sample text and negative sample text of the word segmentation training text according to preset rules includes:
[0068] Select one segmented training text from the segmented training texts arbitrarily and without repetition as the target segmented training text;
[0069] The segmented training texts that have the same text labels as the target segmented training text are used as the first positive sample texts of the target segmented training text.
[0070] Randomly select a text label from the text labels corresponding to the target word segmentation training text as the text label to be trained, and take the word segmentation training text with the text label to be trained as the second positive sample text of the target word segmentation training text.
[0071] Data augmentation processing is performed on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text;
[0072] The word segmentation training text that has an inclusion relationship with the text tags in the target word segmentation training text and is trained for different text tags is used as the first negative sample text of the target word segmentation training text.
[0073] The word segmentation training text that is completely identical to the text label except for the text label to be trained is used as the second negative sample text of the target word segmentation training text;
[0074] The word segmentation training text that has the highest encoding vector similarity to the target word segmentation training text but different text labels is used as the third negative sample text of the target word segmentation training text. In an optional embodiment of the present invention, by constructing positive and negative sample texts, each word segmentation training text has positive and negative contrast texts, thereby improving the difference between the various word segmentation training texts. For example, in the financial industry, positive and negative sample texts are constructed for the following word segmentation training texts containing text labels: Word segmentation training text 1 (pension insurance, work injury insurance, medical insurance), Word segmentation training sample 2 (car insurance, cargo insurance), Word segmentation training text 3 (pension insurance, work injury insurance), Word segmentation training text 4 (car insurance, liability insurance, medical insurance), Word segmentation training text 5 (pension insurance, work injury insurance, accident insurance), Word segmentation training text 6 (pension insurance, work injury insurance, medical insurance), Word segmentation training text 6 (pension insurance, work injury insurance, medical insurance), Word segmentation training text 7 (pension insurance, work injury insurance, accident insurance), Word segmentation training text 8 (pension insurance, work injury insurance, medical insurance), Word segmentation training text 9 (pension insurance, work injury insurance, medical insurance), Word segmentation training text 1 ... (Medical insurance), where, according to the construction rules of positive sample text, word segmentation training text 1 and word segmentation training text 6 are positive sample texts for each other; when training the model for medical insurance text tags, word segmentation training text 1 and word segmentation training text 4 are positive sample texts for each other; according to the construction rules of negative sample text, when training the model for medical insurance text tags, word segmentation training text 1 and word segmentation training text 3 are negative sample texts for each other, and word segmentation training text 3 and word segmentation training text 6 are negative sample texts for each other; when training the model for medical insurance text tags and accident insurance text tags, word segmentation training text 1 and word segmentation training text 5 are negative sample texts for each other.
[0075] Furthermore, as an optional embodiment of the present invention, reference is made to... Figure 2 As shown, the data augmentation processing performed on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text includes:
[0076] S11. Back-translate the target word segmentation training text using a preset translation dictionary to obtain the back-translated target word segmentation training text;
[0077] S12. Extract text keywords from the target word segmentation training text;
[0078] S13. Match the text keywords with phrases in a preset thesaurus to obtain successfully matched thesaurus phrases;
[0079] S14. Replace the text keywords in the target word segmentation training text with the successfully matched synonym groups to obtain the replaced target word segmentation training text;
[0080] S15. Integrate the back-translated target segmentation training text and the replaced target segmentation training text to obtain the third positive sample text of the target segmentation training text.
[0081] In this embodiment of the invention, the preset translation dictionary can be a publicly available and commonly used translation software. The preset thesaurus consists of the phrase itself and its synonyms; for example, the synonym for age can be years.
[0082] In an optional embodiment of the present invention, translation software can be used to translate the Chinese target word segmentation training text into English target word segmentation training text, and then another translation software can be used to translate the English target word segmentation training text back into Chinese target word segmentation training text, thus obtaining the back-translated target word segmentation training text. This enriches the diversity of positive sample texts for the target word segmentation training text. For example, a certain sentence in the pension insurance description text, “This insurance product accepts the insured age range of 60 years and above,” can be translated into “This insurance product accepts the insured age range is above 60 years old” using publicly available Chinese-English translation software. Then, by using different translation software, the back-translated pension insurance description text, “This insurance product accepts the insured age range of 60 years and above,” can be obtained.
[0083] In another optional embodiment of the present invention, the target word segmentation training text and the word segmentation training text can be encoded using a preset BERT model to obtain target text encoding vectors and word segmentation text encoding vectors. Further, the similarity between the target text encoding vector and the word segmentation text encoding vector is calculated using the cosine similarity formula to obtain the encoding vector similarity. Then, a third negative sample text of the target word segmentation training text is selected from the word segmentation training text.
[0084] S2. Using a preset text classification model, predict the text similarity between the word segmentation training text and the positive sample text and the negative sample text, respectively.
[0085] In this embodiment of the invention, the preset text classification model can be a BERT model (Bidirectional Encoder Representations from Transformer). In addition, the preset text classification model can be applied to a variety of different fields, such as: insurance description text classification, financial product description text classification, and medical text classification.
[0086] In this embodiment of the invention, a preset text classification model is used to predict the text similarity between the word segmentation training text and the positive sample text and the negative sample text, respectively. This enables the preset text classification model to be trained using a contrastive learning method, so that the model can obtain better differential representation when encoding text.
[0087] Furthermore, as an optional embodiment of the present invention, reference is made to... Figure 3 As shown, the step of using a preset text classification model to predict the text similarity between the segmented training text and the positive sample text and the negative sample text includes:
[0088] S21. Add the corresponding classification identifier to the beginning of the word segmentation training text, the positive sample text, and the negative sample text respectively to obtain the word segmentation training text, the positive sample text, and the negative sample text with classification identifiers.
[0089] S22. The word segmentation training text, positive sample text and negative sample text with classification identifiers are encoded by the encoding layer in the preset text classification model to obtain the word segmentation text classification identifier encoding vector, the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector.
[0090] S23. Using the multilayer perceptron module in the preset text classification model, calculate the similarity between the word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector, respectively, to obtain the text similarity between the word segmentation training text and the positive sample text and the negative sample text.
[0091] In this embodiment of the invention, the encoding layer can be a module for extracting text data representation vectors. The classification identifier can be a special label [CLS]. The multilayer perceptron module can be an artificial neural network consisting of input, output, and hidden layers.
[0092] In an optional embodiment of the present invention, there are word segmentation training text “This insurance product includes pension insurance, work injury insurance and medical insurance”, positive sample text “This insurance product includes pension insurance, work injury insurance and medical insurance”, and negative sample text “This insurance product includes pension insurance and work injury insurance”. First, classification identifiers are added to the word segmentation training text, the positive sample text and the negative sample text respectively to obtain word segmentation training text with classification identifier “[CLS]This insurance product includes pension insurance, work injury insurance and medical insurance”, positive sample text with classification identifier “[CLS]This insurance product includes pension insurance, work injury insurance and medical insurance”, and negative sample text with classification identifier “[CLS]This insurance product includes pension insurance and work injury insurance”.
[0093] Furthermore, in an optional embodiment of the present invention, the segmented training text, the positive sample text, and the negative sample text with classification identifiers are encoded to obtain vector representations of the segmented training text, the positive sample text, and the negative sample text. This makes it easier and more effective to calculate the text similarity between the segmented training text and the positive sample text and the negative sample text, thereby improving the efficiency of text classification model training.
[0094] Further, as another optional embodiment of the present invention, the step of using the multilayer perceptron module in the preset text classification model to calculate the similarity between the segmented text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector, respectively, to obtain the text similarity between the segmented training text and the positive sample text and the negative sample text, includes:
[0095] The word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector are vector-connected using the preset connection function in the multilayer perceptron module to obtain a positive sample comparison connection vector; and the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector are vector-connected using the connection function to obtain a negative sample comparison connection vector.
[0096] The first preset parameter in the multilayer perceptron module is used to multiply the positive sample comparison connection vector and the negative sample comparison connection vector respectively to obtain the positive comparison multiplication vector and the negative comparison multiplication vector;
[0097] Using the preset activation function in the multilayer perceptron module, the positive sample similarity between the word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector in the positive contrast multiplication vector, and the negative sample similarity between the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector in the negative contrast multiplication vector are calculated respectively.
[0098] The positive sample similarity and the negative sample similarity are normalized using the second preset parameters in the multilayer perceptron module to obtain the text similarity between the word segmentation training text and the positive sample text and the negative sample text.
[0099] In this embodiment of the invention, the preset connection function can be the concat function. The preset activation function can be the sigmoid function.
[0100] In detail, embodiments of the present invention can calculate the similarity between the segmented text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector using the following similarity calculation formula:
[0101] S = W2 * σ(W1 * concat(q, d))
[0102] Wherein, S represents the similarity between the word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector or the negative sample classification identifier encoding vector; W2 represents the second preset parameter; σ represents the preset activation function; W1 represents the first preset parameter; concat represents the preset connection function; q represents the word segmentation text classification identifier encoding vector; and d represents the positive sample classification identifier encoding vector or the negative sample classification identifier encoding vector.
[0103] S3. Calculate the first loss value between the similarity and the preset target similarity using a preset first loss function, and adjust the parameters of the preset text classification model according to the first loss value until the first loss value is less than the first preset threshold, thereby obtaining a preliminarily trained text classification model.
[0104] In this embodiment of the invention, the preset first loss function refers to a function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of the random event. The preset target similarity refers to the text similarity between the segmented training text and the positive and negative sample texts, set manually. For example, in the financial insurance industry, researchers can set the target similarity of segmented training text 1 (pension insurance, work injury insurance, medical insurance) and its corresponding positive sample text—segmented training text 6 (pension insurance, work injury insurance, medical insurance)—to 1, and the target similarity of segmented training text 1 (pension insurance, work injury insurance, medical insurance) and its corresponding negative sample text—segmented training text 5 (pension insurance, work injury insurance, accident insurance)—to 0. The first preset threshold refers to the maximum acceptable loss value of the preset text classification model set by researchers according to business needs.
[0105] In an optional embodiment of the present invention, a first loss value is calculated between the similarity and the preset target similarity using a preset first loss function. The training effect of the preset text classification model is then determined to have reached the expected target by judging whether the first loss value is less than a first preset threshold, thereby achieving initial training of the text classification model. When the first loss value is still greater than the first preset threshold, the parameter values in the preset text classification model need to be optimized and adjusted, and the process returns to the step of predicting the text similarity between the word segmentation training text and the positive sample text and the negative sample text using the preset text classification model, until the first loss value is less than the first preset threshold.
[0106] S4. Based on the number of text tags, adjust the number of multilayer perceptrons in the multilayer perceptron module of the initially trained text classification model to obtain an improved multilayer perceptron module, and use the initially trained text classification model to extract the feature vector of the word segmentation training text to obtain the word segmentation text feature vector.
[0107] In an optional embodiment of the present invention, since a single multilayer perceptron can typically only be trained on one text label, if the trained text classification model is to be used to classify new segmented text using text labels, each text label needs to be trained accordingly. Therefore, the number of multilayer perceptrons in the multilayer perceptron module of the initially trained text classification model needs to be adjusted to match the number of text labels. For example, in the financial insurance industry, the segmented training text 1 contains three labels: pension insurance, work injury insurance, and medical insurance. Therefore, the number of multilayer perceptrons in the multilayer perceptron module of the traditional text classification model can be adjusted to three so that each text label can be trained accordingly, thereby ensuring that the trained text classification model can achieve the task of classifying text labels on segmented text data.
[0108] Furthermore, in an optional embodiment of the present invention, since positive and negative sample texts are input for comparison when training the text classification model, when extracting the feature vector of the word segmentation training text using the preliminarily trained text classification model, it is also necessary to fill the input of the positive and negative sample texts with filler text, so that the word segmentation text feature vector extracted by the preliminarily trained text classification model is more accurate.
[0109] In detail, as an optional embodiment of the present invention, the step of extracting the feature vector of the word segmentation training text using the pre-trained text classification model to obtain the word segmentation text feature vector includes:
[0110] The word segmentation training text is encoded using the encoding layer in the preliminarily trained text classification model to obtain the word segmentation text encoding vector;
[0111] The word segmentation text encoding vector is subjected to position index encoding to obtain the word segmentation text position encoding vector;
[0112] The word segmentation text position encoding vector is combined with the word segmentation text encoding vector to obtain the word segmentation text feature vector.
[0113] In this embodiment of the invention, the encoder may be a device that encodes and converts data into a signal form that can be used for communication, transmission and storage.
[0114] This invention improves the feature vector by performing vector encoding, position encoding, and vector combination operations on the segmented training text, thereby increasing the feature differences between the segmented training texts and improving the accuracy of text classification.
[0115] S5. Calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector using the improved multilayer perceptron module.
[0116] In this embodiment of the invention, the predicted label may be a text label of the segmented text feature vector obtained by the improved multilayer perceptron module.
[0117] In this embodiment of the invention, the improved multilayer perceptron module is used to calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector. This facilitates the subsequent comparison and adjustment of the parameters in the improved multilayer perceptron module, thereby obtaining the trained text classification model.
[0118] Further, as an optional embodiment of the present invention, the step of calculating the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector using the improved multilayer perceptron module includes:
[0119] The third, fourth and fifth preset parameters in the improved multilayer perceptron module are used to perform linear transformations on the word segmentation text feature vectors to obtain query vectors, key vectors and numerical vectors.
[0120] The similarity matrix is obtained by multiplying the query vector with the transpose of the key vector.
[0121] The similarity matrix is normalized to obtain a normalized matrix;
[0122] The activation matrix is obtained by performing activation calculation on the normalized matrix;
[0123] The dot product of the activation matrix and the numerical vector yields the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector.
[0124] In this embodiment of the invention, the third, fourth and fifth preset parameters can be an improved multilayer perceptron module parameter matrix obtained through multiple training and optimization processes.
[0125] In an optional embodiment of the present invention, by performing score normalization calculation on the similarity matrix, the matrix gradient becomes more stable and reliable, thereby improving the calculation efficiency of the probability distribution of the text labels of the corresponding word segmentation training text in the word segmentation text feature vector.
[0126] S6. Calculate the second loss value between the predicted probability and the preset target probability using a preset second loss function, and adjust the parameters of the improved multilayer perceptron module according to the second loss value until the second loss value is less than the second preset threshold, thereby obtaining the trained text classification model.
[0127] In this embodiment of the invention, the preset second loss function refers to a function that maps the values of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of the random event. The preset target probability refers to the probability distribution of known text tags for each segmented training text. For example, the probability distribution of the segmented training text (pension insurance, work injury insurance, medical insurance) is such that the probability is 1 in the multilayer perceptron layers corresponding to the three text tags (pension insurance, work injury insurance, and medical insurance), and 0 in the other multilayer perceptron layers. The first preset threshold refers to the maximum acceptable loss value set by researchers for the improved multilayer perceptron module based on business needs.
[0128] In an optional embodiment of the present invention, a second loss value is calculated between the predicted probability and the preset target probability using a preset second loss function. The training effect of the improved multilayer perceptron module is then determined to have reached the expected target by judging whether the second loss value is less than a second preset threshold, thereby achieving complete training of the text classification model. When the second loss value is still greater than the second preset threshold, the parameter values in the improved multilayer perceptron module need to be optimized and adjusted, and the step of calculating the predicted probability of the text label of the word segmentation training text corresponding to the word segmentation text feature vector using the improved multilayer perceptron module is returned until the second loss value is less than the second preset threshold.
[0129] This invention constructs positive and negative sample texts for word segmentation training and trains a preset text classification model using these texts. This results in a better difference performance when extracting features from financial text data. Furthermore, the preset text classification model predicts the text similarity between the word segmentation training text and the positive and negative sample texts. The parameters of the preset text classification model are adjusted based on the loss value between the text similarity and a preset target similarity, ensuring the first loss value is less than a first preset threshold. This yields a pre-trained text classification model with stronger text representation capabilities, thus improving text classification accuracy. Even further, the number of layers in the multilayer perceptron module of the trained text classification model is adjusted, and the improved multilayer perceptron module is trained to obtain a trained multilayer perceptron module. This improves the accuracy of text classification using the multilayer perceptron module, resulting in more accurate classification of financial texts. Therefore, the text classification model training method provided by this invention can improve the accuracy of financial text classification models, thereby making it more convenient and effective for financial workers to find relevant financial texts.
[0130] like Figure 4 The diagram shown is a functional block diagram of the training device for the text classification model of this invention.
[0131] The text classification model training device 100 of the present invention can be installed in an electronic device. Depending on the functions implemented, the text classification model training device 100 may include a positive and negative sample construction module 101, a preliminary model training module 102, and a complete model training module 103. The module mentioned in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0132] In this embodiment, the functions of each module / unit are as follows:
[0133] The positive and negative sample construction module 101 is used to obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules.
[0134] The model preliminary training module 102 is used to predict the text similarity between the word segmentation training text and the positive sample text and the negative sample text using a preset text classification model, calculate the first loss value between the similarity and the preset target similarity using a preset first loss function, and adjust the parameters of the preset text classification model according to the first loss value until the first loss value is less than a first preset threshold, so as to obtain a text classification model that has been preliminarily trained.
[0135] The complete model training module 103 is used to adjust the number of multilayer perceptron modules in the initially trained text classification model according to the number of text tags, to obtain an improved multilayer perceptron module, and to extract the feature vector of the segmented training text using the initially trained text classification model, to obtain the segmented text feature vector. The improved multilayer perceptron module is used to calculate the predicted probability of the text tag of the segmented training text corresponding to the segmented text feature vector, and to calculate the second loss value between the predicted probability and the preset target probability using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, to obtain the trained text classification model.
[0136] like Figure 5 The diagram shown is a schematic representation of the electronic device used in this invention to implement the training method for a text classification model.
[0137] The electronic device may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a training program for a text classification model.
[0138] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Furthermore, the memory 11 can include both internal and external storage units of the electronic device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device, such as the code of a text classification model training program, but also to temporarily store data that has been output or will be output.
[0139] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules (such as training programs for text classification models) stored in the memory 11, and calls data stored in the memory 11 to perform various functions of the electronic device and process data.
[0140] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The communication bus 12 is configured to enable communication between the memory 11 and at least one processor 10, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus.
[0141] Figure 5 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 5 The structure shown does not constitute a limitation on the electronic device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0142] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0143] Optionally, the communication interface 13 may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device and other electronic devices.
[0144] Optionally, the communication interface 13 may further include a user interface, which may be a display, an input unit (such as a keyboard), or, optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.
[0145] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0146] The training program for the text classification model stored in the memory 11 of the electronic device is a combination of multiple computer programs. When run in the processor 10, it can achieve the following:
[0147] Obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules;
[0148] The text similarity between the word segmentation training text and the positive sample text and the negative sample text is predicted using a preset text classification model.
[0149] The first loss value between the similarity and the preset target similarity is calculated using a preset first loss function, and the parameters of the preset text classification model are adjusted according to the first loss value until the first loss value is less than a first preset threshold, thus obtaining a text classification model that has been initially trained.
[0150] Based on the number of text tags, the number of multilayer perceptron modules in the initially trained text classification model is adjusted to obtain an improved multilayer perceptron module. The feature vector of the word segmentation training text is then extracted using the initially trained text classification model to obtain the word segmentation text feature vector.
[0151] The improved multilayer perceptron module is used to calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector;
[0152] The second loss value between the predicted probability and the preset target probability is calculated using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, thus obtaining a trained text classification model.
[0153] Specifically, the processor 10's implementation method of the above-mentioned computer program can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0154] Furthermore, if the modules / units integrated into the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or volatile. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0155] Embodiments of the present invention may also provide a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following:
[0156] Obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules;
[0157] The text similarity between the word segmentation training text and the positive sample text and the negative sample text is predicted using a preset text classification model.
[0158] The first loss value between the similarity and the preset target similarity is calculated using a preset first loss function, and the parameters of the preset text classification model are adjusted according to the first loss value until the first loss value is less than a first preset threshold, thus obtaining a text classification model that has been initially trained.
[0159] Based on the number of text tags, the number of multilayer perceptron modules in the initially trained text classification model is adjusted to obtain an improved multilayer perceptron module. The feature vector of the word segmentation training text is then extracted using the initially trained text classification model to obtain the word segmentation text feature vector.
[0160] The improved multilayer perceptron module is used to calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector;
[0161] The second loss value between the predicted probability and the preset target probability is calculated using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, thus obtaining a trained text classification model.
[0162] Furthermore, the computer's usable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store the operating system, applications required for at least one function, etc.; and the data storage area may store data created based on the use of blockchain nodes, etc.
[0163] In the several embodiments provided by this invention, it should be understood that the disclosed electronic devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0164] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0165] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0166] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0167] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0168] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0169] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. The term "second class" is used to indicate names and does not indicate any specific order.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A training method for a text classification model, characterized in that, The method includes: Obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules; The text similarity between the word segmentation training text and the positive sample text and the negative sample text is predicted using a preset text classification model. The first loss value between the similarity and the preset target similarity is calculated using a preset first loss function, and the parameters of the preset text classification model are adjusted according to the first loss value until the first loss value is less than a first preset threshold, thus obtaining a text classification model that has been initially trained. Based on the number of text tags, the number of multilayer perceptron modules in the initially trained text classification model is adjusted to obtain an improved multilayer perceptron module. The feature vector of the word segmentation training text is then extracted using the initially trained text classification model to obtain the word segmentation text feature vector. The improved multilayer perceptron module is used to calculate the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector; The second loss value between the predicted probability and the preset target probability is calculated using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, thus obtaining a trained text classification model.
2. The training method for the text classification model as described in claim 1, characterized in that, The step of predicting the text similarity between the segmented training text and the positive sample text and the negative sample text using a preset text classification model includes: Add corresponding classification identifiers to the words of the word segmentation training text, the positive sample text, and the negative sample text respectively to obtain word segmentation training text, positive sample text, and negative sample text with classification identifiers; The word segmentation training text, positive sample text, and negative sample text with classification identifiers are encoded using the encoding layer in the preset text classification model to obtain the word segmentation text classification identifier encoding vector, the positive sample classification identifier encoding vector, and the negative sample classification identifier encoding vector. The similarity between the segmented text classification identifier encoding vector and the positive sample classification identifier encoding vector and the negative sample classification identifier encoding vector is calculated using the multilayer perceptron module in the preset text classification model, respectively, to obtain the text similarity between the segmented training text and the positive sample text and the negative sample text.
3. The training method for the text classification model as described in claim 2, characterized in that, The step of using the multilayer perceptron module in the preset text classification model to calculate the similarity between the segmented text classification identifier encoding vector and the encoding vectors of the positive sample classification identifier and the encoding vectors of the negative sample classification identifier, respectively, to obtain the text similarity between the segmented training text and the positive sample text and the negative sample text, includes: The word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector are vector-connected using the preset connection function in the multilayer perceptron module to obtain a positive sample comparison connection vector; and the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector are vector-connected using the connection function to obtain a negative sample comparison connection vector. The first preset parameter in the multilayer perceptron module is used to multiply the positive sample comparison connection vector and the negative sample comparison connection vector respectively to obtain the positive comparison multiplication vector and the negative comparison multiplication vector; Using the preset activation function in the multilayer perceptron module, the positive sample similarity between the word segmentation text classification identifier encoding vector and the positive sample classification identifier encoding vector in the positive contrast multiplication vector, and the negative sample similarity between the word segmentation text classification identifier encoding vector and the negative sample classification identifier encoding vector in the negative contrast multiplication vector are calculated respectively. The positive sample similarity and the negative sample similarity are normalized using the second preset parameters in the multilayer perceptron module to obtain the text similarity between the word segmentation training text and the positive sample text and the negative sample text.
4. The training method for the text classification model as described in claim 1, characterized in that, The step of constructing positive and negative sample texts of the word segmentation training text according to preset rules includes: Select one segmented training text from the segmented training texts arbitrarily and without repetition as the target segmented training text; The segmented training texts that have the same text labels as the target segmented training text are used as the first positive sample texts of the target segmented training text. Randomly select a text label from the text labels corresponding to the target word segmentation training text as the text label to be trained, and take the word segmentation training text with the text label to be trained as the second positive sample text of the target word segmentation training text. Data augmentation processing is performed on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text; The word segmentation training text that has an inclusion relationship with the text tags in the target word segmentation training text and is trained for different text tags is used as the first negative sample text of the target word segmentation training text. The word segmentation training text that is completely identical to the text label except for the text label to be trained is used as the second negative sample text of the target word segmentation training text; The word segmentation training text that has the highest encoding vector similarity to the target word segmentation training text and has a different text label is used as the third negative sample text of the target word segmentation training text.
5. The training method for the text classification model as described in claim 4, characterized in that, The step of performing data augmentation processing on the target word segmentation training text to obtain the third positive sample text of the target word segmentation training text includes: The target word segmentation training text is back-translated using a preset translation dictionary to obtain the back-translated target word segmentation training text. Extract text keywords from the target word segmentation training text; The text keywords are matched with phrases in a preset thesaurus to obtain successfully matched thesaurus phrases; The text keywords in the target word segmentation training text are replaced with the successfully matched synonym phrases to obtain the replaced target word segmentation training text. By integrating the back-translated target segmentation training text and the replaced target segmentation training text, a third positive sample text of the target segmentation training text is obtained.
6. The training method for the text classification model as described in claim 1, characterized in that, The step of extracting feature vectors from the word-segmented training text using the pre-trained text classification model to obtain word-segmented text feature vectors includes: The word segmentation training text is encoded using the encoding layer in the preliminarily trained text classification model to obtain the word segmentation text encoding vector; The word segmentation text encoding vector is subjected to position index encoding to obtain the word segmentation text position encoding vector; The word segmentation text position encoding vector is combined with the word segmentation text encoding vector to obtain the word segmentation text feature vector.
7. The training method for the text classification model as described in claim 1, characterized in that, The step of calculating the predicted probability of the text label of the word segmentation training text corresponding to the word segmentation text feature vector using the improved multilayer perceptron module includes: The third, fourth and fifth preset parameters in the improved multilayer perceptron module are used to perform linear transformations on the word segmentation text feature vectors to obtain query vectors, key vectors and numerical vectors. The similarity matrix is obtained by multiplying the query vector with the transpose of the key vector. The similarity matrix is normalized to obtain a normalized matrix; The activation matrix is obtained by performing activation calculation on the normalized matrix; The dot product of the activation matrix and the numerical vector yields the predicted probability of the text label of the segmented training text corresponding to the segmented text feature vector.
8. A training device for a text classification model, characterized in that, The device includes: The positive and negative sample construction module is used to obtain the word segmentation training text and the text tags corresponding to the word segmentation training text, and construct the positive sample text and negative sample text of the word segmentation training text according to preset rules; The model preliminary training module is used to predict the text similarity between the word segmentation training text and the positive sample text and the negative sample text using a preset text classification model, calculate the first loss value between the similarity and the preset target similarity using a preset first loss function, and adjust the parameters of the preset text classification model according to the first loss value until the first loss value is less than a first preset threshold, so as to obtain the text classification model that has been preliminarily trained. The complete model training module is used to adjust the number of multilayer perceptron modules in the initially trained text classification model according to the number of text tags, to obtain an improved multilayer perceptron module, and to extract the feature vector of the segmented training text using the initially trained text classification model, to obtain the segmented text feature vector. The improved multilayer perceptron module is used to calculate the predicted probability of the text tag of the segmented training text corresponding to the segmented text feature vector, and to calculate the second loss value between the predicted probability and the preset target probability using a preset second loss function. The parameters of the improved multilayer perceptron module are adjusted according to the second loss value until the second loss value is less than a second preset threshold, to obtain the trained text classification model.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor to enable the at least one processor to perform the training method of the text classification model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the text classification model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Text classification method and system based on label information and text characteristics and medium
CN109492101A
Short text classification method and system introducing negative example similarity information
CN114648061A