Limit tag text classification method and device, computer equipment and storage medium

By clustering and pre-training the label description text of the classification label, combined with the comparative learning training of positive and negative sample pairs, the problem of low accuracy caused by semantic ambiguity in the classification of limit label text is solved, and the effect of improving classification accuracy is achieved.

CN120217065APending Publication Date: 2025-06-27SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311812640.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the limit label text classification, due to the increase in the number of classification labels, there are semantically fuzzy classification labels, making it difficult to accurately capture subtle label semantic differences, thereby reducing the classification accuracy of the model.

Method used

By clustering the label description text of the classification label, K label text clusters are obtained, and the label description text in each label text cluster is input into the text classification model for pre-training, and the pre-trained classification model is obtained. Then, the positive label and negative label of the training text sample are determined, and the positive sample pair and negative sample pair are constructed. Then, the pre-trained classification model is compared and learned to train based on the constructed positive sample pair and negative sample pair to obtain the target classification model.

Benefits of technology

The text classification model is trained through the K label text clusters obtained by clustering, and a pre-trained classification model is obtained with good convergence. Then, a positive sample pair and a negative sample pair are constructed to compare and train the pre-trained classification model to obtain a target classification model with high recognition accuracy, which improves the accuracy of limit label text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217065A_ABST
    Figure CN120217065A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a limit tag text classification method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the clustering of tag description texts of classification tags, and obtaining K tag text clusters, respectively inputting the label description text in each label text cluster into a text classification model for pre-training to obtain a pre-trained classification model, then determining a positive label and a negative label of a training text sample, constructing a positive sample pair and a negative sample pair, and performing comparative learning training on the pre-trained classification model to obtain a positive sample pair and a negative sample pair; and obtaining a target classification model for classifying the target text. It can be seen that the text classification model is trained through the K label text clusters obtained through clustering, the pre-training classification model with good convergence is obtained, then the positive sample pair and the negative sample pair are constructed to conduct comparative learning training on the pre-training classification model, and the target classification model with high recognition precision is obtained; the purpose of improving the accuracy of limit tag text classification can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular, to an extreme label text classification method, device, computer device, and storage medium. Background Art

[0002] With the continuous development of society, there is an explosive growth of a large amount of text data, and furthermore, the classification labels for text classification are also increasing. How to improve the accuracy of text classification has become one of the key problems urgently to be addressed in the field of natural language processing technology.

[0003] Extreme label text classification refers to a classification problem with a very large number of categories. Generally speaking, the number of categories is greater than 1000, which leads to the need to find the classification category with the highest matching degree or relevance to a piece of text among at least 1000 classification categories.

[0004] Currently, in extreme label text classification, due to the continuous increase in the number of classification labels, there are inevitably classification labels with semantic ambiguity. Therefore, when facing extreme label classification, it is difficult to accurately select the correct label from the classification labels with semantic ambiguity because it is impossible to accurately capture the subtle label semantic differences, resulting in a low classification accuracy of the model. Summary of the Invention

[0005] Embodiments of the present invention provide an extreme label text classification method, device, computer device, and storage medium to solve the problem of low accuracy in extreme label text classification.

[0006] An extreme label text classification method, the method comprising:

[0007] Obtaining a training text sample and a label description text corresponding to a classification label;

[0008] Clustering the label description text corresponding to the classification label to obtain K label text clusters;

[0009] Respectively inputting the label text clusters into a text classification model for pre-training to obtain a pre-trained classification model;

[0010] Respectively determining the positive label and the negative label of the training text sample, and constructing positive sample pairs and negative sample pairs;

[0011] Inputting the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target text.

[0012] For the above method, optionally, the respectively determining the positive label and the negative label of the training text sample, and constructing positive sample pairs and negative sample pairs, includes:

[0013] Obtain the training text vector of the training text sample and the label text vector of the label description text;

[0014] Calculate the target similarity between the training text vector and each label text vector respectively;

[0015] Determine the positive label of the training sample, and screen out the negative labels from the classification labels according to the magnitude of the target similarity;

[0016] Construct the positive sample pair and the negative sample pair by combining the positive label and the negative label with the training text sample respectively.

[0017] For the above method, optionally, the determining the positive label of the training sample, and screening out the negative labels from the classification labels according to the magnitude of the target similarity includes:

[0018] Manually give the positive label of the training text sample;

[0019] Determine the similarity distribution relationship between the classification labels according to the magnitude of the target similarity;

[0020] Uniformly sample the classification labels according to the similarity distribution relationship to obtain a preset number of the classification labels as negative labels.

[0021] For the above method, optionally, the negative sample pair includes hard negative sample pairs;

[0022] Before constructing the positive sample pair and the negative sample pair by combining the positive label and the negative label with the training text sample respectively, it further includes:

[0023] According to the hierarchical relationship between the classification labels, select one or more classification labels from the sibling nodes of the node where the positive label is located as hard labels; the hard labels are used to construct the hard negative sample pair with the training text sample.

[0024] For the above method, optionally, clustering the label description texts corresponding to the classification labels to obtain K label text clusters includes:

[0025] Perform word segmentation on the label description of each classification label respectively to obtain multiple word tokens corresponding to the classification label;

[0026] Perform feature extraction on the word tokens to obtain the word token feature values corresponding to each word token;

[0027] Calculate the average feature value of the word token feature values corresponding to each classification label respectively;

[0028] Using a clustering algorithm, cluster the label description texts corresponding to the classification labels according to the magnitudes of the feature average values to obtain K label text clusters.

[0029] For the above method, optionally, the step of inputting the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts includes:

[0030] Input the positive sample pairs and the corresponding negative sample pairs into the pre-trained classification model respectively to obtain training text vectors, positive label vectors, and negative label vectors;

[0031] Calculate the loss value for the training text vectors, the positive label vectors, and the negative label vectors to obtain a triplet loss value;

[0032] When the triplet loss value is greater than 0, update the model parameters of the pre-trained classification model according to the triplet loss value to continue the contrastive learning training for the pre-trained classification model;

[0033] When the triplet loss value converges, obtain the trained target classification model.

[0034] For the above method, optionally, the target classification model includes a two-tower model.

[0035] An extreme label text classification device, characterized in that the device includes:

[0036] A text acquisition unit, configured to acquire training text samples and label description texts corresponding to classification labels;

[0037] A label clustering unit, configured to cluster the label description texts corresponding to the classification labels to obtain K label text clusters;

[0038] A pre-training unit, configured to input the label text clusters into a text classification model for pre-training respectively to obtain a pre-trained classification model;

[0039] A sample pair construction unit, configured to respectively determine the positive labels and negative labels of the training text samples and construct positive sample pairs and negative sample pairs;

[0040] A model training unit, configured to input the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts.

[0041] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for classifying extreme label texts described in any one of the above is implemented.

[0042] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for classifying extreme label texts described in any one of the above is implemented.

[0043] For the above extreme label text classification method, device, computer device and storage medium, by clustering the label description texts of classification labels, K label text clusters are obtained. Respectively, based on the label description texts in each label text cluster, they are input into a text classification model for pre-training to obtain a pre-trained classification model. Then, the positive labels and negative labels of the training text samples are determined, and positive sample pairs and negative sample pairs are constructed. Furthermore, the pre-trained classification model is subjected to contrastive learning training according to the constructed positive sample pairs and negative sample pairs to obtain a target classification model for classifying target texts. It can be seen that in the present invention, the K label text clusters obtained by clustering are used to train the text classification model to obtain a pre-trained classification model with better convergence. Then, positive sample pairs and negative sample pairs are constructed to perform contrastive learning training on the pre-trained classification model to obtain a target classification model with higher recognition accuracy, which can achieve the purpose of improving the accuracy of extreme label text classification. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart of an implementation of the extreme label text classification method in an embodiment of the present invention;

[0046] Figure 2 It is a partial flowchart of the implementation of the extreme label text classification method in an embodiment of the present invention;

[0047] Figure 3 It is a partial flowchart of the implementation of the extreme label text classification method in an embodiment of the present invention;

[0048] Figure 4 It is a partial flowchart of the implementation of the extreme label text classification method in an embodiment of the present invention;

[0049] Figure 5It is a partial implementation flowchart of the extreme label text classification method in an embodiment of the present invention;

[0050] Figure 6 It is a partial implementation flowchart of the extreme label text classification method in an embodiment of the present invention;

[0051] Figure 7 It is a structural schematic diagram of the extreme label text classification device in an embodiment of the present invention;

[0052] Figure 8 It is a structural schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0055] It should also be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0056] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0057] In addition, in the description of the specification and appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance.

[0058] References to "one embodiment" or "some embodiments" etc. described in the specification of the present invention mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0059] The present invention discloses an extreme label text classification method, device, computer device and storage medium. By clustering the label description texts of classification labels, K label text clusters are obtained. Based on the label description texts in each label text cluster respectively, they are input into a text classification model for pre-training to obtain a pre-trained classification model. Then, the positive labels and negative labels of training text samples are determined, and positive sample pairs and negative sample pairs are constructed. Furthermore, the pre-trained classification model is contrastively trained according to the constructed positive sample pairs and negative sample pairs to obtain a target classification model for classifying target texts. It can be seen that the present invention trains the text classification model through the K label text clusters obtained by clustering to obtain a pre-trained classification model with better convergence. Then, positive sample pairs and negative sample pairs are constructed to contrastively train the pre-trained classification model to obtain a target classification model with higher recognition accuracy, which can achieve the purpose of improving the accuracy of extreme label text classification. The following is illustrated by specific embodiments.

[0060] As Figure 1 shown, it is a flowchart of the implementation of an extreme label text classification method disclosed in an embodiment of the present invention. This method is applicable to electronic devices with natural language processing capabilities, such as devices like mobile phones, tablet computers, personal computers or servers, etc. The extreme label text classification method in this embodiment can specifically include the following steps:

[0061] S101: Obtain the training text samples and the label description texts corresponding to the classification labels.

[0062] Specifically, in this embodiment, a large number of texts can be obtained from a text database as training text samples, and label description texts of a large number of at least 1000 classification labels are prepared. The label description texts of up to millions of classification labels can be used to train the target classification model.

[0063] It should be understood that when the number of label description texts of classification labels reaches more than 1000, there will inevitably be label description texts with relatively close content, that is, there are semantically ambiguous texts between the label description texts. How to enable the model to learn and accurately capture the differences between the semantically ambiguous label description texts is one of the main problems to be solved by the present invention.

[0064] S102: Cluster the label description texts corresponding to the classification labels to obtain K label text clusters.

[0065] Cluster the label description texts corresponding to all classification labels through a classification algorithm to obtain K label text clusters. Among them, there is a relatively high semantic similarity between the label description texts in each label text cluster, and the label description texts in the label text cluster correspond to the clustering category of the label text cluster.

[0066] In a specific implementation, in this embodiment, the label description texts corresponding to the classification labels can be clustered by the KMeans clustering algorithm, and according to the set K value, the label description texts corresponding to the classification labels are clustered into K label text clusters.

[0067] Specifically, in this embodiment, the label description texts corresponding to the classification labels can be clustered through the following steps to obtain K label text clusters, specifically as Figure 2 shown:

[0068] S201: Perform word segmentation processing on the label descriptions of each classification label respectively to obtain multiple word elements corresponding to the classification label.

[0069] Perform word segmentation processing on the label description text through a word segmentation tool to obtain multiple word elements corresponding to the classification label, and so on, to obtain multiple word elements corresponding to each classification label.

[0070] In a specific implementation, the word segmentation tools used for word segmentation processing in this embodiment include but are not limited to jieba word segmentation tool, SnowNLP word segmentation tool, THULAC word segmentation tool, and NLPIR word segmentation tool, etc., and no specific limitation is made in this embodiment.

[0071] For example, taking "I want to go to Peking University" as an example, after performing word segmentation processing, four word elements "I", "want to go", "Beijing", and "University" can be obtained. Thus, word segmentation processing of the label description text can be realized to obtain multiple word elements corresponding to the classification label.

[0072] S202: Extract features from the word elements to obtain the word element feature values corresponding to each word element.

[0073] For multiple tokens corresponding to each classification label, feature extraction is performed on each token respectively to obtain token feature values. And so on, token feature values of each token corresponding to each classification label are obtained for subsequent steps of clustering the label description text based on the obtained token feature values.

[0074] Specifically, in this embodiment, the word2vec model can be used to perform feature extraction on multiple tokens corresponding to the classification label to obtain a feature matrix, and then the token feature values of each token are calculated according to the feature matrix.

[0075] S203: Calculate the feature average values of the token feature values corresponding to each classification label respectively.

[0076] For each classification label, after obtaining the token feature values of each token corresponding to the classification label, the feature average value is calculated, that is, the token feature values of each token corresponding to the classification label are summed up and then divided by the number of tokens to obtain the feature average value.

[0077] S204: Through a clustering algorithm, cluster the label description texts corresponding to the classification labels according to the magnitudes of the feature average values to obtain K label text clusters.

[0078] Through the clustering algorithm, determine K clustering center points, and then cluster according to each clustering center point respectively according to the magnitudes of the feature average values to obtain K label text clusters.

[0079] Specifically, the clustering algorithm in this embodiment can be the KMeans clustering algorithm. Determine K clustering center points through the KMeans clustering algorithm, and then calculate the distances from the feature average values of each label description text to each clustering center point, and then assign the label description text to the clustering center point with the closest distance to it. Thus, K label text clusters are obtained. In addition, new clustering center points can be recalculated based on the obtained K label text clusters, and then the distances from the feature average values of each label description text to each new clustering center point are recalculated to obtain K new label text clusters. And so on, until the K label text clusters meet the convergence condition, and the obtained K label text clusters are used as the final output result of the clustering algorithm. It can be understood that the label text clusters in this embodiment are obtained through the KMeans clustering algorithm. Therefore, the label description texts in each label text cluster are texts with relatively high similarity.

[0080] In addition, the value of K in this embodiment can be specified according to actual needs or determined through testing to obtain a suitable K value. For example, through hierarchical clustering, different K values are used to pre-classify the sample set to obtain classification results, and then the classification results for different K values are evaluated. The quality of the classification results can be evaluated by the average diameter of the classes. Usually, when the number of classes decreases, the average diameter will increase; when the number of classes increases and exceeds a certain value, the average diameter will remain unchanged, and this value is the optimal K value. Thus, a suitable K value can be determined.

[0081] In summary, in this embodiment, through the KMeans clustering algorithm, K labeled text clusters are obtained for K-class learning of the text classification model, which can preheat and fine-tune the text classification model so that the text classification model can converge in subsequent learning, achieving the purpose of improving the classification performance and training efficiency of the model.

[0082] S103: Input the labeled text clusters into the text classification model for pre-training to obtain a pre-trained classification model.

[0083] The obtained K labeled text clusters are respectively input into the text classification model for pre-training. One labeled text cluster can be regarded as a batch, and each batch is sent into the text classification model for pre-training. After each batch is sent into the text classification model and the training is completed, the text classification model has completed one epoch of training. And so on, the text classification model is trained for N1 epochs to obtain a pre-trained classification model that has completed pre-training. Among them, the value of N1 in this embodiment can be set according to actual needs and is not limited in this embodiment.

[0084] In specific implementation, the text classification model in this embodiment can be a two-tower model. The obtained K labeled text clusters are respectively input into the two-tower model for pre-training, and the two-tower model is preheated and fine-tuned during the pre-training process to obtain a pre-trained model, that is, the two-tower model after pre-training.

[0085] S104: Determine the positive label and negative label of the training text sample respectively to construct positive sample pairs and negative sample pairs.

[0086] Among them, the positive label refers to the correct classification label of the training text sample, and the negative label can be any other classification label except the positive label of the training text sample.

[0087] After determining the positive label and negative label of the training text sample, label the positive label for the training text sample to obtain positive sample pairs, and label the negative label for the training text sample to obtain negative sample pairs.

[0088] In a specific implementation, in this embodiment, the positive and negative labels of the training text samples can be determined manually, or the positive and negative labels of the training text samples can be determined by calculating the similarity. Specifically, it can be as follows:

[0089] First, in this embodiment, the negative label can be determined by the occurrence frequency of the classification labels.

[0090] Manually determine the positive label of each training text sample, and screen out the classification labels with a lower occurrence frequency in the positive labels as the negative labels, including the classification labels with an occurrence frequency of 0 in the positive labels, that is, the classification labels that never appear in the positive labels of the training text samples. For example, those with an occurrence frequency of 20 times or less are determined to have a low occurrence frequency. Accordingly, a set of labels with a low occurrence frequency can be obtained, and a classification label is randomly selected from the label set as the negative label. Accordingly, the positive and negative labels of the training text samples can be determined.

[0091] Second, in this embodiment, the distance from the label description text of the classification label to the training text sample can be calculated, and the classification label with the farthest distance is screened out as the negative label of the training text.

[0092] Specifically, the distances in this embodiment include but are not limited to cosine similarity, Euclidean distance, or Manhattan distance. The following takes the calculation of cosine similarity as an example for illustration:

[0093] Obtain the sample vector of the training text sample and the label text vector of the label description text of each classification label, and then calculate the cosine similarity between the training text vector and each label vector respectively through the cosine similarity calculation formula, and then screen out the classification label corresponding to the label vector with the smallest cosine similarity as the negative label of the training text sample. In addition, in this embodiment, the method of manual annotation can be used to annotate a classification label for the training text sample as the positive label.

[0094] S105: Input the positive sample pair and the negative sample pair into the pre-trained classification model for contrast training to obtain a target classification model for classifying the target text.

[0095] Input positive sample pairs and negative sample pairs into a pre-trained classification model for contrast training. During the training process, calculate the model loss function. When the model loss function meets the preset conditions, determine that the pre-training of the classification model is completed to obtain the target classification model. After that, the target text and the label description text of the classification label can be input into the target classification model to classify the target text. For example, after inputting the target text and the label description text of the classification label into the target classification model, obtain the target vector of the target text and the label text vector of the label description text, calculate the similarity between the target vector and each label text vector, and select the classification label corresponding to the label text vector with the largest similarity as the classification category of the target text.

[0096] In a specific implementation, the model loss function in this embodiment can be a triplet loss function. During the training process of the pre-trained classification model, use the triplet loss function to determine whether the pre-trained classification model is sufficiently trained. When the pre-trained classification model is sufficiently trained, obtain the target classification model.

[0097] Furthermore, in this embodiment, the following steps can be used to determine whether the pre-trained classification model is sufficiently trained through the triplet loss function, as Figure 3 shown:

[0098] S301: Input the positive sample pair and the corresponding negative sample pair into the pre-trained classification model respectively to obtain the training text vector, the positive label vector, and the negative label vector.

[0099] It should be understood that the training text samples in the positive sample pair and the corresponding negative sample are the same. Input the training text sample, the label description text corresponding to the positive label in the positive sample pair, and the label description text corresponding to the negative label in the negative sample pair into the pre-trained classification model to obtain the training text vector, the positive label vector, and the negative label vector.

[0100] In a specific implementation, the pre-trained classification model in this embodiment is a two-tower model. The two-tower model has two model branches. Input the training text sample into one of the model branches for encoding to obtain the training text vector, and input the label description text corresponding to the positive label and the label description text corresponding to the negative label into the other model branch for encoding to obtain the positive label vector and the negative label vector.

[0101] S302: Calculate the loss value for the training text vector, the positive label vector, and the negative label vector to obtain the triplet loss value.

[0102] S303: When the triplet loss value is greater than 0, update the model parameters of the pre-trained classification model according to the triplet loss value to continue the contrast training of the pre-trained classification model.

[0103] Input the training text vector, positive label vector, and negative label vector into the triplet loss value calculation formula to obtain the triplet loss value.

[0104] Specifically, the triplet loss value calculation formula in this embodiment can be shown as follows:

[0105] L(a, p, n) = max(||f(a) - f(p)|| 2 - ||f(a) - f(n)|| 2 + α, 0)

[0106] where α is a constant, which is the distance between the positive label vector and the negative label vector, a represents the text vector, p represents the positive label vector, and n represents the negative label vector.

[0107] It should be understood that the distance between the positive label vector and the training text vector should be less than the distance between the negative label vector and the training text vector. Therefore, when the difference between the distance between the positive label vector and the training text vector and the distance between the negative label vector and the training text vector is greater than α, the triplet loss value is 0, and there is no need to update or adjust the parameters of the pre-trained classification model; when the difference between the distance between the positive label vector and the training text vector and the distance between the negative label vector and the training text vector is less than α, the triplet loss value is a positive number, indicating that the vector obtained by encoding the pre-trained classification model does not meet the expectation. It is necessary to adjust the parameters of the pre-trained classification model, and input the training text sample, the label description text corresponding to the positive label, and the label description text corresponding to the negative label into the pre-trained classification model for training, that is, update the parameters of the pre-trained classification model according to the triplet loss value to continue the contrast training of the pre-trained classification model.

[0108] S304: When the triplet loss value converges, obtain the trained target classification model.

[0109] Specifically, in this embodiment, the positive sample pair and the corresponding negative sample pair can be input into the pre-trained classification model for iterative training, and when the triplet loss value is greater than 0, continuously update the parameters of the pre-trained classification model until the triplet loss value converges, and determine the trained pre-trained classification model as the target classification model.

[0110] In summary, the present invention discloses an extreme label text classification method. This method clusters the label description texts of classification labels to obtain K label text clusters, and respectively inputs the label description texts in each label text cluster into a text classification model for pre-training to obtain a pre-trained classification model. Then, the positive labels and negative labels of the training text samples are determined, positive sample pairs and negative sample pairs are constructed, and further, the pre-trained classification model is contrastively trained according to the constructed positive sample pairs and negative sample pairs to obtain a target classification model for classifying target texts. It can be seen that the present invention trains the text classification model through the K label text clusters obtained by clustering to obtain a pre-trained classification model with better convergence, and then constructs positive sample pairs and negative sample pairs to contrastively train the pre-trained classification model to obtain a target classification model with higher recognition accuracy, which can achieve the purpose of improving the accuracy of extreme label text classification.

[0111] Based on Figure 1 In the specific implementation of, step S104 can be specifically implemented as the following steps, specifically as Figure 4 shown:

[0112] S401: Obtain the training text vector of the training text sample and the label text vector of the label description text.

[0113] Generate the training text vector from the training text sample through vector representation technology, and generate the label text vector from the label description text.

[0114] Specifically, the vector representation technology in this embodiment includes but is not limited to TF-IDF, word embeddings or BERT. Among them, taking BERT as an example for illustration, the training text sample and the label description text are respectively input into the BERT model to obtain the training text vector corresponding to the training text sample and the label text vector corresponding to the label description text.

[0115] S402: Calculate the target similarity between the training text vector and each label text vector respectively.

[0116] For each training text vector, input the training text vector and each label text vector into the similarity calculation formula respectively to calculate the target similarity between the training text vector and each label text vector.

[0117] In the specific implementation, the target similarity in this embodiment can be the pre-similarity. For each training text vector, input the training text vector and each label text vector into the cosine similarity calculation formula respectively to calculate the cosine similarity between the training text vector and each label text vector.

[0118] S403: Determine the positive label of the training text sample, and screen out the negative labels from the classification labels according to the magnitude of the target similarity.

[0119] Among them, the positive label of the training text sample can be directly given a positive label manually, and then according to the magnitude of the target similarity, the specified classification label is selected as the negative label.

[0120] Specifically, in this embodiment, the positive label and the negative label can be obtained through the following steps, as Figure 5 shown:

[0121] S501: Manually give the positive label of the training text sample.

[0122] In a specific implementation, the positive label in this embodiment can be obtained by manually annotating the training text sample. That is to say, a classification label is manually selected for the training text sample, and this classification label is marked as the positive label of the training text.

[0123] S502: Determine the similarity distribution relationship between classification labels according to the magnitude of the target similarity.

[0124] Specifically, in this embodiment, each classification label can be sorted in descending order or ascending order according to the magnitude of the target similarity, thereby obtaining the similarity distribution relationship between classification labels. That is to say, according to the ascending or descending order of the target similarity, the classification labels are sorted to obtain the order relationship between classification labels, that is, to determine the similarity distribution relationship between classification labels.

[0125] For example, taking the descending order as an example, the greater the target similarity, the more forward the corresponding classification label sequence will be. Conversely, the smaller the target similarity, the more backward the corresponding classification label sequence will be. Thus, the similarity distribution relationship between classification labels is obtained.

[0126] It should be understood that for each training text vector, the cosine similarity between the training text vector and each label text vector will be calculated respectively, and then the similarity distribution relationship between the corresponding classification labels is obtained. The subsequent steps are executed respectively according to the similarity distribution relationship between the classification labels corresponding to each training text sample.

[0127] S503: Uniformly sample the classification labels according to the similarity distribution relationship to obtain a preset number of classification labels as negative labels.

[0128] In a specific implementation, in this embodiment, the classification labels can be uniformly sampled according to the number of classification labels according to the similarity distribution relationship to obtain a preset number of classification labels as negative labels, or the similarity distribution relationship can be divided into intervals to obtain multiple similarity sub-intervals, and then the classification labels in each similarity sub-interval are randomly sampled respectively to obtain a preset number of classification labels as negative labels.

[0129] For example, taking 1001 classification labels as an example, among which there is 1 classification label as the positive label of the training text sample, and 20 classification labels need to be sampled from the remaining 1000 classification labels as negative labels. Then, one classification label can be sampled every 49 classification labels to obtain a classification label as a negative label, that is, the 50th classification label is the negative label. Accordingly, uniform sampling of classification labels can be achieved according to the similarity distribution relationship, and a preset number of classification labels can be obtained as negative labels.

[0130] Again, taking 1001 classification labels as an example, among which there is 1 classification label as the positive label of the training text sample, and 20 classification labels need to be sampled from the remaining 1000 classification labels as negative labels. Then, according to the order of the target similarity, every 50 classification labels can be divided into a similarity sub-interval, and then one classification label can be randomly sampled from each similarity sub-interval as a negative label. Accordingly, uniform sampling of classification labels can be achieved according to the similarity distribution relationship, and a preset number of classification labels can be obtained as negative labels.

[0131] S404: Construct positive sample pairs and negative sample pairs by combining the positive label and the negative label with the training text sample respectively.

[0132] It can be understood that the positive sample pair in this embodiment is the sample pair composed of the positive label and the training text sample, and the negative sample pair is the sample pair composed of the negative label and the training text sample.

[0133] Specifically, in this embodiment, positive sample pairs and negative sample pairs can be obtained by annotating the training text sample with positive labels or negative labels, and then the obtained positive sample pairs and negative sample pairs can be used for training the pre-trained classification model.

[0134] In summary, in this embodiment, positive labels and negative labels are determined by calculating the target similarity, and then positive sample pairs and negative sample pairs are constructed, and special processing is performed on the difficult negative sample pairs in the negative sample pairs, which is beneficial to improving the ability of the model to discriminate similar semantic classification labels.

[0135] In one implementation, the negative sample pairs in this embodiment include difficult negative sample pairs, and the following steps may also be included before step S404, such as Figure 6 shown as:

[0136] S405: According to the hierarchical relationship between classification labels, select one or more classification labels from the sibling nodes of the node where the positive label is located as difficult labels.

[0137] Among them, the difficult labels are used to construct difficult negative sample pairs with the training text sample.

[0138] It should be understood that the hard negative sample pair includes a training text sample and a hard label, while the positive sample includes a training text sample and a positive label, where the label description text of the hard label is very close to the label description text of the positive label. Additionally, when the number of classification labels reaches a certain level, there will be a certain hierarchical relationship between the classification labels. Accordingly, in this hierarchical relationship, the label description texts of the classification labels included within the sibling nodes of the same parent node are all very close, or there is semantic ambiguity. Therefore, the classification label within the node where the positive label is located can be selected as the hard label, or the classification label within the sibling node of the node where the positive label is located can be selected as the hard label.

[0139] In summary, in this embodiment, by uniformly sampling the classification labels according to the similarity distribution relationship, it can be ensured that there are hard negative labels among the negative labels. Through the hard negative sample pairs composed of the hard negative labels and the training text samples, the pre-trained classification model can be contrastively trained, enabling the pre-trained classification model to more accurately capture the semantic differences between the labels. Especially in scenarios with fuzzy boundaries, it is beneficial to improve the classification accuracy of the finally trained target classification model.

[0140] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0141] As Figure 7 shown, it is a schematic structural diagram of an extreme label text classification device disclosed in an embodiment of the present invention. This device is applicable to electronic devices with natural language processing capabilities, such as mobile phones, tablet computers, personal computers, or servers, etc.

[0142] The device in this embodiment may specifically include the following units:

[0143] Text acquisition unit 701: used to acquire the training text sample and the label description text corresponding to the classification label;

[0144] Label clustering unit 702, used to cluster the label description texts corresponding to the classification labels to obtain K label text clusters;

[0145] Pre-training unit 703, used to respectively input the label text clusters into the text classification model for pre-training to obtain the pre-trained classification model;

[0146] Sample pair construction unit 704, used to respectively determine the positive label and the negative label of the training text sample, and construct positive sample pairs and negative sample pairs;

[0147] The model training unit 705 is configured to input positive sample pairs and negative sample pairs into a pre-trained classification model for contrast training to obtain a target classification model for classifying target texts.

[0148] In summary, the present invention discloses an extreme label text classification device. By clustering the label description texts of classification labels, K label text clusters are obtained. Based on the label description texts in each label text cluster, they are respectively input into a text classification model for pre-training to obtain a pre-trained classification model. Then, the positive label and negative label of the training text sample are determined, and positive sample pairs and negative sample pairs are constructed. Furthermore, based on the constructed positive sample pairs and negative sample pairs, the pre-trained classification model is subjected to contrast training to obtain a target classification model for classifying target texts. It can be seen that the present invention trains the text classification model through the K label text clusters obtained by clustering to obtain a pre-trained classification model with better convergence. Then, positive sample pairs and negative sample pairs are constructed to pre-train the classification model for contrast training to obtain a target classification model with higher recognition accuracy, which can achieve the purpose of improving the accuracy of extreme label text classification.

[0149] In one implementation, the sample pair construction unit 704 can be used to:

[0150] Obtain the training text vector of the training text sample and the label text vector of the label description text;

[0151] Calculate the target similarity between the training text vector and each label text vector respectively;

[0152] Determine the positive label of the training text sample, and screen out negative labels from the classification labels according to the magnitude of the target similarity;

[0153] Construct positive sample pairs and negative sample pairs by combining the positive label and negative label with the training text sample respectively.

[0154] In one implementation, the sample pair construction unit 704 can also be used to:

[0155] Manually assign the positive label of the training text sample;

[0156] Determine the similarity distribution relationship between classification labels according to the magnitude of the target similarity;

[0157] Uniformly sample the classification labels according to the similarity distribution relationship to obtain a preset number of classification labels as negative labels.

[0158] In one implementation, the negative sample pairs include hard negative sample pairs; the sample pair construction unit 704 can also be used to:

[0159] According to the hierarchical relationship between classification labels, select one or more classification labels from the sibling nodes of the node where the positive label is located as difficult labels; the difficult labels are used to construct difficult negative sample pairs with training text samples.

[0160] In one implementation, the label clustering unit 702 can be used for:

[0161] Tokenize the label descriptions of each classification label respectively to obtain multiple tokens corresponding to the classification label;

[0162] Perform feature extraction processing on the tokens to obtain the token feature values corresponding to each token;

[0163] Calculate the average feature values of the token feature values corresponding to each classification label respectively;

[0164] Through the clustering algorithm, cluster the label description texts corresponding to the classification labels according to the magnitude of the average feature values to obtain K label text clusters.

[0165] In one implementation, the model training unit 705 can be used for:

[0166] Input the positive sample pairs and the corresponding negative sample pairs into the pre-trained classification model respectively to obtain training text vectors, positive label vectors and negative label vectors;

[0167] Calculate the loss value for the training text vectors, positive label vectors and negative label vectors to obtain a triplet loss value;

[0168] When the triplet loss value is greater than 0, update the model parameters of the pre-trained classification model to continue the contrast training of the pre-trained classification model;

[0169] When the triplet loss value converges, obtain the trained target classification model.

[0170] In one implementation, the target classification model includes a two-tower model.

[0171] For the specific limitations of the extreme label text classification device, reference can be made to the relevant limitations of the extreme label text classification method in the above text, which will not be elaborated here. Each module in the above extreme label text classification device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0172] In one implementation, the embodiments of the present application disclose a computer device, which can be a server, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an extreme label text classification method.

[0173] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0174] Obtain a training text sample and a label description text corresponding to a classification label;

[0175] Cluster the label description texts corresponding to the classification labels to obtain K label text clusters;

[0176] Input the label text clusters into a text classification model for pre-training to obtain a pre-trained classification model;

[0177] Determine the positive label and negative label of the training text sample respectively, and construct positive sample pairs and negative sample pairs;

[0178] Input the positive sample pairs and negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts.

[0179] In one implementation, determining the positive label and negative label of the training text sample respectively and constructing positive sample pairs and negative sample pairs includes:

[0180] Obtain the training text vector of the training text sample and the label text vector of the label description text;

[0181] Calculate the target similarity between the training text vector and each label text vector respectively;

[0182] Determine the positive label of the training text sample, and screen out the negative label from the classification labels according to the magnitude of the target similarity;

[0183] Construct positive sample pairs and negative sample pairs by combining the positive label and negative label with the training text sample respectively.

[0184] In one implementation, determining the positive label of the training sample and screening out the negative label from the classification labels according to the magnitude of the target similarity includes:

[0185] Manually assign positive labels to the training text samples;

[0186] Determine the similarity distribution relationship between classification labels according to the magnitude of the target similarity;

[0187] Uniformly sample the classification labels according to the similarity distribution relationship to obtain a preset number of classification labels as negative labels.

[0188] In one implementation, the negative sample pairs include hard negative sample pairs;

[0189] Before constructing positive sample pairs and negative sample pairs by respectively combining positive labels and negative labels with the training text samples, it further includes:

[0190] According to the hierarchical relationship between classification labels, select one or more classification labels from the sibling nodes of the node where the positive label is located as hard labels; the hard labels are used to construct hard negative sample pairs with the training text samples.

[0191] In one implementation, cluster the label description texts corresponding to the classification labels to obtain K label text clusters, including:

[0192] Perform word segmentation on the label descriptions of each classification label respectively to obtain multiple word tokens corresponding to the classification labels;

[0193] Perform feature extraction processing on the word tokens to obtain the word token feature values corresponding to each word token;

[0194] Calculate the average feature values of the word token feature values corresponding to each classification label respectively;

[0195] Through a clustering algorithm, cluster the label description texts corresponding to the classification labels according to the magnitude of the average feature values to obtain K label text clusters.

[0196] In one implementation, input the positive sample pairs and negative sample pairs into a pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts, including:

[0197] Input the positive sample pairs and the corresponding negative sample pairs into the pre-trained classification model respectively to obtain training text vectors, positive label vectors, and negative label vectors;

[0198] Calculate the loss value for the training text vectors, positive label vectors, and negative label vectors to obtain a triplet loss value;

[0199] When the triplet loss value is greater than 0, update the model parameters of the pre-trained classification model according to the triplet loss value to continue the contrastive learning training of the pre-trained classification model;

[0200] When the triplet loss value converges, obtain the trained target classification model.

[0201] In one implementation, the target classification model includes a two-tower model.

[0202] In one implementation, the embodiments of the present application disclose a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor in a computer device, the computer device is enabled to execute each step of any one of the embodiments of an extreme label text classification method disclosed in the present invention. The computer-readable storage medium can be non-volatile or volatile.

[0203] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0204] Obtain the training text samples and the label description texts corresponding to the classification labels;

[0205] Cluster the label description texts corresponding to the classification labels to obtain K label text clusters;

[0206] Input the label text clusters into the text classification model for pre-training to obtain a pre-trained classification model;

[0207] Determine the positive labels and negative labels of the training text samples respectively, and construct positive sample pairs and negative sample pairs;

[0208] Input the positive sample pairs and negative sample pairs into the pre-trained classification model for contrast training to obtain a target classification model for classifying target texts.

[0209] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0210] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0211] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A method for classifying extreme label texts, characterized in that, The method includes: Obtaining a training text sample and a label description text corresponding to a classification label; Clustering the label description texts corresponding to the classification labels to obtain K label text clusters; Respectively inputting the label text clusters into a text classification model for pre-training to obtain a pre-trained classification model; Respectively determining the positive label and negative label of the training text sample, and constructing positive sample pairs and negative sample pairs; Inputting the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts.

2. The extreme label text classification method according to claim 1, wherein The respectively determining the positive label and negative label of the training text sample, and constructing positive sample pairs and negative sample pairs includes: Obtaining a training text vector of the training text sample and a label text vector of the label description text; Respectively calculating the target similarity between the training text vector and each label text vector; Determining the positive label of the training text sample, and screening out the negative label from the classification labels according to the magnitude of the target similarity; Constructing the positive sample pairs and the negative sample pairs by respectively combining the positive label and the negative label with the training text sample.

3. The extreme label text classification method according to claim 2, characterized in that, The determining the positive label of the training sample sample, and screening out the negative label from the classification labels according to the magnitude of the target similarity includes: Manually giving the positive label of the training text sample; Determining the similarity distribution relationship between the classification labels according to the magnitude of the target similarity; Uniformly sampling the classification labels according to the similarity distribution relationship to obtain a preset number of the classification labels as negative labels.

4. The extreme label text classification method according to claim 2, wherein The negative sample pairs include hard negative sample pairs; Before the constructing the positive sample pairs and the negative sample pairs by respectively combining the positive label and the negative label with the training text sample, it further includes: According to the hierarchical relationship between the classification labels, selecting one or more classification labels from the sibling nodes of the node where the positive label is located as hard labels; The hard labels are used to construct the hard negative sample pairs with the training text sample.

5. The extreme label text classification method according to claim 1, characterized in that The clustering the label description texts corresponding to the classification labels to obtain K label text clusters includes: Respectively performing word segmentation on the label descriptions of each classification label to obtain a plurality of word tokens corresponding to the classification label; Performing feature extraction processing on the word tokens to obtain a word token feature value corresponding to each word token; Respectively calculating the feature average value of the word token feature values corresponding to each classification label; Through a clustering algorithm, clustering the label description texts corresponding to the classification labels according to the magnitude of the feature average value to obtain K label text clusters.

6. The extreme label text classification method according to claim 1, characterized in that The inputting the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts includes: Respectively inputting the positive sample pairs and the corresponding negative sample pairs into the pre-trained classification model to obtain a training text vector, a positive label vector, and a negative label vector; Calculate the loss value for the training text vector, the positive label vector, and the negative label vector to obtain the triplet loss value; When the triplet loss value is greater than 0, update the model parameters of the pre-trained classification model according to the triplet loss value to continue the contrastive learning training for the pre-trained classification model; When the triplet loss value converges, obtain the trained target classification model.

7. The extreme label text classification method according to any one of claims 1-6, characterized in that, The target classification model includes a two-tower model.

8. An extreme label text classification device, characterized in that, The device includes: A text acquisition unit for acquiring the training text sample and the label description text corresponding to the classification label; A label clustering unit for clustering the label description texts corresponding to the classification labels to obtain K label text clusters; A pre-training unit for inputting the label text clusters into a text classification model for pre-training to obtain a pre-trained classification model; A sample pair construction unit for respectively determining the positive label and the negative label of the training text sample and constructing positive sample pairs and negative sample pairs; A model training unit for inputting the positive sample pairs and the negative sample pairs into the pre-trained classification model for contrastive learning training to obtain a target classification model for classifying target texts.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the extreme label text classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the extreme label text classification method according to any one of claims 1 to 7.