Two-stage text classification method and device, computer equipment and storage medium

Through the two-stage text classification method, filtering and matching alternative tags in the limit label text classification, the semantic ambiguity caused by the huge number of classification categories is solved and the accuracy of text classification is improved.

CN120217064APending Publication Date: 2025-06-27SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311808597.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In extreme label text classification, due to the large number of classification categories, there is semantic ambiguity between classification categories, which in turn reduces the accuracy of text classification.

Method used

Using a two-stage text classification method, firstly, K alternative labels with the highest similarity to the target text are filtered out from the classification label through the first text encoder, and then the target text is constructed into a combined text with each alternative label, input it into the second text encoder, calculate the matching probability of each alternative label, and finally select the alternative label with the greatest matching probability as the classification category of the target text.

Benefits of technology

Through the two-stage text classification method, the accuracy of limit label text classification can be effectively improved and classification errors caused by semantic ambiguity can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217064A_ABST
    Figure CN120217064A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a two-stage text classification method, which comprises the following steps of: inputting a target text and a label description text of a classification label into a first text encoder to obtain K alternative labels with the highest similarity with the target text; respectively constructing the target text and the label description text of each alternative label to obtain K combined texts; inputting each combined text into a second text encoder to obtain a matching probability of the target text and each alternative label; and selecting the alternative label with the maximum matching probability as the classification category of the target text. According to the method, the alternative labels with relatively high similarity with the target text are screened out from the classification labels through the first text encoder, and then the target text and each alternative label are finely arranged through the second encoder, so that the matching probability of the target text and each alternative label is obtained; and selecting the alternative label with the highest matching probability as the classification category of the target text, so that the classification precision of the limit label text classification task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular, to a two-stage text classification method, apparatus, computer device, and storage medium. Background Art

[0002] With the continuous progress of science and technology, the explosive growth of a large amount of text data has brought new challenges to current text processing technologies. Among them, text classification, as one of the key tasks of text processing, has become the focus of attention in this field.

[0003] Extreme label text classification refers to a classification problem with a very large number of categories. Generally speaking, the number of categories is greater than 1000, which means that it is necessary to find the classification category with the highest matching degree or relevance to a piece of text among at least 1000 classification categories.

[0004] However, in the extreme label text classification problem, due to the extremely large number of classification categories, there are semantic ambiguity problems between many classification categories, resulting in low accuracy of extreme label text classification and difficulty in ensuring precision.

[0005] Therefore, there is a problem of low accuracy in extreme label text classification. Summary of the Invention

[0006] Embodiments of the present invention provide a two-stage text classification method, apparatus, computer device, and storage medium to solve the problem of how to improve the accuracy of extreme label text classification.

[0007] A two-stage text classification method, the method comprising:

[0008] Inputting the target text and the label description text of the classification label into a first text encoder to obtain K alternative labels with the highest similarity to the target text;

[0009] Constructing K combined texts by respectively combining the target text with the label description text of each of the alternative labels;

[0010] Inputting each of the combined texts into a second text encoder to obtain the matching probability between the target text and each of the alternative labels; the second text encoder and the first text encoder are different encoders;

[0011] Selecting the alternative label with the maximum matching probability as the classification category of the target text.

[0012] In the above method, optionally, the step of inputting the target text and the label description text of the classification label into a first text encoder to obtain K alternative labels with the highest similarity to the target text includes:

[0013] Input the target text and the label description text of the classification label into the first text encoder to obtain the target vector of the target text and the label vectors of each classification label;

[0014] Calculate the target vector respectively with each label vector to obtain the target similarity between the target vector and each label vector;

[0015] According to the magnitudes of the target similarities, screen out the top K classification labels with the highest similarity to the target vector as candidate labels.

[0016] For the above method, optionally, the step of calculating the target vector respectively with the label vectors of each classification label to obtain the target similarity between the target vector and the label vectors of each classification label includes:

[0017] Perform cosine similarity calculation between the target vector and the label vectors of each classification label by the encoder to obtain the first similarity;

[0018] Perform word frequency similarity calculation between the target text and the label description text of each classification label to obtain the second similarity;

[0019] Calculate the first similarity and the second similarity to obtain the target similarity.

[0020] For the above method, optionally, the first text encoder is trained in the following way:

[0021] Obtain training text samples and the corresponding positive labels and negative labels for the training text samples; the positive labels are the correct classification labels for the training text samples, and the negative labels are the incorrect classification labels for the training text samples;

[0022] Input the training text samples, the label description text corresponding to the positive labels, and the label description text corresponding to the negative labels into the first text encoder to obtain the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels;

[0023] Calculate the text vectors, the positive label vectors, and the negative label vectors to obtain a triplet loss value;

[0024] When the triplet loss value is greater than 0, update the parameters of the first text encoder according to the triplet loss value to continue training the first text encoder;

[0025] When the triplet loss value converges, obtain the trained first encoder.

[0026] In the above method, optionally, the step of inputting the training text sample, the label description text corresponding to the positive label, and the label description text corresponding to the negative label into the first text encoder to obtain the text vector of the training text sample, the positive label vector of the positive label, and the negative label vector of the negative label includes:

[0027] Input the training text sample into the first encoder branch of the first text encoder to obtain the text vector of the training text sample;

[0028] Input the label description text corresponding to the positive label and the label description text corresponding to the negative label into the second encoder branch of the first text encoder to obtain the positive label vector and the negative label vector.

[0029] In the above method, optionally, the second text encoder can be trained as follows:

[0030] Input the training text sample and the label description text of the classification label into the first text encoder to obtain N sample labels corresponding to the training text sample;

[0031] Construct N training sample pairs by combining the training text sample with each corresponding sample label;

[0032] Input the training sample pairs into the second text encoder for training to obtain the trained second text encoder.

[0033] In the above method, optionally, before inputting the training sample pairs into the second text encoder for training to obtain the trained second text encoder, it further includes:

[0034] Calculate the word frequency similarity between the training text sample and the label description text of the classification label to obtain M sample labels corresponding to the training text sample;

[0035] Construct M training sample pairs by combining the training text sample with each corresponding sample label; the training sample pairs are used to train the second text encoder.

[0036] A two-stage text classification device, the device includes:

[0037] An alternative label acquisition unit, configured to input the target text and the label description text of the classification label into the first text encoder to obtain the K alternative labels with the highest similarity to the target text;

[0038] A combined text acquisition unit, configured to construct K combined texts by combining the target text with the label description text of each alternative label;

[0039] A matching probability acquisition unit, configured to input each of the combined texts into a second text encoder to obtain the matching probability between the target text and each of the alternative labels; the second text encoder and the first text encoder are different encoders;

[0040] A classification category determination unit, configured to select the alternative label with the highest matching probability as the classification category of the target text.

[0041] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the two-stage text classification method described in any one of the above is implemented.

[0042] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the two-stage text classification method described in any one of the above is implemented.

[0043] In the above two-stage text classification method, device, computer device and storage medium, the first text encoder is used to screen out the K alternative labels with the highest similarity to the target text from the classification labels, and then the target text is respectively combined with each alternative label to obtain combined texts. The obtained combined texts are input into the second text encoder to obtain the alternative label with the highest matching probability with the target text, and this alternative label is used as the classification category of the target text. It can be seen that in this embodiment, the first text encoder is used to screen out the K alternative labels with the highest similarity to the target text (the alternative labels are also relatively similar to each other), and then the second text encoder calculates the matching probability of the combined texts constructed by the target text and each alternative label, and further uses the alternative label with the highest matching probability as the final classification category of the target text. Thus, the purpose of improving the classification accuracy of the extreme label text can be achieved. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a flowchart for implementing the two-stage text classification method in an embodiment of the present invention;

[0046] Figure 2 It is a partial flowchart for implementing the two-stage text classification method in an embodiment of the present invention;

[0047] Figure 3 It is a partial implementation flowchart of the two-stage text classification method in an embodiment of the present invention;

[0048] Figure 4 It is a partial implementation flowchart of the two-stage text classification method in an embodiment of the present invention;

[0049] Figure 5 It is a partial implementation flowchart of the two-stage text classification method in an embodiment of the present invention;

[0050] Figure 6 It is a partial implementation flowchart of the two-stage text classification method in an embodiment of the present invention;

[0051] Figure 7 It is a partial implementation flowchart of the two-stage text classification method in an embodiment of the present invention;

[0052] Figure 8 It is a schematic structural diagram of a two-stage text classification device in an embodiment of the present invention;

[0053] Figure 9 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] It should be understood that when used in the specification and the appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0056] It should also be understood that the term " / and / " used in the specification and the appended claims of the present invention refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0057] As used in the specification of the present invention and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be construed, depending on the context, to mean "once determined" or "in response to determining" or "once [described condition or event] is detected" or "in response to detecting [described condition or event]".

[0058] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0059] The reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0060] The present invention discloses a two-stage text classification method, apparatus, computer device, and storage medium. The method screens out the K candidate labels with the highest similarity to the target text from the classification labels through a first text encoder, then constructs combined texts by combining the target text with each candidate label respectively, inputs the obtained combined texts into a second text encoder to obtain the candidate label with the highest matching probability with the target text, and uses this candidate label as the classification category of the target text. It can be seen that in this embodiment, the first text encoder is used to screen out the K candidate labels with the highest similarity to the target text (the candidate labels are also relatively similar to each other), and then the second text encoder calculates the matching probability of the combined texts constructed by the target text and each candidate label, and further uses the candidate label with the highest matching probability as the final classification category of the target text. Thus, the purpose of improving the classification accuracy of the extreme label text can be achieved. The following is illustrated by specific embodiments.

[0061] The present invention screens out alternative labels with a high similarity to the target text from the classification labels through a first text encoder, and then uses a second encoder to perform fine ranking on the target text and each alternative label to obtain the matching probability between the target text and each alternative label, and selects the alternative label with the highest matching probability as the classification category of the target text. Thus, the text classification of the target text is completed through a two-stage text classification to improve the classification accuracy of the extreme label text classification task.

[0062] As Figure 1 shown, it is a flowchart of the implementation of a two-stage text classification method disclosed in an embodiment of the present invention. This method is applicable to electronic devices with text processing capabilities, such as mobile phones, tablet computers, laptop computers, personal computers, and servers, etc. The method in this embodiment can specifically include the following steps:

[0063] S101: Input the target text and the label description text of the classification label into the first text encoder to obtain the K alternative labels with the highest similarity to the target text.

[0064] Input the target text and the label description text of the classification label into the first text encoder to obtain the similarity between the target text and each classification label, and then select K classification labels from high to low according to the size of the similarity as the alternative labels of the target text.

[0065] Further, in this embodiment, the K alternative labels can be specifically obtained through the following steps, as Figure 2 shown:

[0066] S201: Input the target text and the label description text of the classification label into the first text encoder to obtain the target vector of the target text and the label vector of each classification label.

[0067] Input the target text into the first text encoder to enable the first text encoder to encode the target text to obtain the target vector corresponding to the target text; input the label description text of the classification label into the first text encoder to enable the first text encoder to encode the label description text of each classification label to obtain the label vector of each classification label.

[0068] Specifically, in the specific implementation, the first text encoder in this embodiment can be a dual encoder (DualEncoder). Input the target text into one encoder branch of the dual encoder to obtain the target vector of the target text, and input the label description text of the classification label into the other encoder branch of the dual encoder to obtain the label vector. Accordingly, the target vector of the target text and the label vector of each classification label can be obtained.

[0069] Among them, the dual encoder in this embodiment can be asFigure 3 As shown, the dual - tower encoder consists of two Roberta models. One is used as the input of the target text, and the other is used as the input of the label description text of the classification label. The triple - loss function is used to determine whether the dual - tower encoder is sufficiently trained during the training process.

[0070] S202: Calculate the target vector with the label vector of each classification label respectively to obtain the target similarity between the target vector and the label vector of each classification label.

[0071] Calculate the similarity between the target vector and each label vector in turn to obtain the similarity between the target vector and each label vector, and obtain the target similarity between the target and the label vector of each classification label.

[0072] In specific implementation, the target similarity in this embodiment includes but is not limited to cosine similarity, and may also include MB25 similarity, Manhattan distance, etc. This embodiment does not specifically limit the target similarity. That is to say, using other methods to calculate similarity to implement the method in this embodiment all fall within the protection scope of the present invention.

[0073] Furthermore, the target similarity in this embodiment can be specifically obtained through the following steps, as Figure 3 shown:

[0074] S301: Calculate the cosine similarity between the target vector and the label vector of each classification label to obtain the first similarity.

[0075] Calculate the cosine similarity between the target vector and each label vector in turn to obtain the similarity between the target vector and each label vector, and obtain the first similarity between the target and the label vector of each classification label.

[0076] In specific implementation, the method for calculating the first similarity in this embodiment includes but is not limited to cosine similarity calculation. This embodiment does not specifically limit the first similarity. That is to say, using other methods to calculate similarity to implement the method in this embodiment all fall within the protection scope of the present invention.

[0077] S302: Calculate the word - frequency similarity between the target text and the label description text of each classification label to obtain the second similarity.

[0078] In specific implementation, the method for calculating the first similarity in this embodiment includes but is not limited to BM25 similarity calculation. This embodiment does not specifically limit the first similarity. That is to say, using other methods to calculate similarity to implement the method in this embodiment all fall within the protection scope of the present invention.

[0079] S303: Calculate the first similarity and the second similarity to obtain the target similarity.

[0080] Input the first similarity and the second similarity into the target similarity calculation formula to obtain the target similarity.

[0081] Among them, the similarity calculation formula can be shown as follows:

[0082] S = W1 × S1 + W2 × S2

[0083] Among them, W1 and W2 represent weights, S1 represents the first similarity, S2 represents the second similarity, and S also represents the target similarity.

[0084] It should be understood that W1 and W2 in the similarity calculation formula can be set according to actual needs, and W1 and W2 can be set to be the same or different. For example, W1 and W2 can both be set to 0.5, or W1 can be set to 0.2 and W2 can be set to 0.8. In this embodiment, the specific values of W1 and W2 are not limited.

[0085] In summary, in the embodiment of the present invention, when calculating the target similarity, the BM25-based statistical representation of vocabulary and cosine similarity are referred to, which is beneficial to improving the accuracy of the first text encoder to obtain the alternative labels corresponding to the target text.

[0086] S203: According to the magnitude of the target similarity, screen out the K classification labels with the highest similarity to the target vector as alternative labels.

[0087] Sort the similarities between the target vector and each label vector in descending or ascending order. When sorting in descending order, sequentially select the classification labels corresponding to the K label vectors as the classification labels of the target text. When sorting in ascending order, inversely select the classification labels corresponding to the K label vectors as the classification labels of the target text.

[0088] It should be noted that the value of K in this embodiment can be set according to actual needs. The magnitude of K includes but is not limited to 10, 15, etc. No specific limitation is made in this embodiment.

[0089] S102: Construct K combined texts by respectively combining the target text with the label description texts of each alternative label.

[0090] Construct K combined texts by respectively combining the target text with the label description texts of each alternative label. The number of combined texts is the same as the number of alternative labels. Among them, the combined text can be understood as the target text pair composed of the target text and the label description text of the alternative label, that is, a pair of texts composed of the target text and the label description text of the alternative label.

[0091] In a specific implementation, in the combined text of this embodiment, the target text and the label description text can be separated by [SEP]. Thus, K combined texts can be constructed.

[0092] Among them, the combined text can be as follows:

[0093] Target text [SEP] Label description text

[0094] It can be seen that the above combined text separates the target text and the label description text by [SEP]. Thus, the first text encoder can determine whether the text in the combined text is the target text or the label description text according to [SEP].

[0095] S103: Input each combined text into the second text encoder to obtain the matching probability between the target text and each alternative label.

[0096] It should be understood that the second text encoder and the first text encoder are different encoders.

[0097] Obtain the target text in the combined text and the label description text of the sample label, convert the target text into a target vector, convert the label description text into a label vector, calculate the matching probability between the target vector and the label vector, and obtain the matching probability between the target text and the alternative label.

[0098] In a specific implementation, the second text encoder in this embodiment can be a CrossEncoder. Input the combined text into the CrossEncoder, and through the CrossEncoder, calculate the matching probability between the target text and the label description text in each input combined text, and perform the subsequent steps of text classification according to the magnitude of the matching probability.

[0099] S104: Select the alternative label with the highest matching probability as the classification category of the target text.

[0100] Sort the combined texts according to the magnitude of the matching probability, and select the alternative label corresponding to the label description text in the combined text with the highest matching probability as the classification category of the target text.

[0101] The present invention discloses a two-stage text classification method, apparatus, computer device, and storage medium. The method screens out the top K candidate labels with the highest similarity to the target text from the classification labels through a first text encoder, then constructs combined texts by combining the target text with each candidate label respectively, inputs the obtained combined texts into a second text encoder to obtain the candidate label with the highest matching probability with the target text, and uses this candidate label as the classification category of the target text. It can be seen that in this embodiment, the top K candidate labels with the highest similarity to the target text are screened out through the first text encoder (the candidate labels are also relatively similar to each other), and then the second text encoder calculates the matching probability for the combined texts constructed by the target text and each candidate label, and further uses the candidate label with the highest matching probability as the final classification category of the target text. Thus, the purpose of improving the classification accuracy of extreme label texts can be achieved.

[0102] In one implementation, the first text encoder in this embodiment can be trained in the following manner, as Figure 4 shown:

[0103] S401: Obtain training text samples and the corresponding positive labels and negative labels of the training text samples.

[0104] Among them, the positive label is the correct classification label of the training text sample, and the negative label is the incorrect classification label of the training text sample. It can be understood that in this embodiment, each training text sample is assigned a correct classification label (i.e., the positive label) and an incorrect classification label (i.e., the negative label). The negative label can refer to any other classification label except the positive label.

[0105] Specifically, in this embodiment, a positive label and a negative label can be manually labeled for each training sample. For example, given a training text sample, a positive label is labeled for the training text sample according to the specific text content of the training text sample, and then a classification label randomly selected from other classification labels except the positive label is used as the negative label. It should be noted that for a training text sample, a positive label can be determined to form a positive sample pair, and one or more negative labels can be randomly selected to form one or more negative sample pairs.

[0106] S402: Input the training text samples, the label description text corresponding to the positive labels, and the label description text corresponding to the negative labels into the first text encoder to obtain the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels.

[0107] Input the training text samples, the label description text corresponding to the positive label, and the label description text corresponding to the negative label into the first text encoder, so that the first text encoder encodes the training text samples, the label description text corresponding to the positive label, and the label description text corresponding to the negative label, and obtains the text vector of the training text sample, the positive label vector of the positive label, and the negative label vector of the negative label.

[0108] It should be understood that training the first text encoder means training the encoding ability of the first text encoder for the training text samples, the label description text corresponding to the positive label, and the label description text corresponding to the negative label, that is, the ability of the first text encoder to encode text into text vectors, and training the first text encoder with the training text samples, the label description text corresponding to the positive label, and the label description text corresponding to the negative label.

[0109] S403: Calculate the text vector, the positive label vector, and the negative label vector to obtain the triplet loss value.

[0110] S404: When the triplet loss value is greater than 0, update the parameters of the first text encoder according to the triplet loss value to continue training the first text encoder.

[0111] Use the text vector as the anchor point, the positive label vector as the positive example, and the negative label vector as the negative example and input them into the triplet loss function to obtain the triplet loss value.

[0112] Among them, the triplet loss function can be as follows:

[0113] L(a, p, n) = max(‖f(a) - f(p)‖ 2 - ‖f(a) - f(n)‖ 2 + α, 0)

[0114] Among them, α is a constant, which is the distance between the positive label vector and the negative label vector, a represents the text vector, p represents the positive label vector, and n represents the negative label vector.

[0115] It should be understood that the distance between the positive label vector and the text vector should be less than the distance between the negative label vector and the text vector. Therefore, when the difference between the distance between the positive label vector and the text vector and the distance between the negative label vector and the text vector is greater than α, the triplet loss value is 0, and there is no need to update or adjust the parameters of the first text encoder according to the triplet loss value; when the difference between the distance between the positive label vector and the text vector and the distance between the negative label vector and the text vector is less than α, the triplet loss value is a positive number, indicating that the vector obtained by encoding the first text encoder does not meet the expectation, and the parameters of the first text encoder need to be adjusted. Then, input the training text sample, the label description text corresponding to the positive label, and the label description text corresponding to the negative label into the first text encoder for training, that is, update the parameters of the first text encoder according to the triplet loss value to continue training the first text encoder.

[0116] S405: When the triplet loss value converges, the trained first text encoder is obtained.

[0117] Specifically, in this embodiment, the positive sample pair and the corresponding negative sample pair can be input into the first text encoder for iterative training. When the triplet loss value is greater than 0, continuously update the parameters of the first text encoder until the triplet loss value converges, and then determine that the training of the first text encoder is completed.

[0118] In Figure 4 the specific implementation, step S402 in this embodiment can be specifically implemented through the following steps, as Figure 5 shown:

[0119] S501: Input the training text sample into the first encoder branch of the first text encoder to obtain the text vector of the training text sample.

[0120] S502: Input the label description text corresponding to the positive label and the label description text corresponding to the negative label into the second encoder branch of the first text encoder to obtain the positive label vector and the negative label vector.

[0121] It should be understood that the first text encoder in this embodiment can have two encoder branches. One encoder branch is used to encode the training text sample to obtain the text vector, and the other encoder branch is used to encode the label description text corresponding to the positive label and the label description text corresponding to the negative label to obtain the positive label vector and the negative label vector.

[0122] In a specific implementation, the first question encoder in this embodiment can be implemented based on a two-tower encoder. The training text samples are input into the first encoder branch of the first text encoder, and the first encoder branch encodes the training text samples to obtain the text vectors of the training text samples. The label description text corresponding to the positive label and the label description text corresponding to the negative label are input into the second encoder branch of the first text encoder, so that the second encoder branch encodes the label description text corresponding to the positive label to obtain the positive label vector, and encodes the label description text corresponding to the negative label to obtain the negative label vector.

[0123] In addition, in this embodiment, two independent encoders can be used to encode the training text samples, the label description text of the positive label, and the label description text of the negative label. One encoder encodes the training text samples to obtain the text vectors, and the other encoder encodes the label description text of the positive label and the label description text of the negative label to obtain the positive label vector and the negative label vector.

[0124] It can be seen that in this embodiment, by encoding the training text samples respectively through the two branches of the two-tower encoder, and encoding the label text corresponding to the positive label and the label description text corresponding to the negative label, it is possible to learn the representations of queries and documents in a low-dimensional space, thereby making the similarity calculation more efficient.

[0125] Based on Figure 1 the specific implementation, the second text encoder in this embodiment can be trained in the following manner, as Figure 6 shown:

[0126] S601: Input the training text samples and the label description text of the classification labels into the first text encoder to obtain N sample labels corresponding to the training text samples.

[0127] Input the training text samples and the label description text of the classification labels into the first text encoder to obtain the similarity between the training text samples and each classification label, and then, according to the magnitude of the similarity, select N classification labels from high to low in sequence as the sample labels of the training text samples.

[0128] Input the training text samples and the label description text of the classification labels into the first text encoder to obtain the text vectors of the training text samples and the label vectors of each classification label, calculate the text vectors respectively with each label vector to obtain the sample similarities between the text vectors and each label vector, and according to the magnitude of the sample similarities, screen out the N classification labels with the highest similarity to the text vectors as the sample labels.

[0129] S602: Construct N training sample pairs by respectively combining the training text samples with each corresponding sample label.

[0130] N training sample pairs are constructed by respectively combining the training text samples with the label description texts of each sample label, and the number of training sample pairs is the same as the number of sample labels.

[0131] In a specific implementation, in the training sample pairs in this embodiment, the training text sample and the label description text in the training sample pair can be separated by [SEP]. Thus, N training sample pairs can be constructed.

[0132] S603: Input the training sample pairs into the second text encoder for training to obtain a trained second text encoder.

[0133] It should be understood that in this embodiment, the training text samples can be paired with each sample label to construct N training sample pairs. The sample labels are obtained based on calculating the similarity with the training text samples. Therefore, the similarity between the training text samples and the label descriptions of the sample labels is relatively high, that is, there is a certain semantic ambiguity between the training text samples and the label descriptions of the sample labels. Therefore, by using the training sample pairs with high semantic similarity to train the second text encoder, the second text encoder can directly learn the interactions and relationships between the inputs to generate accurate representations or matching probabilities, and a second text encoder with higher accuracy can be trained.

[0134] In a specific implementation, the second text encoder in this embodiment can be a cross encoder. Inputting the training text pairs into the cross encoder can enable the second text encoder to learn the interactions and relationships between the training text samples and the label description texts in the input training text pairs to obtain a trained cross encoder.

[0135] In addition, in this embodiment, when training the cross encoder, the binary classification loss function is calculated to determine whether the cross encoder is fully trained. For example, during the training of the cross encoder, when the loss value calculated by the binary classification loss function no longer decreases, it is determined that the cross encoder is fully trained.

[0136] It should be noted that the second text encoder in this embodiment includes but is not limited to the cross encoder. The solutions formed by using other encoders as the second text encoder in this embodiment all fall within the protection scope of the present invention.

[0137] Based on Figure 6 In a specific implementation, the following steps may further be included before step S603, as Figure 7 shown:

[0138] S604: Calculate the word frequency similarity between the training text samples and the label description texts of the classification labels to obtain M sample labels corresponding to the training text samples.

[0139] Calculate the word frequency similarity between the training text samples and the label description texts of each classification label respectively, and then sort each classification label according to the magnitude of the word frequency similarity, and select M classification labels as the sample labels of the training text samples.

[0140] In a specific implementation, the word frequency similarity in this embodiment can be the BM25 similarity. Calculate the word frequency similarity between the training text samples and the label description texts of each classification label respectively through the BM25 algorithm to obtain the word frequency similarity between the training text samples and each classification label, and then select M classification labels as the sample labels of the training text samples according to the word frequency similarity.

[0141] It should be noted that K, N, and M in this embodiment all refer to specific numerical values. The numerical magnitudes of K, N, and M can be set according to actual needs, and the numerical magnitudes of K, N, and M can be the same or different, which are not limited in this embodiment.

[0142] S605: Construct M training sample pairs by respectively combining the training text samples with each corresponding sample label.

[0143] Among them, the training sample pairs are used to train the second text encoder.

[0144] Construct M training sample pairs by respectively combining the training text samples with the label description texts of each sample label. The number of training sample pairs is the same as the number of sample labels.

[0145] In a specific implementation, the training text sample and the label description text in the training sample pairs in this embodiment can be separated by [SEP]. Thus, M training sample pairs can be constructed. Input the obtained M training sample pairs into the second text encoder for training to obtain the trained second text encoder.

[0146] In summary, in this embodiment, not only are training sample pairs obtained by calculating semantic similarity to enable the second text encoder to learn semantic-based statistical representation encoding, but also training sample pairs are obtained by calculating word frequency similarity to enable the second text encoder to learn vocabulary-based statistical representation encoding, which can greatly improve the classification accuracy of the second text encoder for extreme label text classification.

[0147] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0148] Such as Figure 8As shown in the figure, it is a structural diagram of a two-stage text classification device disclosed in an embodiment of the present invention. This device is applicable to electronic devices with text processing capabilities, such as mobile phones, tablet computers, laptop computers, personal computers, and servers, etc.

[0149] Specifically, the device in this embodiment may include the following units:

[0150] An alternative label acquisition unit 801, configured to input the target text and the label description text of the classification label into a first text encoder to obtain the K alternative labels with the highest similarity to the target text;

[0151] A combined text acquisition unit 802, configured to construct K combined texts by respectively combining the target text with the label description text of each alternative label;

[0152] A matching probability acquisition unit 803, configured to input each combined text into a second text encoder to obtain the matching probability between the target text and each alternative label; the second text encoder and the first text encoder are different encoders;

[0153] A classification category determination unit 804, configured to select the alternative label with the highest matching probability as the classification category of the target text.

[0154] In summary, the present invention discloses a two-stage text classification device. By using a first text encoder, K alternative labels with the highest similarity to the target text are screened out from the classification labels, and then the target text is respectively combined with each alternative label to construct combined texts. The obtained combined texts are input into a second text encoder to obtain the alternative label with the highest matching probability with the target text, and this alternative label is used as the classification category of the target text. It can be seen that in this embodiment, K alternative labels with the highest similarity to the target text (the alternative labels are also relatively similar to each other) are screened out through the first text encoder, and then the second text encoder calculates the matching probability of the combined texts constructed by the target text and each alternative label, and further uses the alternative label with the highest matching probability as the final classification category of the target text. Thus, the purpose of improving the classification accuracy of the extreme label text can be achieved.

[0155] In one implementation, the alternative label acquisition unit 801 may be configured to:

[0156] Input the target text and the label description text of the classification label into the first text encoder to obtain the target vector of the target text and the label vectors of each classification label;

[0157] Calculate the target vector respectively with each label vector to obtain the target similarity between the target vector and each label vector;

[0158] Select the top K classification labels with the highest similarity to the target vector as candidate labels according to the magnitude of the target similarity.

[0159] In one implementation, the candidate label acquisition unit 801 can be used for:

[0160] Perform cosine similarity calculation between the target vector and the label vectors of each classification label to obtain the first similarity;

[0161] Perform word frequency similarity calculation between the target text and the label description texts of each classification label to obtain the second similarity;

[0162] Calculate the first similarity and the second similarity to obtain the target similarity.

[0163] In one implementation, the first text encoder is trained as follows:

[0164] Obtain the training text samples and the corresponding positive labels and negative labels for the training text samples; the positive label is the correct classification label for the training text sample, and the negative label is the incorrect classification label for the training text sample;

[0165] Input the training text samples, the label description texts corresponding to the positive labels, and the label description texts corresponding to the negative labels into the first text encoder to obtain the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels;

[0166] Calculate the text vectors, the positive label vectors, and the negative label vectors to obtain the triplet loss value;

[0167] When the triplet loss value is greater than 0, update the parameters of the first text encoder according to the triplet loss value to continue training the first text encoder;

[0168] When the triplet loss value converges, obtain the trained first text encoder.

[0169] In one implementation, the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels can be obtained as follows:

[0170] Input the training text samples into the first encoder branch of the first text encoder to obtain the text vectors of the training text samples;

[0171] Input the label description texts corresponding to the positive labels and the label description texts corresponding to the negative labels into the second encoder branch of the first text encoder to obtain the positive label vectors and the negative label vectors.

[0172] In one implementation, the second text encoder is trained as follows:

[0173] Input the training text samples and the label description text of the classification labels into the first text encoder to obtain N sample labels corresponding to the training text samples;

[0174] Construct N training sample pairs by combining the training text samples with each corresponding sample label respectively;

[0175] Input the training sample pairs into the second text encoder for training to obtain a trained second text encoder.

[0176] In one implementation, the second text encoder can be trained as follows:

[0177] Calculate the word frequency similarity between the training text samples and the label description text of the classification labels to obtain M sample labels corresponding to the training text samples;

[0178] Construct M training sample pairs by combining the training text samples with each corresponding sample label respectively; the training sample pairs are used to train the second text encoder.

[0179] For the specific limitations of the two-stage text classification device, reference can be made to the relevant limitations of the two-stage text classification method in the above text, which will not be elaborated here. Each module in the above two-stage text classification device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0180] In one implementation, an embodiment of the present application discloses a computer device, which can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a two-stage text classification method.

[0181] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0182] Input the target text and the label description text of the classification label into the first text encoder to obtain the top K alternative labels with the highest similarity to the target text;

[0183] Construct K combined texts by combining the target text with the label description text of each alternative label respectively;

[0184] Input each combined text into the second text encoder to obtain the matching probability between the target text and each alternative label; the second text encoder and the first text encoder are different encoders;

[0185] Select the alternative label with the highest matching probability as the classification category of the target text.

[0186] In one implementation, before inputting the training sample pairs in the first sample set into the second text encoder for training to obtain the trained second text encoder, it further includes:

[0187] Calculate the word frequency similarity between the training text sample and the label description text of the classification label to obtain M sample labels corresponding to the training text sample;

[0188] Construct M training sample pairs by combining the training text sample with each corresponding sample label respectively; the training sample pairs are used to train the second text encoder.

[0189] In one implementation, Embodiment 4 of the application discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor in a computer device, the computer device can execute each step of any embodiment of a two-stage text classification method disclosed in the present invention. The computer-readable storage medium can be non-volatile or volatile.

[0190] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0191] Input the target text and the label description text of the classification label into the first text encoder to obtain the top K alternative labels with the highest similarity to the target text;

[0192] Construct K combined texts by combining the target text with the label description text of each alternative label respectively;

[0193] Input each combined text into the second text encoder to obtain the matching probability between the target text and each alternative label; the second text encoder and the first text encoder are different encoders;

[0194] Select the alternative label with the highest matching probability as the classification category of the target text.

[0195] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0196] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0197] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A two-stage text classification method, characterized in that, The method includes: Inputting the target text and the label description text of the classification label into a first text encoder to obtain K alternative labels with the highest similarity to the target text; Constructing K combined texts by respectively combining the target text with the label description text of each alternative label; Inputting each combined text into a second text encoder to obtain the matching probability between the target text and each alternative label; the second text encoder and the first text encoder are different encoders; Selecting the alternative label with the maximum matching probability as the classification category of the target text.

2. The two-stage text classification method according to claim 1, wherein, The step of inputting the target text and the label description text of the classification label into a first text encoder to obtain K alternative labels with the highest similarity to the target text includes: Inputting the target text and the label description text of the classification label into a first text encoder to obtain the target vector of the target text and the label vectors of each classification label; Calculating the target vector respectively with each label vector to obtain the target similarity between the target vector and each label vector; According to the magnitudes of the target similarities, screening out K classification labels with the highest similarity to the target vector as alternative labels.

3. The two-stage text classification method according to claim 2, wherein The step of calculating the target vector respectively with the label vectors of each classification label to obtain the target similarity between the target vector and the label vectors of each classification label includes: Performing cosine similarity calculation between the target vector and the label vectors of each classification label through the encoder to obtain a first similarity; Performing word frequency similarity calculation between the target text and the label description text of each classification label to obtain a second similarity; Calculating the first similarity and the second similarity to obtain the target similarity.

4. The two-stage text classification method according to claim 1, wherein, The first text encoder is trained in the following manner: Obtaining training text samples and the corresponding positive labels and negative labels for the training text samples; the positive label is the correct classification label of the training text sample, and the negative label is the incorrect classification label of the training text sample; Inputting the training text samples, the label description text corresponding to the positive labels, and the label description text corresponding to the negative labels into a first text encoder to obtain the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels; Calculating the text vectors, the positive label vectors, and the negative label vectors to obtain a triplet loss value; When the triplet loss value is greater than 0, updating the parameters of the first text encoder according to the triplet loss value to continue training the first text encoder; When the triplet loss value converges, obtaining the trained first text encoder.

5. The two-stage text classification method according to claim 4, characterized in that, The step of inputting the training text samples, the label description text corresponding to the positive labels, and the label description text corresponding to the negative labels into a first text encoder to obtain the text vectors of the training text samples, the positive label vectors of the positive labels, and the negative label vectors of the negative labels includes: Input the training text sample into the first encoder branch of the first text encoder to obtain the text vector of the training text sample; Input the label description text corresponding to the positive label and the label description text corresponding to the negative label into the second encoder branch of the first text encoder to obtain the positive label vector and the negative label vector.

6. The two-stage text classification method according to claim 1, characterized in that, The second text encoder can be trained as follows: Input the training text sample and the label description text of the classification label into the first text encoder to obtain N sample labels corresponding to the training text sample; Construct N training sample pairs by respectively combining the training text sample with each corresponding sample label; Input the training sample pairs into the second text encoder for training to obtain the trained second text encoder.

7. The two-stage text classification method according to claim 6, wherein Before inputting the training sample pairs into the second text encoder for training to obtain the trained second text encoder, it further includes: Perform word frequency similarity calculation on the training text sample and the label description text of the classification label to obtain M sample labels corresponding to the training text sample; Construct M training sample pairs by respectively combining the training text sample with each corresponding sample label; the training sample pairs are used to train the second text encoder.

8. A two-stage text classification device, characterized in that, The device includes: An alternative label acquisition unit, configured to input the target text and the label description text of the classification label into the first text encoder to obtain K alternative labels with the highest similarity to the target text; A combined text acquisition unit, configured to construct K combined texts by respectively combining the target text with the label description text of each alternative label; A matching probability acquisition unit, configured to input each combined text into the second text encoder to obtain the matching probability between the target text and each alternative label; the second text encoder and the first text encoder are different encoders; A classification category determination unit, configured to select the alternative label with the maximum matching probability as the classification category of the target text.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the two-stage text classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the two-stage text classification method according to any one of claims 1 to 7.