Text multi-label classification method and system and medium
By constructing a high-quality sample set and vector representation model and combining large language models to process text data, the problem of matching spoken text and professional terms in the professional field is solved, and efficient text multi-label classification is achieved.
Patent Information
- Application Number
- CN202510378524.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing text classification model is difficult to effectively match the labels formed by colloquial text data and professional terms in the professional field, resulting in low classification accuracy and uneven number of labels, affecting the efficiency and accuracy of classification tasks.
By constructing a high-quality sample set, using vector characterization model and classification model, matching text data and labels based on vector similarity, using a large language model for preprocessing and sample expansion, realizing text multi-label classification.
It improves the accuracy of text data classification, adapts to complex label systems, reduces the limit on the number of tags, and improves the efficiency and accuracy of classification tasks.
Smart Images

Figure CN120296598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and in particular, to a method for text multi-label classification, a system for text multi-label classification, and a computer-readable storage medium. Background Art
[0002] In a specific business field, for example, when consumers make complaints and reports, after inputting relevant text data, users often also select tags corresponding to the text data.
[0003] However, the tags available for users to select may include multiple types, the number of tags under each type ranges from several hundred to thousands, and it often involves professional field classification. Therefore, the matching accuracy between the tags selected by users and the input text data is very low. If the input text data is matched based on the tags selected by the users themselves, it will affect the efficiency of subsequent task distribution according to the tags and the accuracy of statistical analysis.
[0004] Existing text classification models can classify the input text after sufficient training. However, in classification tasks involving professional fields, the colloquial text data input by users is usually difficult to match with the categories involving professional field classification. Secondly, due to the very low matching accuracy between the tags selected by users themselves and the input text data, the text classification model cannot obtain enough data with accurate tags for training. Thirdly, the number of tags under each type is very large, while the number of text data corresponding to each tag is unbalanced, greatly reducing the classification accuracy of the text classification model. In addition, the online analysis process of existing text classification models often directly discriminates the categories after inputting the text, so the number of category tags for classification is limited.
[0005] In order to overcome the above-mentioned defects existing in the prior art, there is an urgent need in this field for a method for text multi-label classification that can determine the category of text data based on high-quality samples, solve the matching problem between colloquial text data and tags formed based on professional terms, improve the accuracy of classification tasks, and does not limit the number of classification tags, and is more suitable for the actual business field with a complex tag system. Summary of the Invention
[0006] The following gives a brief overview of one or more aspects to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects, and is neither intended to identify key or decisive elements of all aspects nor to attempt to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0007] To overcome the above-mentioned defects existing in the prior art, the present invention provides a text multi-label classification method, a text multi-label classification system, and a computer-readable storage medium, which can determine the category of text data based on high-quality samples, solve the matching problem between colloquial text data and labels formed based on technical terms, improve the accuracy of the classification task, and do not limit the number of classification labels, being better applicable to the actual business fields with complex label systems.
[0008] Specifically, according to the above-mentioned text multi-label classification method provided by the first aspect of the present invention, the method includes the steps of: converting a text data sample into a first entry including a feature vector, the feature vector being generated by a vector feature model, and the vector feature model being pre-trained through a first sample set; based on the feature vector of the text data sample, determining at least one similar entry of the first entry in a first database according to vector similarity, the first database being constructed by entries transformed from each sample of the first sample set, and the entry including a feature vector and a label category; comparing the first entry with the similar entries through a classification model, and determining the similar entries of the same category as the first entry as the entries of the same category; and counting the frequencies of the label categories of the entries of the same category, and determining the N label categories with the highest frequencies as the category of the text data sample.
[0009] Preferably, in an embodiment of the present invention, the first sample set for training the vector feature model is constructed by quality samples, and the steps for determining the quality samples include: converting each sample of an original data sample set into a second entry including a feature vector and a label category, the feature vector being generated by a vector feature model, and the vector feature model being pre-trained using the original data samples; storing a plurality of the second entries in a second database; and based on the second database, determining at least one selected label for each sample according to vector similarity, and taking the sample whose selected label includes the label category as a quality sample.
[0010] Preferably, in an embodiment of the present invention, the steps for determining the selected label include: calculating the distance between the feature vector of the selected sample and the feature vectors of other second entries in the second database; selecting the M second entries with the closest distances and counting the frequencies of the corresponding label categories; and according to the frequencies of the counted label categories, selecting the K label categories with the highest frequencies as the selected label of the selected sample.
[0011] Preferably, in an embodiment of the present invention, the determination of the quality samples further includes an iterative step: converting the quality samples into second entries including representation vectors and label categories, where the representation vectors are generated by an iterated vector representation model, and the iterated vector representation model is pre-trained with the quality samples; storing a plurality of the second entries in a second database; based on the second database, determining at least one election label for each sample according to vector similarity, and using the sample whose election label includes the label category as the new quality sample; and repeating the above steps until the quality samples no longer change.
[0012] Preferably, in an embodiment of the present invention, the construction of the first sample set further includes a sample adjustment process, and the sample adjustment process includes the steps of: counting the categories of the quality samples to determine the number of samples in each category; screening the samples in the categories with sufficient samples and expanding the samples in the categories with insufficient samples; where the step of sample expansion includes: randomly selecting at least two input samples from the screened samples included in each category with sufficient samples; and the large language model performs sample expansion based on the input samples, the prompt template, and the category of the samples to be generated.
[0013] Preferably, in an embodiment of the present invention, the offline training step of the vector representation model includes: constructing positive and negative sample triplets based on the input sample set; inputting the positive and negative sample triplets into the vector representation model to generate vectors; and fine-tuning the vector representation model based on the generated vectors and the triplet loss function.
[0014] Preferably, in an embodiment of the present invention, the step of determining at least one similar entry of the first entry includes: calculating the distance between the representation vector of the first entry and the representation vectors of each entry in the first database; and selecting the Q entries with the closest distance as the similar entries of the first entry.
[0015] Preferably, in an embodiment of the present invention, it further includes preprocessing of the samples, and the preprocessing steps include: refining keywords and abstracts of the original text of the samples through a large language model; and the generation of the representation vectors includes: performing vector representation on the original text, the keywords, and the abstracts respectively through the vector representation model.
[0016] In addition, the above-mentioned text multi-label classification system provided by the second aspect of the present invention includes a memory and a processor. Computer instructions are stored on the memory. The processor is connected to the memory and is configured to execute the computer instructions stored on the memory to implement the text multi-label classification method provided by any one of the above embodiments.
[0017] In addition, computer instructions are stored on the computer-readable storage medium provided according to the third aspect of the present invention. When the computer instructions are executed by a processor, the text multi-label classification method provided in any one of the above embodiments is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] After reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings, the above features and advantages of the present invention can be better understood. In the drawings, the components are not necessarily drawn to scale, and components with similar related characteristics or features may have the same or similar reference numerals.
[0019] Figure 1 A schematic diagram of a text multi-label classification system provided according to some embodiments of the present invention is shown;
[0020] Figure 2 A flowchart of a text multi-label classification method provided according to some embodiments of the present invention is shown;
[0021] Figure 3 A schematic diagram of the construction of a first sample set provided according to some embodiments of the present invention is shown;
[0022] Figure 4 A schematic diagram of the fine-tuning process of a vector representation model provided according to some embodiments of the present invention is shown; and
[0023] Figure 5 A schematic diagram of a classification model provided according to some embodiments of the present invention is shown.
[0024] REFERENCE NUMERALS:
[0025] 100: Text multi-label classification system;
[0026] 110: Memory;
[0027] 111: Computer-readable storage medium;
[0028] 120: Processor;
[0029] S210-S240: Steps;
[0030] 410: Vector representation model;
[0031] 420: Triplet loss function;
[0032] 510, 520: Text content;
[0033] 511, 521: Original text;
[0034] 512, 522: Abstract;
[0035] 513, 523: Keywords;
[0036] 531, 532: Vector representation models;
[0037] e11, e12, e13: Representation vectors;
[0038] e21, e22, e23: Representation vectors;
[0039] 541, 542: Attention mechanism layers;
[0040] 550: MLP structure; and
[0041] 560: Output result. Detailed implementation manners
[0042] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Note that the aspects described below in conjunction with the accompanying drawings and specific embodiments are merely exemplary and should not be construed as imposing any limitation on the protection scope of the present invention.
[0043] In the description of the present invention, it should be noted that, unless otherwise clearly specified and defined, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] In addition, the "upper", "lower", "left", "right", "top", "bottom", "horizontal", and "vertical" used in the following description should be understood as the orientations shown in this section and the related drawings. This relative term is only for convenience of description and does not represent that the device described needs to be manufactured or operated in a specific orientation, so it should not be construed as a limitation on the present invention.
[0045] It can be understood that although the terms "first", "second", "third", etc. can be used here to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first component, region, layer, and / or part discussed below can be referred to as the second component, region, layer, and / or part without departing from some embodiments of the present invention.
[0046] As described above, existing text classification models can classify the input text after sufficient training. However, in classification tasks related to professional fields, the colloquial text data input by users is usually difficult to match the categories related to professional field classification. Secondly, due to the low matching accuracy between the labels selected by the users themselves and the input text data, the text classification model cannot obtain sufficient data with accurate labels for training. Thirdly, the number of labels under each type is large, while the number of text data corresponding to each label is unbalanced, greatly reducing the classification accuracy of the text classification model.
[0047] To overcome the above-mentioned defects of the prior art, the present invention provides a text multi-label classification method, a text multi-label classification system, and a computer-readable storage medium, which can determine the category of text data based on high-quality samples, solve the matching problem between colloquial text data and labels formed based on professional terms, improve the accuracy of classification tasks, and do not limit the number of classification labels, and are better applicable to the actual business fields with complex label systems.
[0048] In some non-limiting embodiments, the above-mentioned text multi-label classification method provided by the first aspect of the present invention can be implemented via the above-mentioned text multi-label classification system provided by the second aspect of the present invention.
[0049] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of a text multi-label classification system provided according to some embodiments of the present invention.
[0050] As Figure 1 shown, the text multi-label classification system 100 may be configured with a memory 110 and a processor 120. The memory 110 includes but is not limited to the above-mentioned computer-readable storage medium 111 provided by the third aspect of the present invention, on which computer instructions are stored. The processor 120 is connected to the memory 110 and is configured to execute the computer instructions stored on the memory 110 to implement the text multi-label classification method provided by the first aspect of the present invention.
[0051] The working principle of the above text multi-label classification system will be described below in conjunction with some embodiments of text multi-label classification methods. Those skilled in the art can understand that these embodiments of text multi-label classification methods are only some non-limiting implementation manners provided by the present invention, aiming to clearly show the main concept of the present invention and provide some specific solutions for the public to implement, rather than limiting all functions or all working manners of the text multi-label classification system. Similarly, the text multi-label classification system is also a non-limiting implementation manner provided by the present invention, and does not limit the execution subject and execution order of each step in these text multi-label classification methods.
[0052] Please refer to Figure 2 , Figure 2 which shows a flowchart of a text multi-label classification method provided according to some embodiments of the present invention.
[0053] As Figure 2 shown, the text multi-label classification method may include step S210: converting a text data sample into a first entry including a feature vector, the feature vector being generated by a vector feature model, and the vector feature model being pre-trained through a first sample set. Based on this vector feature model, the text multi-label classification system can capture and extract fine-grained semantic features of natural language text through deep learning techniques.
[0054] The vector feature model is pre-trained through a first sample set. The text multi-label classification system may first determine quality samples to construct a high-quality first sample set.
[0055] Please refer to Figure 3 , Figure 3 which shows a schematic diagram of the construction of a first sample set provided according to some embodiments of the present invention.
[0056] As Figure 3 shown, the text multi-label classification system may first convert each sample in the original data sample set into a second entry including a feature vector and a label category. The feature vector is generated by a vector feature model, where the vector feature model is pre-trained using the original data samples.
[0057] In some embodiments, the offline pre-training step of the vector feature model may first construct positive and negative sample triples based on the input samples.
[0058] The text multi-label classification system may construct a first positive and negative sample triple according to the input sample and the label category corresponding to the input sample selected by the user:
[0059] (data, category, other category),
[0060] Among them, the data is the text content of the input sample, the category is the label category corresponding to the input sample, the other category is another label category other than the label category corresponding to the input sample, the category is the positive sample, and the other category is the negative sample.
[0061] Similarly, the text multi-label classification system can also construct a second positive and negative sample triple:
[0062] (Data, same-category data, other-category data),
[0063] Among them, the data is the text content of the input sample, the same-category data is the text content of another sample under the label category corresponding to the input sample in the original data, the other-category data is the text content of another sample under other label categories other than the label category corresponding to the input sample in the original data, the same-category data is the positive sample, and the other-category data is the negative sample.
[0064] In some embodiments, the text multi-label classification system can control the proportion between various categories when constructing the triple dataset. For example, when constructing the positive and negative sample triple, the mode or median can be selected as the number of data in the positive and negative sample triple under each category according to the statistical distribution of the number of samples under each category in the original data. For example, 10 pieces of data can be selected under a certain category to construct the positive and negative sample triple.
[0065] In the selection of negative samples, if the label of the data is a hierarchical label category, then, in the positive and negative sample triple dataset constructed based on the selected data, the negative samples of half of the positive and negative sample triples are randomly selected from all categories other than the label category corresponding to the text data sample, and the negative samples of the other half of the positive and negative sample triples are randomly selected under a certain major category other than the label category corresponding to the text data sample.
[0066] For example, the constructed positive and negative sample triplets are the first positive and negative sample triplets. The text multi-label classification system can determine the median to be 100 based on the statistical distribution of the number of samples in each category of the original data. Then, 100 samples are randomly selected in each category to construct the positive and negative sample triplets. If the number of samples in a certain label category is more than 100, the extra samples are removed; if the number of samples in a certain label category is less than 100, all the samples are used to construct the positive and negative sample triplets. Specifically, the text multi-label classification system can randomly select a sample in the "product quality" category as the data of the positive and negative sample triplets. The text content of this sample can be: "There is a problem with the quality of XX product. It broke just three days after purchase", and the category is "product quality". Based on the selected sample, the text multi-label classification system can select negative samples to construct multiple positive and negative sample triplets. For example, the negative samples can be randomly selected from other label categories except the "product quality" category; or, the negative samples can also be randomly selected from sub-categories including "overbearing clauses", "contract breach", or "contract fraud" under the large category of "contract disputes". Similarly, the constructed positive and negative sample triplets are the second positive and negative sample triplets. The text content of the above sample is used as the data of the positive and negative sample triplets. After that, samples can be selected in the "product quality" category as the same-category data. The selected samples as the same-category data are not limited to the aforementioned randomly selected 100 samples, but are selected from all samples in the "product quality" category.
[0067] After that, the above-mentioned first positive and negative sample triplets and the second positive and negative sample triplets are proportioned according to a preset ratio to form a triplet dataset, and then each positive and negative sample triplet in the triplet dataset is input into the vector representation model to generate vectors. In some embodiments, the preset ratio of the first positive and negative sample triplets and the second positive and negative sample triplets can be 1:1.
[0068] Please refer to Figure 4 , Figure 4 which shows a schematic diagram of the fine-tuning process of the vector representation model provided according to some embodiments of the present invention.
[0069] As Figure 4 shown, the vector representation model 410 can be a pre-trained model of the Encoder (encoder) class, such as BERT (Bidirectional Encoder Representation from Transformers) or RoBERTa (an improved version of BERT).
[0070] The text multi-label classification system can fine-tune the vector representation model through a contrastive learning-based vector representation algorithm. Specifically, as Figure 4As shown, c1, c2, and c3 are positive and negative sample triples input into the vector representation model 410. In the above embodiments, when the positive and negative sample triples are the first positive and negative sample triples, c1 can be data, c2 can be a category, and c3 can be another category. Correspondingly, when the positive and negative sample triples are the second positive and negative sample triples, c1 can be data, c2 can be data of the same category, and c3 can be data of another category.
[0071] Based on the input c1, c2, and c3, the vector representation model 410 generates vectors v1, v2, and v3. Based on the generated vectors v1, v2, and v3, the vector representation model is fine-tuned based on the triplet loss function 420 so that the distance between the data and the positive samples is closer and the distance between the data and the negative samples is farther.
[0072] Here, the specific formula of the triplet loss function 420 is as follows:
[0073]
[0074] Among them, ∈ is a hyperparameter, which is a constant greater than 0 and can be set and adjusted during fine-tuning.
[0075] Please continue to refer to Figure 3 , through the above pre-trained vector representation model, the text multi-label classification system can convert each sample of the original data sample set into a second entry including a representation vector and a label category. In some embodiments, the text multi-label classification system can use the representation vector as the key, and the text content of each sample in the original data sample set and the label category selected by the user as the value to form a second entry. Then, multiple second entries are stored in the second database.
[0076] For example, the text multi-label classification system can select a sample, the text content of which can be: "The quality of XX commodity is problematic. It broke within three days of purchase", and the label category is "Commodity quality". The text content of this sample is converted into a representation vector key_1 through the pre-trained vector representation model, and the text content and the label category are used as value_1. Thus, this sample is converted into a second entry including the representation vector key_1 and the value value_1.
[0077] In some preferred embodiments, the second database can also include category entries. The text multi-label classification system can also perform vector representation on the label category through the vector representation model to obtain a representation vector key, and then use a null value and the label category as the value to form a category entry to be stored in the second database, so as to further consider the semantics of the label category itself during the text classification task.
[0078] For example, a text multi-label classification system may select the "product quality" category from the label categories available for user selection. Then, taking this "product quality" as the text content, it is transformed into a representation vector key_2 through a pre-trained vector representation model, and the label category "product quality" is taken as value_2. Thus, this label category is transformed into a category entry including the representation vector key_2 and the value value_2.
[0079] As Figure 3 shown, the text multi-label classification system can perform a matching election to determine quality samples. Specifically, the text multi-label classification system can implement the matching election based on the second database by determining at least one election label for each sample according to the vector similarity.
[0080] In some embodiments, the text multi-label classification system can first calculate the distance between the representation vector of the selected sample and the representation vectors of other second entries in the second database to determine the vector similarity between the representation vectors. Then, select the M second entries with the closest distance and count the frequencies of the corresponding label categories. Then, according to the frequencies of the label categories of the M second entries counted, select the K label categories with the highest frequencies as the election labels of the selected sample.
[0081] Taking the above sample as an example, the text data content of the selected sample is: "There is a problem with the quality of XX product. It broke just three days after purchase", and the label category is "product quality". The second entry transformed from this selected sample includes the representation vector key_1 and the value value_1. The text multi-label classification system can first calculate the distance between the representation vector key_1 of the selected sample and the representation vectors of other second entries in the second database. Then, select the 10 second entries with the closest distance and count the frequencies of the corresponding label categories. The label categories corresponding to the 10 second entries with the closest distance can be "food quality", "product quality", "food quality", "product quality", "product quality", "product transportation", "product quality", "product transportation", "product transportation", and "product quality". Among them, "product quality" appears 5 times, "product transportation" appears 3 times, and "food quality" appears 2 times. According to the frequencies of the label categories counted, select the 2 label categories with the highest frequencies as the election labels of the selected sample, that is, select "product quality" and "product transportation" as the election labels of this selected sample.
[0082] After that, the text multi-label classification system can, according to the election label results obtained from the matching election, take the samples with election labels including label categories as quality samples, that is, select the data whose election label results are consistent with the label categories in the original data as confident data.
[0083] For example, in the above example, since the selected sample's election labels "product quality" and "product transportation" include the label category "product quality" of the selected sample, the text multi-label classification system can use the selected sample as a quality sample for constructing the first sample set.
[0084] For another example, the text data content of the selected sample may be: "A certain restaurant has a bad attitude, verbally abuses customers, and insults them", and the label category is "product quality". The text data content of this sample is converted into a representation vector key_3 through a pre-trained vector representation model, and the text data content and the label category are used as value_3. Thus, this sample is converted into the second entry including the representation vector key_3 and the value value_3. Then, the text multi-label classification system can select the 10 second entries closest to the representation vector key_3 and count the frequencies of the corresponding label categories. The label categories corresponding to the 10 closest second entries can be "catering service", "catering service", "catering service", "catering service", "catering service", "food quality", "food quality", "food quality", "catering service", and "catering service". Among them, "catering service" appears 7 times and "food quality" appears 3 times. According to the counted frequencies of the label categories, select the 2 label categories with the highest frequencies as the election labels of the selected sample, that is, select "catering service" and "food quality" as the election labels of the selected sample. In this example, since the election labels "catering service" and "food quality" of the selected sample do not include the label category "product quality" of the selected sample, the selected sample will not be used as a quality sample to construct the first sample set.
[0085] In this way, by selecting the data whose election label results are consistent with the label categories in the original data as the confidence data, the text multi-label classification system can solve the matching problem between the colloquial text data and the label categories formed by professional terms.
[0086] Due to the differences between colloquial language and professional terms, the text multi-label classification system can also preprocess the input text data samples to screen out the irrelevant information contained in the text. In some preferred embodiments, the text multi-label classification system can extract the keywords, abstracts, or key information of the original text of the input sample through a large language model, and the extracted keywords, abstracts, or key information can be used alone or in combination.
[0087] For example, the original text may include the content: "On a certain date, for a certain purpose, XX products were purchased on a certain website. XX products are clearly fragile, but the manufacturer did not properly package them. After being transported by express delivery, the express packaging was significantly damaged. After opening, no defects were seen, but the product was damaged three days after use." The text multi-label classification system can extract the keywords of the original text through a large language model: "a certain website", "XX products", "express delivery", "packaging damage", "no defects", etc., and extract the abstract: "On a certain date, XX fragile products were purchased on a certain website. The manufacturer did not properly package the products. After being transported by express delivery, the express packaging was significantly damaged. No defects were seen when the product was opened. However, the product was damaged three days after use."
[0088] After that, based on the keywords and abstracts obtained after preprocessing, through the vector representation model, vector representations are made respectively based on the original text, keywords, and abstracts. Then, the text multi-label classification system can store the representation vectors converted from the original text, keywords, and abstracts respectively as the representation vectors for the second purpose, or store the representation vectors converted from the original text, keywords, and abstracts after concatenation as the representation vectors for the second purpose.
[0089] More preferably, the determination of quality samples can also include an iterative step. As Figure 3 shown, after determining the quality samples, the text multi-label classification system can use the quality samples to pre-train the vector representation model to iterate the vector representation model. The quality samples are converted into a second entry including a representation vector and a label category through the iterated vector representation model, and the converted second entry is stored in the second database. The second database can clear the second entries and category entries formed by converting each sample of the original data set before, and only store the second entries of the quality samples and the category entries of the label categories formed by the iterated vector representation model pre-trained by the quality samples.
[0090] Then, based on the new second database, the text multi-label classification system can determine at least one elected label for each sample according to the vector similarity, and use the samples with the elected label including the label category as new quality samples. Then, the new quality samples are compared with the quality samples in the previous iteration round to determine whether the quality samples converge. If the quality samples change, it means that the quality samples do not converge, and the text multi-label classification system can continue to iterate. If the quality samples no longer change, it means that the quality samples converge, and the text multi-label classification system stops iterating.
[0091] Please continue to refer to Figure 3 , the construction of the first sample set can also include a sample adjustment process to keep the number of samples in each category balanced. The sample adjustment process includes sample screening and sample expansion.
[0092] The text multi-label classification system can first count the categories of quality samples to determine the number of samples under each category. In some embodiments, when the number of samples under a category exceeds a preset value, it can be determined that the category is a category with sufficient samples. When the number of samples under a category is lower than the preset value, it can be determined that the category is a category with insufficient samples.
[0093] Then, the text multi-label classification system can screen the samples of the categories with sufficient samples and expand the samples of the categories with insufficient samples.
[0094] Sample screening can be achieved by clustering or random selection. In some embodiments, the text multi-label classification system can vectorize the samples under the categories with sufficient samples. Then, the formed representation vectors are clustered into n categories, where n is the number of samples to be screened. Then, select the samples at the n clustering centers and remove the remaining samples. Alternatively, in some other embodiments, the text multi-label classification system can randomly select n samples under the categories with sufficient samples.
[0095] Sample expansion can be achieved by the large language model based on retrieval enhancement and few-shot mechanism. In some embodiments, the text multi-label classification system can randomly select at least two input samples from the samples under the categories with sufficient samples after the above screening. Then, through the large language model, based on the input samples, the prompt template, and the category of the samples to be generated, sample expansion is performed.
[0096] In some embodiments, the prompt template of the large language model can be set as follows:
[0097] "<Category>: {label1}<Text>: {text1}<Category>: {label2}<Text>: {text2}<Category>: {label3}<Text>:"
[0098] Among them, label1 / text1 and label2 / text2 are the label categories and text data contents of two randomly selected input samples; label3 is the label category for which sample expansion is required.
[0099] For example, a text multi-label classification system can randomly select at least two input samples from the samples under the categories with sufficient filtered samples mentioned above. One of the input samples can have a label category of "product quality", and the text data content can be: "There is a problem with the quality of XX product. It broke just three days after purchase."; Another input sample can have a label category of "contract breach", and the text data content can be: "XX manufacturer refuses to pay after receiving the goods. The amount of the payment is very high, and subsequent production cannot continue." The label category that needs to be sample-expanded can be "product defect". Thus, the prompt words input into the large language model are formed as follows:
[0100] "<category>: {product quality}<text>: {There is a problem with the quality of XX product. It broke just three days after purchase}<category>: {contract breach}<text>: {XX manufacturer refuses to pay after receiving the goods. The amount of the payment is very high, and subsequent production cannot continue.}<category>: {product defect}<text>:"
[0101] The large language model can form text data content for the "product defect" category according to the above prompt word template. <text3>。For example, forming text data content <text3>: "The purchased hand warmer exploded during charging, causing huge losses."
[0102] In this way, through the above sample adjustment process including sample screening and sample augmentation, samples under categories with sufficient samples are screened, and samples under categories with insufficient samples are augmented, so that the number of samples under each category is balanced.
[0103] Text classification models based on machine learning or deep learning often rely on high-quality training data. In specific business fields, the input text data is usually low-quality data that is partial to colloquial language and has low label accuracy. In the prior art, those skilled in the art obtain high-quality samples by manually annotating texts, which is time-consuming, laborious, inefficient, and inaccurate. The multi-label text classification method provided by the present invention constructs the first sample set through the construction process as Figure 3 shown, obtains multiple versions of gradually optimized vector representation models through iteration, produces election label results according to the vector representation models, compares the election label results with the label categories of the samples to determine quality samples, and then obtains the finally converged quality sample results through iterative screening to solve the sample quality problem, thereby determining high-quality samples, so that the multi-label text classification method can perform model training and iteration without relying on manually annotated texts, effectively improving efficiency and accuracy.
[0104] Thus, the multi-label text classification system can construct the first sample set through the construction process as Figure 3 shown, improving the accuracy of the multi-label text classification method.
[0105] Then, the multi-label text classification system can perform pre-training of the vector representation model based on the first sample set to determine the final version of the vector representation model. Based on this vector representation model, the input text data sample is converted into the first entry including the representation vector.
[0106] Specifically, the pre-training of the vector representation model can be as Figure 4 shown. The multi-label text classification system can first construct positive and negative sample triples according to the text content and label categories of the input quality samples. The constructed positive and negative sample triples can be the first positive and negative sample triples (data, category, other category) or the second positive and negative sample triples (data, same-category data, other-category data). The first positive and negative sample triples and the second positive and negative sample triples can form a triple data set according to a preset ratio. Here, the preset ratio can be selected as the optimal ratio according to the effect of pre-training. Then, the positive and negative sample triples in the triple data set are input into the vector representation model to generate vectors, and the vector representation model is fine-tuned through the triplet loss function.
[0107] Please continue to refer to Figure 2 , the multi-label text classification method may include step S220: based on the representation vectors of text data samples, determine at least one similar entry of the first entry in the first database according to vector similarity. The first database is constructed by entries transformed from each sample of the first sample set, and the entry includes a representation vector and a label category.
[0108] The multi-label text classification system may transform each quality sample in the above first sample set into an entry including a representation vector and a label category through a vector representation model. Store each entry transformed from each quality sample in the first sample set into the first database to construct the first database reflecting the mapping from text content to label categories.
[0109] Preferably, the first database may also include category entries. The category entries may be determined according to the vector representation of the label category by the vector representation model. Specifically, the multi-label text classification system may perform vector representation on the label category through the vector representation model to obtain the representation vector key, and then use null value and the label category as the value to form category entries to be stored in the first database, so as to further consider the semantics of the label category itself in the text classification task process.
[0110] More preferably, due to the differences between spoken language and professional terms, the multi-label text classification system may also preprocess the quality samples, such as refining the keywords, abstracts or key information of the original text of the input samples through a large language model to screen out the irrelevant information contained in the text. Then, through the vector representation model, perform vector representation based on the original text, keywords and abstracts respectively and store them as the representation vectors of the samples in the first database.
[0111] After that, the multi-label text classification system may calculate the distance between the representation vector of the first entry transformed from the input text data sample and the representation vectors of each entry in the first database to determine the vector similarity between the representation vector of the first entry and the representation vectors of each entry. Then, select the Q entries with the closest distance as the similar entries of the first entry.
[0112] As Figure 2 shown, the multi-label text classification method may include step S230: compare the first entry and the similar entries through a classification model, and determine the similar entries of the same category as the first entry as the entries of the same category.
[0113] Since the expression of the label category is relatively professional while the expression of the text content is relatively spoken, therefore, the classification model provided by the present invention does not directly judge the relationship between the text content and the label category, but judges whether the text content of the text and the text content of the similar entries belong to the same category. In this way, the multi-label text classification system can improve the classification accuracy of the multi-label text classification method by further predicting whether the text and its similar text belong to the same category.
[0114] The text multi-label classification system can input the first entry and similar entries into the classification model respectively, and based on the classification model, determine whether the text corresponding to the first entry and the text corresponding to the similar entries belong to the same category. The classification model can be a two-tower model, and the classification model can include an attention layer and an MLP (Multilayer Perceptron) structure.
[0115] Please refer to Figure 5 , Figure 5 which shows a schematic diagram of the classification model provided according to some embodiments of the present invention.
[0116] As Figure 5 shown, the text multi-label classification system can input the text content 510 corresponding to the first entry and the text content 520 corresponding to a similar entry into the classification model. In Figure 5 the embodiment shown, the text multi-label classification system can obtain the keywords and summaries of the text content 510 and the text content 520 through a large language model. Then, the original text 511, summary 512, and keywords 513 corresponding to the text content 510 are input into the vector representation model 531 to generate the representation vector e11 corresponding to the original text 511, the representation vector e12 corresponding to the summary 512, and the representation vector e13 corresponding to the keywords 513; and the original text 521, summary 522, and keywords 523 corresponding to the content 520 are input into the vector representation model 532 to generate the representation vector e21 corresponding to the original text 521, the representation vector e22 corresponding to the summary 522, and the representation vector e23 corresponding to the keywords 523. The vector representation model 531 and the vector representation model 532 can be vector representation models pre-trained through a first sample set.
[0117] The representation vectors e11, e12, and e13 are input into the attention mechanism layer 541 for fusion and then input into the MLP structure 550, and the representation vectors e21, e22, and e23 are input into the attention mechanism layer 542 for fusion and then input into the MLP structure 550.
[0118] In an example, the process of the attention mechanism layer 541 and the attention mechanism layer 542 fusing the input representation vectors can include concatenating the input representation vectors and inputting them into a fully connected layer, and then outputting the weights of each representation vector according to the softmax function, and outputting the fused representation vector through weighted summation.
[0119] The text multi-label classification system can discriminate the text content 510 corresponding to the first entry and the text content 520 corresponding to the similar entries of this entry according to the MLP structure 550, and obtain the output result 560. In response to the text content 510 corresponding to the first entry and the text content 520 corresponding to the similar entries of this entry being of the same category, the output result 560 can be 1; in response to the text content 510 corresponding to the first entry and the text content 520 corresponding to the similar entries of this entry being of different categories, the output result 560 can be 0. The similar entries with the output result 560 being 1 are determined as the similar entries of the same category as the first entry.
[0120] After that, please continue to refer to Figure 2 , the text multi-label classification method may include step S240: counting the frequencies of the label categories of the similar entries, and determining the top N label categories with the highest frequencies as the categories of the text data samples.
[0121] In one embodiment, the text content of the text data sample input to the text multi-label classification system may be: "The waiter was not hygienic during the cooking process, did not wear a mask or gloves, directly touched the food with hands, and there was also a foreign object in a certain dish served". The text multi-label classification system converts the text data sample into a first entry including a feature vector, and the feature vector is converted from the text content. Then, according to the first database constructed based on the first sample set, Q similar entries of the first entry are determined. Then, one similar entry among the Q similar entries and the first entry are sequentially input into the classification model to determine the similar entries of the first entry. In this embodiment, the similar entries may be 5, corresponding to the label categories: "Food Quality", "Food Service", "Food Quality", "Food Service", and "Food Service" respectively. Among them, the frequency of the label category "Food Service" is 3 times, and the frequency of the label category "Food Quality" is 2 times. The text multi-label classification system can determine the top 1 label category with the highest frequency as the category of this text data sample, that is, "Food Service" can be output as the category of this text data sample.
[0122] The text multi-label classification method provided by the present invention uses text semantic matching to achieve content classification, and does not limit the number of label categories used for classification. Further, in some embodiments, the label categories are hierarchical labels. The text multi-label classification method provided by the present invention can directly and accurately match to the smallest category through text semantic matching, and can adapt to the dynamic changes of the label system.
[0123] In summary, the text multi-label classification method provided by the present invention can perform pre-training of the vector representation model based on a high-quality first sample set, so as to achieve an accurate vector representation of the text data sample input to the text multi-label classification system. Combining with the first database constructed by the high-quality first sample set, the classification model is used to determine the category of the text data, effectively addressing the problem that it is difficult to match the label categories formed by low-quality text content and professional terms, and improving the accuracy of text data classification. Moreover, the text multi-label classification method provided by the present invention does not limit the number of classification labels, and is better applicable to the actual business fields with complex label systems.
[0124] Although the above methods are illustrated and described as a series of actions for simplicity of explanation, it should be understood and appreciated that these methods are not limited by the order of the actions, because according to one or more embodiments, some actions may occur in a different order and / or concurrently with other actions that are illustrated and described herein or that are not illustrated and described herein but are understandable to those skilled in the art.
[0125] Those skilled in the art will understand that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0126] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality in a generalized form. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in different ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present invention.
[0127] The various illustrative logical modules and circuits described in connection with the embodiments disclosed herein can be implemented or executed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0128] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0129] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for multi-label classification of texts, characterized in that, Including steps: Convert the text data sample into a first entry including a characterization vector, which is generated by a vector characterization model pre-trained through a first sample set; Based on the characterization vector of the text data sample, determine at least one similar entry of the first entry in the first database according to vector similarity. The first database is constructed by entries transformed from each sample of the first sample set, and the entry includes a characterization vector and a label category; Compare the first entry and the similar entries through a classification model, and determine the similar entries of the same category as the first entry as the same-category entries; and Count the frequencies of the label categories of the same-category entries, and determine the top N label categories with the highest frequencies as the categories of the text data sample.
2. The text multi-label classification method according to claim 1, characterized in that The first sample set for training the vector characterization model is constructed from quality samples. The steps for determining the quality samples include: Convert each sample of the original data sample set into a second entry including a characterization vector and a label category, which is generated by a vector characterization model pre-trained using the original data sample; Store multiple of the second entries in a second database; and Based on the second database, determine at least one election label for each sample according to vector similarity, and use the sample whose election label includes the label category as a quality sample.
3. The text multi-label classification method according to claim 2, characterized in that, The steps for determining the election label include: Calculate the distance between the characterization vector of the selected sample and the characterization vectors of other second entries in the second database; Select the M second entries with the closest distances and count the frequencies of the corresponding label categories; and According to the counted frequencies of the label categories, select the top K label categories with the highest frequencies as the election label of the selected sample.
4. The text multi-label classification method according to claim 2, wherein The determination of the quality samples further includes an iterative step: Convert the quality sample into a second entry including a characterization vector and a label category, which is generated by an iterated vector characterization model pre-trained through the quality sample; Store multiple of the second entries in a second database; Based on the second database, determine at least one election label for each sample according to vector similarity, and use the sample whose election label includes the label category as the new quality sample; and Repeat the above steps until the quality samples no longer change.
5. The text multi-label classification method according to claim 2, wherein The construction of the first sample set further includes a sample adjustment process, and the sample adjustment process includes steps: Statistically analyze the categories of the quality samples to determine the number of samples in each category; Screen the samples in the categories with sufficient samples and expand the samples in the categories with insufficient samples; Among them, the steps for sample expansion include: Randomly select at least two input samples from the screened samples included in each category with sufficient samples; and The large language model performs sample expansion based on the input samples, the prompt template, and the category of the sample to be generated.
6. The text multi-label classification method according to claim 1, wherein, The offline training steps of the vector characterization model include: Construct positive and negative sample triples based on the input sample set; Input the positive and negative sample triplets into the vector representation model to generate vectors; and Based on the generated vectors, fine-tune the vector representation model based on the triplet loss function.
7. The text multi-label classification method according to claim 1, characterized in that, The steps of determining at least one similar entry of the first entry include: Calculate the distances between the representation vector of the first entry and the representation vectors of each entry in the first database; and Select the Q entries with the closest distances as the similar entries of the first entry.
8. The text multi-label classification method according to claim 1, wherein It also includes preprocessing of the samples, and the preprocessing steps include: refining the keywords and summaries of the original text of the samples through a large language model; and The generation of the representation vectors includes: performing vector representation respectively based on the original text, the keywords, and the summaries through the vector representation model.
9. A text multi-label classification system, characterized in that, Includes: A memory storing computer instructions thereon; And A processor connected to the memory and configured to execute the computer instructions stored on the memory to implement the text multi-label classification method according to any one of claims 1 to 8.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instructions are executed by the processor, the text multi-label classification method according to any one of claims 1 to 8 is implemented.