A Multi-Label Classification Data Augmentation Method and System Based on Large Language Models
By building a dual-weighted label relationship network and tail-driven sampling, creative label text is generated, and the long-tail distribution problem in multi-label text classification is solved, and the classification accuracy and generalization ability of large language models are improved.
Patent Information
- Application Number
- CN202410936689.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-12
AI Technical Summary
The existing technology is difficult to effectively solve the long-tail distribution problem of data sets in multi-label text classification, resulting in a decrease in the classification accuracy and generalization ability of large language models.
By building a dual-weighted label relationship network, tail-driven sampling is performed, creative label text is generated, and large language models are used to optimize label matching and style consistency, increase the number of instances of rare labels, and generate multi-label classification enhancement data.
It effectively solves the problem of long-tail distribution, improves the generalization ability of large language models, and ensures that the generated text remains consistent and relevant with the original data.
Smart Images

Figure CN119577435B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a multi-label classification data augmentation method and system based on a large language model. Background Art
[0002] In the field of natural language processing, multi-label text classification is a crucial task that requires assigning multiple labels to a single text simultaneously. For example, in news classification, an article may be classified into multiple categories such as "Politics", "Economy", and "International" at the same time. The challenge of multi-label text classification lies in dealing with a large and diverse label system, which is of great significance for effectively understanding and organizing a large amount of text data.
[0003] One of the main problems faced by multi-label text classification is the long-tail distribution problem of the dataset. In this distribution, most labels are only associated with a small number of texts, resulting in serious data imbalance. This imbalance has a significant impact on the learning effect of the large language model, making it difficult for the large language model to effectively learn the features of a small number of labels from limited samples, thereby reducing the classification accuracy and generalization ability. Traditional solutions, including text rewriting or over-(under-)sampling, have achieved certain results in some simple text classification tasks, but are ineffective in dealing with multi-label text classification tasks. These methods often simply copy or slightly modify existing data points and cannot generate new label combinations or texts, resulting in insufficient data diversity. This limitation makes traditional methods unable to fundamentally solve the long-tail distribution problem of the data, thereby restricting the improvement of model performance. Therefore, how to improve the generalization ability of the large language model in practical applications and solve the long-tail distribution problem of the dataset during multi-label text classification is an urgent problem to be solved at present. Summary of the Invention
[0004] In view of the deficiencies of existing methods and the requirements of practical applications, in order to improve the generalization ability of large language models in practical applications and solve the long-tail distribution problem of datasets in multi-label text classification, on the one hand, the present invention provides a multi-label classification data augmentation method based on a large language model, including the following steps: obtaining an original text dataset, and obtaining a dual-weighted label relationship network according to the original text dataset; performing tail-driven sampling in the dual-weighted label relationship network, and generating creative labels for the original text dataset according to the sampling results; using the large language model and the creative labels to generate creative label texts; and merging the creative label texts to obtain multi-label classification augmented data. The present invention uses existing labels to construct a label relationship network and then performs tail-driven sampling, designs innovative label combinations considering label matching and style consistency, increases the number of instances of rare labels, effectively addresses the long-tail distribution problem while maintaining consistency and relevance with the original data, and improves the generalization ability of large language models in practical applications.
[0005] Optionally, the obtaining of the dual-weighted label relationship network according to the original text dataset satisfies the following formula: G=(V, E, W v , W e ), where G represents the dual-weighted label relationship network, V represents the vertices represented by the labels of the original text dataset, E represents the edges connecting vertex pairs representing the co-occurrence situation of the labels in the dataset, W v represents the weight aggregation of the vertices, W v ={w v (i)|i∈V}, w v (i) represents the occurrence frequency of label i, W e represents the weight aggregation of the edges, W e ={w e (i, j)|i, j∈V}, w e (i, j) represents the co-occurrence intensity between label i and label j. The dual-weighted label relationship network constructed by the present invention according to the original text dataset provides a comprehensive perspective for the characteristics of each label and their interconnections, is conducive to clearly describing label dynamics so as to introduce deep semantic changes, and is further conducive to capturing and expressing the complex relationships between labels.
[0006] Optionally, the tail-driven sampling in the double-weighted label relationship network includes the following steps: constructing a tail-driven sampling label transition probability model, obtaining the transition probability of labels in the double-weighted label relationship network according to the tail-driven sampling label transition probability model; constructing a tail-driven sampling label important scarcity feature model, obtaining the important scarcity feature value of labels in the double-weighted label relationship network according to the tail-driven sampling label important scarcity feature model; combining the transition probability and the important scarcity feature value, evaluating the label transition acceptance degree in the double-weighted label relationship network, and completing the tail-driven sampling according to the evaluation result. The present invention performs tail sampling in the double-weighted label relationship network, realizes targeted label sampling by focusing on tail labels and maintaining the evaluation of label features, and thus effectively performs data augmentation to alleviate the long-tail effect.
[0007] Optionally, the tail-driven sampling label transition probability model satisfies the following formula: where q(i→j) represents the transition probability from label i to label j in the double-weighted label relationship network, w e (i,j) represents the co-occurrence intensity between label i and label j, and neighbors(i) represents the neighbors that can transfer with label i. By calculating the label transition probability, the present invention can flexibly adjust to the target distribution, further meeting the goal of implementing a balanced sampling strategy in the long-tail distribution.
[0008] Optionally, the tail-driven sampling label important scarcity feature model satisfies the following formula: where p(i) represents the important scarcity feature value of label i, S(i) represents the important scarcity score of label i, and neighbors(i) represents the neighbors that can transfer with label i; S(i)=-log(w v (i)), where S(i) represents the important scarcity score of label i, and w v (i) represents the occurrence frequency of label i, and T represents the temperature parameter for adjusting the distribution smoothness. The important scarcity feature value provided by the present invention can effectively reflect the target distribution of labels, reduce the uncertainty in the sampling process, and improve the effectiveness of sampling.
[0009] Optionally, the combination of the transition probability and the important scarcity feature value to evaluate the label transition acceptance degree in the double-weighted label relationship network satisfies the following formula: Among them, α(i→j) represents the label transfer acceptance degree in the double-weighted label relationship network, min(·) represents the minimum value function, p(·) represents the important scarcity feature value of the label, q(i→j) represents the transfer probability from label i to label j in the double-weighted label relationship network, and q(j→i) represents the transfer probability from label j to label i in the double-weighted label relationship network. The present invention guides the sampling result to develop towards the target distribution by evaluating the label transfer acceptance degree, which is further beneficial to solving the long-tail distribution problem.
[0010] Optionally, generating the creative label text by using the large language model and the creative label satisfies the following formula: Among them, represents the large language model target loss function, represents the label matching loss, represents the style consistency loss, and λ represents the large language model hyperparameter. By balancing the importance of label matching and style consistency, the present invention ensures that the generated text is not only relevant to the sampled labels but also consistent in style with the actual text, achieving a balance between accuracy and authenticity.
[0011] Optionally, the label matching loss satisfies the following formula: Among them, represents the label matching loss, n represents the number of texts, and r φ (X i ) represents the possibility that the large language model generates text X i of, represents the temperature coefficient for the i-th label to generate the i-th text, represents the temperature coefficient for the i-th label to generate the j-th text. By standardizing the label matching loss with a specific formula, the present invention ensures that the generated text not only accurately reflects each label in the sampled label combination but also maintains consistency and relevance with the original dataset.
[0012] Optionally, the style consistency loss satisfies the following formula: Among them, represents the style consistency loss, t represents the number of tokens of the text, x t represents the series of tokens of the text, c(Y) represents the composite text composed of the multi-label set associated with the text and the prompt words, and P φ (x t |c(Y)) represents the token x of the text generated by the large language model tThe probability of conforming to the composite text c(Y). Through a specific formula to standardize the style consistency loss, the present invention can help generate text with a style that is coordinated, controllable, and coherent with the original dataset, further ensuring the balance between the accuracy and authenticity of the generated text.
[0013] In a second aspect, in order to efficiently execute a multi-label classification data augmentation method based on a large language model provided by the present invention, the present invention also provides a multi-label classification data augmentation system based on a large language model, including a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute a multi-label classification data augmentation method based on a large language model as described in the first aspect of the present invention. The multi-label classification data augmentation system based on a large language model of the present invention has a compact structure and stable performance, and can stably execute a multi-label classification data augmentation method based on a large language model provided by the present invention, further improving the overall applicability and practical application ability of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flowchart of a multi-label classification data augmentation method based on a large language model of the present invention;
[0015] Figure 2 It is a framework diagram of a multi-label classification data augmentation system based on a large language model of the present invention;
[0016] Figure 3 It is a schematic structural diagram of a multi-label classification data augmentation device provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described here are only for illustrative purposes and do not limit the present invention. In the following description, in order to provide a thorough understanding of the present invention, a large number of specific details are set forth. However, it will be apparent to those of ordinary skill in the art that the present invention does not have to employ these specific details. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the present invention.
[0018] Throughout the specification, references to "one embodiment", "an embodiment", "an example", or "an illustration" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment", "in an embodiment", "an example", or "an illustration" appearing throughout the specification do not necessarily all refer to the same embodiment or example. Additionally, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Further, those of ordinary skill in the art will understand that the diagrams provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0019] In one embodiment, refer to Figure 1 , Figure 1 , which is a flowchart of a multi-label classification data augmentation method based on a large language model according to the present invention. In order to improve the generalization ability of the large language model in practical applications and solve the long-tail distribution problem of the dataset in multi-label text classification, the present invention provides a multi-label classification data augmentation method based on a large language model, as Figure 1 shown, the method includes the following steps:
[0020] S1. Obtain the original text dataset, and obtain a double-weighted label relationship network according to the original text dataset.
[0021] Specifically, obtain the dataset for the multi-label text classification task through public datasets, data competition platforms, web crawlers, or academic research and literature as the original text dataset. The public datasets include Ren-CECps1.0, which is a multi-label Chinese sentiment corpus containing 37,678 Chinese blog sentences and 11 sentiment labels, where each sentence is assigned one or more sentiments; the THUCNews dataset, which is a news classification dataset from which some data can be extracted for multi-label text classification. For example, the news title can be selected as the text, and multiple topic categories of the news can be used as labels; other NLP datasets, and many NLP-related studies provide open-source datasets that may contain the data required for the multi-label text classification task. It is also possible to extract the text and corresponding labels from business data, which requires data processing and annotation according to specific business scenarios; use a web crawler to scrape relevant text data from the Internet and perform label annotation according to certain rules or algorithms.
[0022] It should be understood that the goal of multi-label text classification is to assign a set of label subsets to each text instance x, where L = {l1, l2,..., l NGiven all possible label spaces, the multi-label text classification task can be formalized as learning a mapping function f(x) → 2 L , to predict the power set of L.
[0023] Furthermore, according to the original text dataset, the obtained double-weighted label relationship network satisfies the following formula: G = (V, E, W v , W e ), where G represents the double-weighted label relationship network, V represents the vertices represented by the labels of the original text dataset, E represents the edges connecting vertex pairs representing the co-occurrence of labels in the dataset, and W v represents the weight aggregation of vertices, and W v = {w v (i) | i ∈ V}, w v (i) represents the occurrence frequency of label i, and W e represents the weight aggregation of edges, and W e = {w e (i, j) | i, j ∈ V}, w e (i, j) represents the co-occurrence intensity between label i and label j.
[0024] S2. Conduct tail-driven sampling in the double-weighted label relationship network, and generate creative labels for the original text dataset according to the sampling results.
[0025] In the embodiment, conducting tail-driven sampling in the double-weighted label relationship network includes the following steps:
[0026] S21. Construct a tail-driven sampling label transition probability model, and obtain the transition probability of labels in the double-weighted label relationship network according to the tail-driven sampling label transition probability model.
[0027] When conducting tail-driven sampling in the double-weighted label relationship network, first calculate the probability of moving from the current label i to another label j. Specifically, obtain the transition probability through the tail-driven sampling label transition probability model, and the tail-driven sampling label transition probability model satisfies the following formula: where q(i → j) represents the transition probability of label i to label j in the double-weighted label relationship network, w e (i, j) represents the co-occurrence intensity between label i and label j, and neighbors(i) represents the neighbors that can transfer with label i.
[0028] S22. Construct a tail-driven sampling label important scarcity feature model, and obtain the important scarcity feature values of labels in the double-weighted label relationship network according to the tail-driven sampling label important scarcity feature model.
[0029] When performing tail-driven sampling in the double-weighted label relationship network, it is also necessary to obtain the important scarcity eigenvalue of the labels in the double-weighted label relationship network. Specifically, the important scarcity eigenvalue of the label is calculated through the tail-driven sampling label important scarcity feature model, and the tail-driven sampling label important scarcity feature model satisfies the following formula: where p(i) represents the important scarcity eigenvalue of label i, S(i) represents the important scarcity score of label i, neighbors(i) represents the neighbors that can transfer with label i, and T represents the temperature parameter for adjusting the distribution smoothness.
[0030] Furthermore, the important scarcity score of label i satisfies the following formula: S(i) = -log(w v (i)), where S(i) represents the important scarcity score of label i, and w v (i) represents the occurrence frequency of label i. It can be understood that according to the information entropy principle, the target distribution can be defined to reflect the important scarcity eigenvalue of the label.
[0031] S23. Combine the transition probability and the important scarcity eigenvalue to evaluate the label transition acceptance in the double-weighted label relationship network, and complete the tail-driven sampling according to the evaluation result.
[0032] In the embodiment, combining the transition probability and the important scarcity eigenvalue to evaluate the label transition acceptance in the double-weighted label relationship network satisfies the following formula: where α(i→j) represents the label transition acceptance in the double-weighted label relationship network, min(·) represents the minimum value function, p(·) represents the important scarcity eigenvalue of the label, q(i→j) represents the transition probability from label i to label j in the double-weighted label relationship network, and q(j→i) represents the transition probability from label j to label i in the double-weighted label relationship network. The label transition acceptance in the double-weighted label relationship network is used to evaluate whether to accept the transition from label i to label j, thereby guiding the sampling result to develop towards the target distribution.
[0033] During the entire sampling process, starting from a tail label, and then gradually transitioning to other labels according to the label transition probability and the label transition acceptance. This process continues until a sufficient number of labels are sampled or a predetermined step limit is reached.
[0034] Further, the creative labels for generating the original text dataset according to the sampling results in step S2 can be obtained by randomly combining and selecting several labels from the tail labels to create combinations, or by selecting frequently co-occurring tail labels based on the degree of association to create combinations, or by predicting the tail labels that may co-occur based on a machine learning model or a neural network model and generating label combinations.
[0035] It should be understood that the obtained creative labels increase the number of rare but important label instances, which is beneficial to solving the data long-tail effect.
[0036] S3. Use the large language model and the creative labels to generate creative label text.
[0037] The large language model refers to an advanced language model with a large number of parameters, such as the GPT model and the LLMs model. In the embodiment, using the large language model and the creative labels to generate creative label text satisfies the following formula: Among them, represents the large language model objective loss function, represents the label matching loss, represents the style consistency loss, and λ represents the large language model hyperparameter. The large language model hyperparameter λ balances the importance of label matching and style consistency, and ensures that the generated text is not only relevant to the sampled labels but also consistent in style with the actual text by optimizing the objective loss function of the large language model, achieving a balance between accuracy and authenticity.
[0038] Specifically, the label matching loss satisfies the following formula: Among them, represents the label matching loss, n represents the number of texts, r φ (X i ) represents the probability that the large language model generates text X i , represents the temperature coefficient for the i-th label to generate the i-th text, represents the temperature coefficient for the i-th label to generate the j-th text.
[0039] Further, in order to effectively align the enhanced input X aug with the corresponding label combination Y, use to show the ideal generation and also guide the large language model to distinguish good and bad generations. In the randomly selected subset {X 1 , Y 1 ; X 2 , Y 2 ; …; X n , Y n}, for each Y iAll are unique. The Jaccard similarity is used to evaluate and rank the degree of similarity, from Y 1 and Y 2 , to Y 1 and Y n . For the label Y 1 , its associated text X 1 is regarded as a positive example, while the text from X 2 to X n is regarded as a negative example, showing a gradually decreasing similarity.
[0040] For the set {X 1 , Y 1 ; X 2 , Y 2 ; …; X n , Y n}, first compare X 1 with X 2 , …, X n , then compare X 2 with X 3 , …, X n , with the aim of aligning the X aug generated by the large language model with the label Y.
[0041] Furthermore, where X contains the tokens x1, ..., x |X| , c(Y 1 ) represents the combination of the labels of positive examples and the prompt information.
[0042] Even further, each comparison involves adjusting the suppression of negative example samples through the temperature coefficient and satisfies the following formula: where s(Y i , Y j ) represents the Jaccard similarity between the label sets Y i and Y j .
[0043] In the embodiment, the style consistency loss satisfies the following formula: where, represents the style consistency loss, t represents the number of tokens of the text, x t represents the sequence of tokens of the text, c(Y) represents the composite text composed of the multi-label set associated with the text and the prompt words, and P φ (x t |c(Y)) represents the probability that the token x t of the text generated by the large language model conforms to the composite text c(Y).
[0044] Randomly select a subset {X 1,Y 1 ; …; X n ,Y n}, where X = {x1, …, x |X|} represents text, which consists of a series of tokens x, and Y = {y1, …, y |Y|} represents the multi-label set associated with X. For Y, we concatenate it with the prompt information to create a composite text c(Y), such as "Generate text for label ", as the input to the large language model.
[0045] S4. Merge the creative label text to obtain multi-label classification enhanced data.
[0046] In the embodiment, the creative labels obtained by tail driving are merged with the creative label text generated by the large language model with the objective loss function optimized, which has a style coordinated, controllable and coherent with the original data set, to obtain multi-label classification enhanced data.
[0047] Please refer to Figure 2 , in the embodiment, in order to efficiently execute a multi-label classification data enhancement method provided by the present invention based on a large language model, the present invention also provides a multi-label classification data enhancement system based on a large language model, including: an input device, an output device, a processor, and a memory. The input device, the output device, the processor, and the memory are interconnected. The memory contains program instructions, and the program instructions are used for the steps of the multi-label classification data enhancement method based on the large language model. The multi-label classification data enhancement system based on a large language model of the present invention has a compact structure and stable performance, and can stably execute a multi-label classification data enhancement method based on a large language model of the present invention, further improving the overall applicability and practical application ability of the present invention.
[0048] In an embodiment, the so-called processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The input device may be used to obtain data information. The output device may be used to output the result obtained from the program instructions included in the computer program stored in the memory provided by the present invention. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0049] In yet another alternative embodiment, please refer to Figure 3 , in order to efficiently execute a multi-label classification data augmentation method based on a large language model provided by the present invention, this embodiment also provides a multi-label classification data augmentation device based on a large language model, as Figure 3 shown, including:
[0050] A memory 10 for storing a computer program; a processor 20 for executing the computer program to implement the above-mentioned multi-label classification data augmentation method based on a large language model.
[0051] The memory 10, the processor 20, a communication interface 31, and a communication bus 32. The memory 10, the processor 20, and the communication interface 31 all complete mutual communication through the communication bus 32.
[0052] In an embodiment, the memory 10 is used to store one or more program instructions, and the memory 10 may store program instructions for implementing the following functions: obtaining an original text data set, obtaining a dual-weighted label relationship network according to the original text data set; performing tail-driven sampling in the dual-weighted label relationship network, and generating creative labels for the original text data set according to the sampling result; using a large language model and the creative labels to generate creative label texts; merging the creative label texts to obtain multi-label classification enhanced data.
[0053] In a possible implementation, the memory 10 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store data created during use.
[0054] In addition, the memory 10 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include NVRAM. The memory stores an operating system, operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0055] The processor 20 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic device. The processor 20 may be a microprocessor or any conventional processor, etc. The processor 20 may call the program stored in the memory 10.
[0056] The communication interface 31 may be an interface of a communication module for connecting to other devices or systems.
[0057] Of course, it should be noted that Figure 3 the structure shown does not constitute a limitation on the multi-label classification data augmentation device based on the large language model in this embodiment. In actual applications, the multi-label classification data augmentation device based on the large language model may include more or fewer components than Figure 3 shown, or combine certain components.
[0058] The embodiment also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above multi-label classification data augmentation method based on the large language model are implemented.
[0059] The storage medium may include various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0060] In summary, the present invention uses existing tags to form a tag relationship network and then performs tail-driven sampling, designs innovative tag combinations considering tag matching and style consistency, increases the number of instances of rare tags, effectively addresses the long-tail distribution problem while maintaining consistency and relevance with the original data, and improves the generalization ability of large language models in practical applications. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A multi-label classification data augmentation method based on large language models, characterized in that, The multi-label classification data augmentation method based on a large language model includes the following steps: Obtain the original text dataset, and obtain a dual-weighted label relationship network according to the original text dataset; Perform tail-driven sampling in the dual-weighted label relationship network, and generate creative labels for the original text dataset according to the sampling results; Use the large language model and the creative labels to generate creative label texts; Merge the creative label texts to obtain multi-label classification enhanced data; The performing tail-driven sampling in the dual-weighted label relationship network includes the following steps: Construct a tail-driven sampling label transition probability model, and obtain the transition probability of labels in the dual-weighted label relationship network according to the tail-driven sampling label transition probability model; Construct a tail-driven sampling label important and scarce feature model, and obtain the important and scarce feature values of labels in the dual-weighted label relationship network according to the tail-driven sampling label important and scarce feature model; Combine the transition probability and the important and scarce feature values to evaluate the label transition acceptance in the dual-weighted label relationship network, and complete tail-driven sampling according to the evaluation results; The tail-driven sampling label transition probability model satisfies the following formula: Among them, represents the transition probability from label to label in the double weighted label relationship network, represents the co-occurrence intensity between label and label in the double weighted label relationship network, represents the neighbors that can transition with label in the double weighted label relationship network. The tail-driven sampling label important and scarce feature model satisfies the following formula: Among them, represents the important scarcity feature value of the label , represents the important scarcity score of the label , represents the neighbor that can be transferred with the label , represents the temperature parameter that regulates the distribution smoothness; Among them, represents the important scarcity score of the label , represents the occurrence frequency of the label ; The combining the transition probability and the important and scarce feature values to evaluate the label transition acceptance in the dual-weighted label relationship network satisfies the following formula: Among them, represents the label transfer acceptance degree in the double-weighted label relationship network, represents the minimum value function, represents the important scarcity feature value of the label, represents the label in the double-weighted label relationship network to label transfer probability, represents the label in the double-weighted label relationship network to label transfer probability.
2. The multi-label classification data augmentation method based on a large language model according to claim 1, wherein The obtaining a dual-weighted label relationship network according to the original text dataset satisfies the following formula: Among them, represents a double-weighted label relationship network, represents the vertex represented by the label of the original text dataset, represents the edge connecting vertex pairs representing the co-occurrence of labels in the dataset, represents the weight aggregation of vertices, , represents the label 's occurrence frequency, represents the weight aggregation of edges, , represents the label and the label 's co-occurrence intensity.
3. A multi-label classification data augmentation method based on a large language model according to claim 1, characterized in that, The using the large language model and the creative labels to generate creative label texts satisfies the following formula: Among them, represents the objective loss function of the large language model, represents the label matching loss, represents the style consistency loss, represents the hyperparameters of the large language model.
4. A multi-label classification data augmentation method based on a large language model according to claim 3, characterized in that, The label matching loss satisfies the following formula: Among them, represents the label matching loss, represents the number of texts, represents the possibility of the large language model generating text of, represents the th label generating the th text's temperature coefficient, represents the th label generating the th text's temperature coefficient.
5. A multi-label classification data augmentation method based on a large language model according to claim 3, characterized in that, The style consistency loss satisfies the following formula: Among them, represents the style consistency loss, represents the number of tokens of the text, represents the series tokens of the text, represents the composite text composed of the multi-label set and the prompt words associated with the text, represents the tokens of the text generated by the large language model that conforms to the composite text probability.
6. A multi-label classification data augmentation system based on a large language model, characterized in that, The multi-label classification data augmentation system based on a large language model includes: an input device, an output device, a processor, and a memory. The input device, the output device, the processor, and the memory are interconnected. The memory includes program instructions, and the program instructions are used to execute the multi-label classification data augmentation method according to any one of claims 1-5.
Citation Information
Patent Citations
Multi-label text classification method and system based on dynamic weight contrast learning
CN114580433A
Long-tail image retrieval method, system and equipment and storage medium
CN117056550A