Label data expansion method, device and equipment

By iteratively training the label prediction model and evaluating it with large language model, the problems of high and low efficiency of manpower labeling in the existing technology are solved, and automated labeling is realized, reducing labor cost and improving labeling efficiency and accuracy.

CN119938913APending Publication Date: 2025-05-06CHINA MOBILE ONLINE SERVICES CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411984692.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The reliance on a large amount of manpower for labeling in the prior art leads to high labor costs, low efficiency, and poor consistency of labeling results.

Method used

By obtaining the initial candidate data set, the label prediction model is iteratively trained based on the data set, and the output of the label prediction model is evaluated and expanded using the large language model, and then the data to be marked automatically.

Benefits of technology

It realizes the labor cost of data annotation, improves the labeling efficiency, and ensures the accuracy and consistency of the annotation results through the evaluation of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938913A_ABST
    Figure CN119938913A_ABST
Patent Text Reader

Abstract

The invention provides an annotation data expansion method, device and equipment, and the method comprises the steps: obtaining an initial candidate data set, the initial candidate data set comprises multiple groups of training data, and each group of training data comprises sample to-be-annotated data and a real label corresponding to the sample to-be-annotated data; and performing iterative training on the label prediction model based on the candidate data set, after each iterative training, respectively inputting the test samples into the label prediction model and the large language model, and predicting the consistency of a first test prediction label output by the label prediction model and a second test prediction label output by the large language model. Expanding the candidate data set; and obtaining to-be-labeled data, inputting the to-be-labeled data into the label prediction model after the multi-round iterative training is completed, and determining a label labeling result corresponding to the to-be-labeled data based on a prediction label output by the label prediction model. According to the invention, the labor cost of data annotation can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, device and equipment for expanding labeled data. Background Art

[0002] Text annotation is a basic and critical task in the field of Natural Language Processing (NLP), which has a direct impact on data quality and model performance. In natural language processing and text annotation tasks, it is common to rely on large annotated data sets to train models. Traditional text annotation methods usually rely on a large amount of manpower for annotation, which is not only costly and inefficient, but also often leads to poor consistency in annotation results due to differences in understanding between annotators. In actual application scenarios, especially in the field of communication customer service, obtaining such a large amount of annotated data is not only time-consuming and cumbersome, but also economically costly. Summary of the invention

[0003] The present invention provides a method, device and equipment for expanding labeled data, so as to solve the defects of the prior art that a large amount of human annotation is relied on and the labor cost is high, thereby reducing the labor cost of data annotation.

[0004] The present invention provides a method for expanding labeled data, comprising: Obtaining an initial candidate data set, wherein the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; Iteratively training the label prediction model based on the candidate data set, inputting the test samples into the label prediction model and the large language model respectively after each iterative training, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; The data to be labeled is obtained, and the data to be labeled is input into the label prediction model after multiple rounds of iterative training, and the label labeling result corresponding to the data to be labeled is determined based on the predicted label output by the label prediction model.

[0005] According to a method for expanding labeled data provided by the present invention, there are multiple label prediction models, and determining the label labeling result corresponding to the data to be labeled based on the predicted labels output by the label prediction models includes: After multiple rounds of iterative training are completed, the weight of the label prediction model is determined based on the sample data to be labeled in the candidate data set, and the weight reflects the correctness of the predicted label output by the label prediction model; Based on the predicted labels output by the label prediction model and the weights of each of the label prediction models, determining the label labeling result corresponding to the data to be labeled, According to a method for expanding labeled data provided by the present invention, the candidate data set is expanded based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model, including: Acquire at least one first test sample as the sample data to be labeled, and add the first test prediction label corresponding to the first test sample into training data to the candidate data set; The first test sample is the test sample corresponding to which the first test prediction label and the second test prediction label are consistent.

[0006] According to a method for expanding labeled data provided by the present invention, the candidate data set is expanded based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model, and further includes: Acquire at least one second test sample as the sample data to be labeled, and add the real label corresponding to the second test sample into training data to the candidate data set; The second test sample is the test sample whose corresponding first test prediction label and second test prediction label are inconsistent.

[0007] According to a method for expanding annotated data provided by the present invention, the iterative training of the label prediction model based on the candidate data set includes: In iterative training, the sample data to be labeled is input into the label prediction model, and a first training loss is determined based on the predicted label output by the label prediction model and the label corresponding to the sample data to be labeled in the candidate data set; When the sample data to be labeled in the candidate data set has a corresponding true label, the sample data to be labeled is input into the large language model, and a second training loss is determined based on the predicted label output by the large language model and the true label corresponding to the sample data to be labeled; A total training loss is determined based on the first training loss and the second training loss, and the label prediction model is updated based on the total training loss.

[0008] According to a method for expanding labeled data provided by the present invention, determining the total training loss based on the first training loss and the second training loss includes: Performing weighted fusion on the first training loss and the second training loss to obtain the total training loss; Among them, the fusion weight of the first training loss is positively correlated with the number of iterations, and the fusion weight of the second training loss is negatively correlated with the number of iterations.

[0009] According to a method for expanding labeled data provided by the present invention, determining the weight of the label prediction model based on the sample data to be labeled in the candidate data set includes: Determine the weight of each label prediction model based on the number of sample data to be labeled corresponding to each label prediction model in the candidate data set; Among them, when the predicted label of the target label prediction model for the sample data to be labeled is consistent with the predicted label output by the large language model for the sample data to be labeled, the label prediction model corresponding to the sample data to be labeled is the target label prediction model.

[0010] According to a method for expanding labeled data provided by the present invention, the step of obtaining an initial candidate data set includes: Reading the sample data to be annotated from real historical data; And / or, inputting a preset label into the large language model, and obtaining the sample data to be annotated corresponding to the preset label output by the large language model.

[0011] The present invention also provides a label data expansion device, comprising: A candidate data set construction module is used to obtain an initial candidate data set, wherein the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; An iterative optimization module, configured to iteratively train the label prediction model based on the candidate data set, input the test samples into the label prediction model and the large language model respectively after each iterative training, and expand the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; The labeling module is used to obtain the data to be labeled, input the data to be labeled into the label prediction model after multiple rounds of iterative training, and determine the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any of the above-described methods for expanding labeled data is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for expanding the labeled data as described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, including a computer program, which implements any of the above-mentioned annotation data expansion methods when executed by a processor.

[0015] The present invention provides a method, device and equipment for expanding labeled data, wherein the method comprises: obtaining an initial candidate data set, wherein the initial candidate data set comprises multiple groups of training data, wherein each group of training data comprises sample data to be labeled and a true label corresponding to the sample data to be labeled; iteratively training a label prediction model based on the candidate data set, and after each iterative training, inputting the test samples into the label prediction model and the large language model respectively, and expanding the candidate data set based on the consistency of a first test prediction label output by the label prediction model and a second test prediction label output by the large language model; obtaining data to be labeled, and inputting the data to be labeled into the label prediction model after multiple rounds of iterative training, and determining the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0016] In this way, by introducing the capabilities of the large language model, the large model is used as a teacher model to expand the sample data to be labeled and the expanded sample data to be labeled is used to guide the iterative training of the label prediction model, and the label prediction model is further evaluated through the large language model, so that only a small amount of data to be labeled with real labels is needed to achieve high-performance training of the label prediction model. Finally, the label labeling result of the data to be labeled is obtained through the predicted label output by the label prediction model, so that the data to be labeled can be accurately labeled without a large amount of manpower participating in the labeling, thereby reducing the labor cost of data labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a flow chart of the annotation data expansion method provided by the present invention.

[0019] Figure 2 This is an overall architecture diagram for implementing the annotation data expansion method provided by the present invention.

[0020] Figure 3It is an iterative optimization flow chart of the annotation data expansion method provided by the present invention.

[0021] Figure 4 It is a structural schematic diagram of the annotation data expansion device provided by the present invention.

[0022] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] Combine the following Figure 1-Figure 3 The method for expanding the labeled data provided by the present invention is described. Figure 1 As shown, the annotation data expansion method includes the steps of: S110, obtaining an initial candidate data set, where the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; S120, iteratively training the label prediction model based on the candidate data set, inputting the test samples into the label prediction model and the large language model respectively after each iterative training, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; S130, obtaining the data to be labeled, inputting the data to be labeled into the label prediction model after multiple rounds of iterative training, and determining the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0025] The method provided by the present invention, by introducing the capability of a large language model, uses the large model as a teacher model to expand sample data to be labeled, and uses the expanded sample data to be labeled to guide the iterative training of a label prediction model, and further implements the evaluation of the label prediction model through the large language model, thereby realizing high-performance training of the label prediction model with only a small amount of data to be labeled with real labels, and finally obtains the label labeling result of the data to be labeled by the predicted labels output by each label prediction model, thereby realizing that the data to be labeled can be accurately labeled without the need for a large amount of manpower to participate in the labeling, thereby reducing the labor cost of data labeling.

[0026] In the method provided by the present invention, an initial candidate data set is first obtained, and the labels corresponding to the sample data to be annotated included in the initial candidate data set are real labels, and the real labels are manually annotated labels. In a possible implementation, a small amount of sample data to be annotated can be read, and the corresponding real labels can be obtained by manual annotation, so as to construct the initial candidate data set. However, in some scenarios, the acquisition of sample data to be annotated may also require a lot of manpower. For example, for customer service scenarios, it is necessary to search for sentences with specific meanings in a large number of conversations, or in some scenarios, sentences with specific meanings have a long time span. In these scenarios, the sample data to be annotated cannot be read directly from the database, but need to be manually screened, so that even if an initial candidate data set including a small amount of training data is constructed, a high manpower and time cost is required. In view of this defect, in a possible implementation of the method provided by the present invention, obtaining the initial candidate data set includes: Read sample data to be annotated from real historical data; and / or The preset labels are input into the large language model to obtain the sample data to be annotated corresponding to the preset labels output by the large language model.

[0027] That is to say, the sample data to be labeled in the initial candidate data set can be all extracted from real historical data, or partly extracted from real historical data and partly generated by a large language model. When extracting sample data to be labeled from real historical data and manually labeling it, a small and representative sample set of labeled data is constructed. , taking the scenario of customer call problem identification in the mobile communication customer service field as an example, Indicates the call text data that needs to be labeled, Si represents the i-th sentence of the call, and LRi represents the real customer demand label to which the Si-th sentence belongs.

[0028] Multiple preset labels are pre-set, and the labels corresponding to all sample data to be labeled represent the probabilities of the sample data to be labeled with each preset label. Different preset labels can be set according to different application scenarios. For example, in the customer source problem identification scenario in the field of mobile communication customer service, the preset labels are divided into three levels. The first level contains 9 categories, such as handling, complaints (including complaints), consultation (including inquiries), and reporting of faults; the second level contains 45 categories, such as cancellation, activation, change, etc.; the third level contains 34 categories (some second-level problems do not have third-level problem detailed labels), such as cancellation, termination, downgrade, renewal, upgrade, etc.

[0029] When manually labeling the sample data to obtain the true label, a mode of two people for labeling and one person for quality inspection can be adopted. For example, three experts are required: labeler P1, labeler P2, and quality inspector P3. P1 and P2 are assigned the data independently (they are unaware of each other's labeling) and get the labeling answers respectively. , .when and If they are consistent, P3 can directly select the annotation label of step S1 ;when and If there is any inconsistency, P3 needs to review again , and refer to and The result determines Labels for .

[0030] When generating sample data to be annotated through a large language model, first build a prompt engineering in a structured data format to obtain an annotated data set extracted from real historical data and manually annotated: , construct a QA pair in JSON format (including: question (Question), answer (Answer) and related context (Context)). At the same time, build a task description based on the application scenario (such as the customer call problem recognition scenario in the mobile communication customer service field) and obtain As a Question, get As the Answer, get the entire call text according to the serial number, remove the redundant fields, and keep only the text content as the Context. , From the same serial number conversation ( ), then one of the constructed JSON objects is a QA pair: {"Context":" ", "Questions":[{"Question":" ", "Answer":" "}, {"Question":" ", "Answer":" "}]}.

[0031] Then, based on the statistics of the sample data to be labeled constructed by historical real data, the existing distribution of the number of preset labels is obtained to determine the number of sample data to be labeled that needs to be supplemented for each preset label. The existing distribution of the number of preset labels can be expressed as: ,in, Indicates the 45 combined label categories defined in the application scenario. express The number of labeled data sets corresponding to the combined label categories under the label data set is used to determine the number of sample data to be labeled for different preset labels. For example, the initial candidate data set requires 300 sample data to be labeled for each preset label (in order to obtain better training results). is the i-th combined label, for The number of annotations, for The number of additional annotations required is The large model of the prompting project constructed in step S3.1 is generated of Data, composed of Need to supplement the data set (The corresponding true labels are ). Therefore, the pseudo data set generated by the large model is: . Further, after obtaining the pseudo dataset generated by the large language model Afterwards, it can be directly added to the initial candidate data set, or the labels corresponding to the sample data to be labeled in the pseudo data set generated by the large language model can be manually verified to ensure the accuracy of the real labels in the initial candidate data set.

[0032] For the initial candidate dataset ,Before applying it to train the label prediction model, preprocessing operations such as data cleaning, text vectorization, and necessary feature engineering are performed.

[0033] Although the large language model can output relatively accurate prediction labels, the large language model has a huge number of parameters, and has the disadvantages of long inference time and high cost. In the method provided by the present invention, the label prediction model is trained to automatically label the labeled data. The label prediction model is a small model compared to the large language model, that is, it has fewer parameters and a simpler structure. Figure 2As shown, in the method provided by the present invention, the number of sample data to be labeled included in the initial candidate data set is small. In the process of iteratively training the label prediction model based on the candidate data set, the candidate data set is continuously expanded under the guidance of the large language model (LLM), thereby improving the performance of the label prediction model.

[0034] Specifically, in the method provided by the present invention, there are multiple label prediction models, for example, Figure 3 As shown in Figure 1, three structured label prediction models can be constructed, namely, the few_shot model structure, the text classification model structure, and the text similarity model structure. Of course, it is understandable that the number and structure of label prediction models are not limited to Figure 3 The situation shown in , may also adopt other numbers and other structures.

[0035] In the method provided by the present invention, the label prediction model is trained for multiple rounds of iterations. In each round of iteration training, a specific number of sample data to be labeled are selected from the current candidate data set, or all the sample data to be labeled and the corresponding labels in the current candidate data set are selected to optimize the parameters of the label prediction model. After each round of iteration training is completed, the effect of the label prediction model is evaluated by the consistency of the second test prediction label output by the large language model for the test sample and the first test prediction label output by the label prediction model for the test sample, and the candidate data set is expanded.

[0036] Specifically, the candidate data set is expanded based on the consistency of the first test prediction label and the second test prediction label, including: Acquire at least one first test sample as sample data to be labeled, and add the first test prediction label corresponding to the first test sample into training data to be added to the candidate data set; The first test sample is a test sample whose corresponding first test prediction label and second test prediction label are consistent.

[0037] like Figure 3 As shown in the figure, taking the label prediction model as an example, 300 test samples are randomly obtained from a large number of unlabeled call samples. , three sets of results are predicted by three label prediction models respectively , , ; At the same time, through the large language model Make a prediction and get another set of prediction results . The output results predicted by the three sets of label prediction models , , Respectively with the prediction results of the large language model Comparing one by one and combining the test samples, three sets of comparison data sets can be obtained , , , take the intersection of the three data sets to get the filtered data set .by and For example, let the i-th test sample be The results predicted by the label prediction model Prediction results with large models If consistent, , similarly (The prediction result of the jth test sample label prediction model is consistent with the prediction result of the large model), (The prediction result of the k-th test sample label prediction model is consistent with the prediction result of the large model). , , Take the intersection of the three data sets, set , then filter the data set .

[0038] Based on the screening data set, the candidate data set can be expanded by directly adding all the screening data sets to the candidate data set, thus obtaining the augmented annotated data set. In another possible implementation, based on the quantity distribution of sample data to be labeled under each preset label in the current candidate data, data with preset labels corresponding to a smaller number of sample data to be labeled are selected from the screening data set and added to the candidate data set.

[0039] Furthermore, in a possible implementation, in order to ensure data diversity and the proportion of real data, prevent excessive reliance on large language models, and further improve the training efficiency and performance of the label prediction model, the candidate data set is expanded based on the consistency of the first test prediction label and the second test prediction label, further comprising: At least one second test sample is obtained as sample data to be labeled, and the real label corresponding to the second test sample constitutes training data and is added to the candidate data set; The second test sample is a test sample whose corresponding first test prediction label and second test prediction label are inconsistent.

[0040] By adding the second test sample as the sample data to be labeled into the candidate data set, it is possible to add real labeling results to the candidate data set used for training the label prediction model, thereby preventing excessive reliance on large language models and achieving the effect of ensuring the prediction performance of the label prediction model with only a small amount of manpower cost.

[0041] After each round of iterative training, it is determined whether the set number of iterations has been reached. If the set number of cycles has not been reached, a new round of iterative training is performed on the label prediction model. That is, 300 test samples are randomly obtained from a large number of unlabeled samples. , repeat the steps of expanding the candidate data set to obtain a new round of training data set, , and again determine whether the set number of iterations has been reached. If the set number of cycles has not been reached, then loop step S5.4 until the set target number of cycles is reached, and record the training data set at this time as (The training dataset used in the last iteration of training).

[0042] It is not difficult to see from the above description that in the process of training the label prediction model, some of the candidate data sets used for training have real labels, while others do not, but correspond to labels predicted by the large language model. In the method provided by the present invention, the real labels are used to obtain the corresponding loss of the output of the large language model, thereby improving the accuracy of the training loss, thereby realizing the deep mining of the information of the real labels and improving the performance of the label prediction model obtained by training. Specifically, multiple rounds of iterative training of multiple label prediction models are performed based on the candidate data sets, including: In iterative training, the sample data to be labeled is input into the label prediction model, and the first training loss is determined based on the predicted label output by the label prediction model and the label corresponding to the sample data to be labeled in the candidate data set; When there is a corresponding true label for the sample data to be labeled in the candidate data set, the sample data to be labeled is input into the large language model, and a second training loss is determined based on the predicted label output by the large language model and the true label corresponding to the sample data to be labeled; Based on the first training loss and the second training loss, a total training loss is determined.

[0043] That is to say, for a set of sample data to be labeled with real labels , use the large language model to make predictions and output the predicted value z of each preset label, so as to obtain the predicted label output by the large model For a single sample S1, the loss function Loss_Softmax predicted by the large language model can be calculated using the following process: y is the one-hot encoding of L_Real (the true category label of each sample is converted into a binary vector. The length of this vector is equal to the total number of categories, where only one element is 1, indicating the category to which the sample belongs, and the rest of the elements are 0. For example, if the total number of categories C=5, and the true category of a sample is 3, then the one-hot encoding of the true label of the sample is: , is the predicted probability, and C is the number of categories. Then for each sample, Fork loss function ,in is the original output of the model for the i-th category, is the one-hot encoding of the true label, is the predicted probability.

[0044] The large model As an additional supervisory signal, it guides the training and fine-tuning of the small model. By continuously approaching the difference between the predicted label and the true label, the confidence of the label closest to the true label is adjusted to help the label prediction model better learn the relationship between data features and categories. Specifically, the original loss function of the small model (such as cross entropy loss) ) is combined with the Softmax loss of the large model to form a total loss function ,in is a hyperparameter between 0 and 1, which is used to balance the importance of the two loss functions. Initialize the hyperparameter set of each small model ( value, learning rate, number of samples for training at a time (also called batch size), etc.), and in each iteration, calculate the gradient according to the total loss function, and use the optimization algorithm Stochastic Gradient Descent (SGD) or Adaptive Moment Estimation (Adam) to update the parameter hyperparameter set of the label prediction model, such as Values, learning rates, etc. Through the learned hyperparameter set, the classification model effect of each round is obtained, its performance is evaluated on the validation set, and the best hyperparameter set is selected based on the effect.

[0045] Further, based on the first training loss and the second training loss, a total training loss is determined, including: Perform weighted fusion of the first training loss and the second training loss to obtain the total training loss; Among them, the fusion weight of the first training loss is positively correlated with the number of iterations, and the fusion weight of the second training loss is negatively correlated with the number of iterations. During the iterative training process, since the performance of the initial label prediction model is not high, in the method provided by the present invention, the fusion weight of the first training loss is positively correlated with the number of iterations, and the fusion weight of the second training loss is negatively correlated with the number of iterations, that is, the loss obtained from the prediction result based on the label prediction model increases with the increase of the number of iterations, while the loss obtained from the prediction result based on the large language model decreases with the increase of the number of iterations.

[0046] In the early stage of iterative training, a higher weight is set for the loss of the large language model, so that the large language model can provide guidance in the early stage of training the label prediction model. In the later stage of iterative training, as the performance of the label prediction model gradually improves, the weight of the loss of the label prediction model is increased, thereby improving the training efficiency of the label prediction model and obtaining a label prediction model with better performance.

[0047] In one implementation of the method provided by the present invention, the weight distribution of each label prediction model is automatically guided and adjusted based on the evaluation results of the label prediction model by the large model. The label labeling result corresponding to the data to be labeled is determined based on the predicted label output by the label prediction model, including: After multiple rounds of iterative training, the weight of the label prediction model is determined based on the sample data to be labeled in the candidate data set. The weight reflects the accuracy of the predicted label output by the label prediction model. Based on the predicted labels output by the label prediction model and the weights of each label prediction model, the label annotation results corresponding to the data to be labeled are determined.

[0048] After multiple rounds of iterative training are completed, the accuracy of the predicted labels output by each label prediction model is determined based on the sample data to be labeled in the candidate data used in the last round of iterative training, thereby determining the weight of each label prediction model.

[0049] Specifically, the weight of the label prediction model is determined based on the sample data to be labeled in the candidate data set, including: Determine the weight of each label prediction model based on the number of sample data to be labeled corresponding to each label prediction model in the candidate data set; Among them, when the predicted label of the target label prediction model for the sample data to be labeled is consistent with the predicted label output by the large language model for the sample data to be labeled, the label prediction model corresponding to the sample data to be labeled is the target label prediction model.

[0050] When the set number of cycles is reached and the iterative optimization of the label prediction model is completed, ,in, Represents the sample data to be labeled in the candidate data set as a set of inference samples. Indicates the label of the sample data to be labeled, as a comparison label sample. Get the predicted labels for the sample data to be labeled output by each label prediction model. Taking three label prediction models as an example, three sets of results can be obtained: , , .

[0051] The predicted labels output by the three-group label prediction model , , The prediction results of the large language model for the sample data to be labeled Compare one by one and get three groups of comparison results , , (Initialized to 0). and For example, let i start from 1. When the i-th test sample The results of the small model prediction Prediction results with large models When consistent, +1, loop until i=n, ​​then the initial weight of the few_shot model Similarly, the initial weights of the text classification model , the initial weight of the text similarity model .

[0052] In a possible implementation, after obtaining the initial weights of each label prediction model, the initial weights are normalized to obtain the final weights of each label prediction model. That is, , , .

[0053] After obtaining the weights of each label prediction model, the label prediction model can be used to predict the labels of the data to be labeled, thereby realizing automatic labeling of a large amount of data to be labeled without the need for manual labeling. That is, for the data to be labeled, they are input into the label prediction model respectively to obtain the predicted labels output by each label prediction model respectively, and each predicted label is weighted and fused according to the weight corresponding to the label prediction model to obtain the labeled labels of the data to be labeled. For example, the label prediction model consists of three: ,in , Indicates the annotation label. , and Represent the predicted labels of the three label prediction models respectively.

[0054] The method provided by the present invention extracts similar unlabeled data from a small amount of labeled data by means of a collaborative approach of a large language model and a small language model (label prediction model), thereby improving data utilization, assisting in data labeling for the construction of various scene models, and reducing labeling costs and model construction costs. At the same time, the advantages of large and small models can be fully utilized to improve the quality of labeling. In addition, the large language model is used as a teacher model to score the trained label prediction model, and the label prediction model with a higher score is cyclically iteratively optimized, and the voting fusion weight distribution of each label prediction model is automatically guided and adjusted through the final evaluation results, which not only improves the reliability of labeling, but also ensures the universal scene adaptation of the model, and is applicable to a variety of fields, including mobile communication customer service, e-commerce customer service, financial services, medical health and other field labeling, with strong versatility and adaptability, and also realizes the automation of model guidance and adjustment, reduces manual intervention, and reduces labor costs.

[0055] The following is a description of the label data expansion device provided by the present invention. The label data expansion device described below and the label data expansion method described above can be referred to in correspondence with each other. Figure 4 As shown, the annotation data expansion device provided by the present invention includes: The candidate data set construction module 410 is used to obtain an initial candidate data set, where the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; An iterative optimization module 420 is used to iteratively train the label prediction model based on the candidate data set, and after each iterative training, input the test sample into the label prediction model and the large language model respectively, and expand the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; The labeling module 430 is used to obtain the data to be labeled, input the data to be labeled into the label prediction model after multiple rounds of iterative training, and determine the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0056] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the label data expansion method, which includes: obtaining an initial candidate data set, wherein the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and a true label corresponding to the sample data to be labeled; iteratively training the label prediction model based on the candidate data set, after each iterative training, inputting the test sample into the label prediction model and the large language model respectively, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; obtaining the data to be labeled, inputting the data to be labeled into the label prediction model after multiple rounds of iterative training, and determining the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0057] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0058] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the labeled data expansion method provided by the above methods, and the method includes: obtaining an initial candidate data set, the initial candidate data set includes multiple groups of training data, each group of training data includes sample data to be labeled and a true label corresponding to the sample data to be labeled; iteratively training a label prediction model based on the candidate data set, and after each iterative training, inputting the test samples into the label prediction model and the large language model respectively, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; obtaining the data to be labeled, inputting the data to be labeled into the label prediction model after multiple rounds of iterative training, and determining the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0059] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the labeled data expansion method provided by the above-mentioned methods, the method comprising: obtaining an initial candidate data set, the initial candidate data set comprising multiple groups of training data, each group of training data comprising sample data to be labeled and a true label corresponding to the sample data to be labeled; iteratively training a label prediction model based on the candidate data set, and after each iterative training, inputting the test samples into the label prediction model and the large language model respectively, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; obtaining the data to be labeled, inputting the data to be labeled into the label prediction model after multiple rounds of iterative training, and determining the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

[0060] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0061] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for expanding labeled data, characterized in that: include: Obtaining an initial candidate data set, wherein the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; Iteratively training the label prediction model based on the candidate data set, inputting the test samples into the label prediction model and the large language model respectively after each iterative training, and expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; The data to be labeled is obtained, and the data to be labeled is input into the label prediction model after multiple rounds of iterative training, and the label labeling result corresponding to the data to be labeled is determined based on the predicted label output by the label prediction model.

2. The method for expanding labeled data according to claim 1, characterized in that: There are multiple label prediction models, and determining the label labeling result corresponding to the data to be labeled based on the predicted labels output by the label prediction models includes: After multiple rounds of iterative training are completed, the weight of the label prediction model is determined based on the sample data to be labeled in the candidate data set, and the weight reflects the correctness of the predicted label output by the label prediction model; Based on the predicted labels output by the label prediction model and the weights of each of the label prediction models, a label labeling result corresponding to the data to be labeled is determined.

3. The method for expanding labeled data according to claim 1, characterized in that: The expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model includes: Acquire at least one first test sample as the sample data to be labeled, and add the first test prediction label corresponding to the first test sample into training data to the candidate data set; The first test sample is the test sample corresponding to which the first test prediction label and the second test prediction label are consistent.

4. The method for expanding the labeled data according to claim 3, characterized in that: The method of expanding the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model further includes: Acquire at least one second test sample as the sample data to be labeled, and add the real label corresponding to the second test sample into training data to the candidate data set; The second test sample is the test sample whose corresponding first test prediction label and second test prediction label are inconsistent.

5. The method for expanding labeled data according to claim 1, characterized in that: The performing multiple rounds of iterative training on the label prediction model based on the candidate data set includes: In iterative training, the sample data to be labeled is input into the label prediction model, and a first training loss is determined based on the predicted label output by the label prediction model and the label corresponding to the sample data to be labeled in the candidate data set; When the sample data to be labeled in the candidate data set has a corresponding true label, the sample data to be labeled is input into the large language model, and a second training loss is determined based on the predicted label output by the large language model and the true label corresponding to the sample data to be labeled; A total training loss is determined based on the first training loss and the second training loss, and the label prediction model is updated based on the total training loss.

6. The method for expanding the labeled data according to claim 5, characterized in that: The determining a total training loss based on the first training loss and the second training loss comprises: Performing weighted fusion on the first training loss and the second training loss to obtain the total training loss; Among them, the fusion weight of the first training loss is positively correlated with the number of iterations, and the fusion weight of the second training loss is negatively correlated with the number of iterations.

7. The method for expanding labeled data according to claim 1, characterized in that: The determining the weight of the label prediction model based on the sample data to be labeled in the candidate data set includes: Determine the weight of each label prediction model based on the number of sample data to be labeled corresponding to each label prediction model in the candidate data set; Among them, when the predicted label of the target label prediction model for the sample data to be labeled is consistent with the predicted label output by the large language model for the sample data to be labeled, the label prediction model corresponding to the sample data to be labeled is the target label prediction model.

8. The method for expanding labeled data according to claim 1, characterized in that: The obtaining of the initial candidate data set includes: Reading the sample data to be annotated from real historical data; And / or, inputting a preset label into the large language model, and obtaining the sample data to be annotated corresponding to the preset label output by the large language model.

9. A device for expanding labeled data, characterized in that: include: A candidate data set construction module is used to obtain an initial candidate data set, wherein the initial candidate data set includes multiple sets of training data, each set of training data includes sample data to be labeled and true labels corresponding to the sample data to be labeled; An iterative optimization module, configured to iteratively train the label prediction model based on the candidate data set, input the test samples into the label prediction model and the large language model respectively after each iterative training, and expand the candidate data set based on the consistency of the first test prediction label output by the label prediction model and the second test prediction label output by the large language model; The labeling module is used to obtain the data to be labeled, input the data to be labeled into the label prediction model after multiple rounds of iterative training, and determine the label labeling result corresponding to the data to be labeled based on the predicted label output by the label prediction model.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for expanding labeled data according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • News content core-oriented labeling method, equipment and medium

    CN120470127A

  • Anti-money laundering risk rule generation method and system based on multi-modal large language model

    CN120598567A