Label generation method and device, electronic equipment and computer storage medium

By screening and generating target labels for candidate sample data, the problems of high cost, long time and low accuracy in the prior art label acquisition method are solved, and the efficiency of training large language model and the accuracy of labels are improved.

CN120105087AActive Publication Date: 2025-06-06BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411845772.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-06
Estimated Expiration
2044-12-13

Smart Images

  • Figure CN120105087A_ABST
    Figure CN120105087A_ABST
Patent Text Reader

Abstract

The invention provides a label generation method and device, electronic equipment and a computer storage medium, and the method comprises the steps: obtaining a labeled sample data set and a label-free sample data set, the labeled sample data set comprises a plurality of initial sample data and a real label of each initial sample data, and the label-free sample data set comprises a plurality of initial sample data; the unlabeled sample data set comprises multiple pieces of to-be-labeled sample data; determining a plurality of sample quality evaluation scores corresponding to the plurality of initial sample data; screening the plurality of initial sample data based on the plurality of sample quality evaluation scores to obtain a plurality of candidate sample data; and for each piece of to-be-labeled sample data in the plurality of pieces of to-be-labeled sample data, generating a target label corresponding to each piece of to-be-labeled sample data according to each piece of to-be-labeled sample data and the plurality of pieces of candidate sample data. According to the embodiment provided by the scheme, the tag generation efficiency and the tag accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of label processing technology, and in particular to a label generation method, device, electronic device and computer storage medium. Background Art

[0002] The training of large language models usually requires the model to learn the mapping relationship between input and output from labeled data. A large number of high-quality real labels can help the model identify the pattern between the input text and the corresponding output label, and then adjust the parameters of the model. However, the label acquisition method in related technologies is usually based on manual labeling or semi-supervised learning. These methods face the defects of high cost, long time period, low label accuracy and easy introduction of human bias, which leads to low efficiency of large language model training. Summary of the invention

[0003] The embodiments of the present application provide a label generation method, device, electronic device and computer storage medium, which can improve the label generation efficiency and the label accuracy. The above technical solution is as follows:

[0004] In a first aspect, an embodiment of the present application provides a label generation method, the method comprising:

[0005] Obtain a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled;

[0006] Determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data;

[0007] Based on the multiple sample quality assessment scores, the multiple initial sample data are screened to obtain multiple candidate sample data;

[0008] For each of the plurality of sample data to be labeled, a target label corresponding to each of the sample data to be labeled is generated according to the sample data to be labeled and the plurality of candidate sample data.

[0009] In a possible implementation, the determining of the plurality of sample quality assessment scores corresponding to the plurality of initial sample data includes:

[0010] For each of the plurality of initial sample data, determining an average similarity score between the initial sample data and the plurality of initial sample data;

[0011] Determine the influence score of the initial sample data in the labeled sample data set;

[0012] Determine the information increment score of the initial sample data in the labeled sample data set;

[0013] The sample quality assessment score corresponding to the above-mentioned initial sample data is determined according to the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score, and then multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data are obtained.

[0014] In a possible implementation, the determining of the average similarity score between the initial sample data and the plurality of initial sample data includes:

[0015] Determine the semantic vector corresponding to the initial sample data;

[0016] Generating a plurality of first similarity scores between the initial sample data and the plurality of initial sample data according to the semantic vector corresponding to the initial sample data;

[0017] An average similarity score corresponding to the initial sample data is determined according to the multiple first similarity scores.

[0018] In a possible implementation, determining the influence score of the initial sample data in the labeled sample data set includes:

[0019] Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining the first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0020] The current initial sample data is removed from the labeled sample data set to obtain a labeled sample data subset;

[0021] Predicting the labels of the labeled sample data subset by the large language model to obtain the second label prediction probability of the true labels of all the initial sample data in the labeled sample data subset output by the large language model;

[0022] The influence score of the initial sample data in the labeled sample data set is determined according to the first label prediction probability and the second label prediction probability.

[0023] In a possible implementation, the determining of the information increment score of the initial sample data in the labeled sample data set includes:

[0024] Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining the first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0025] Determine the self entropy data of the initial sample data based on the first label prediction probability;

[0026] Determine a plurality of cross entropy data between the initial sample data and the plurality of initial sample data according to the first label prediction probability;

[0027] The information increment score corresponding to the initial sample data is determined according to the self-entropy data and the multiple cross-entropy data.

[0028] In a possible implementation, determining the sample quality assessment score corresponding to the initial sample data according to the average similarity score, the influence score, and the information increment score includes:

[0029] Obtaining a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score;

[0030] The sample quality assessment score corresponding to the initial sample data is determined according to the first weight, the second weight, the third weight, the average similarity score, the influence score and the information increment score.

[0031] In a possible implementation, the initial sample data and the sample quality assessment scores correspond one to one, and the multiple initial sample data are screened based on the multiple sample quality assessment scores to obtain multiple candidate sample data, including:

[0032] Sorting the multiple initial sample data according to the multiple sample quality assessment scores to obtain a sorting result;

[0033] A first preset number of initial sample data with the highest sample quality evaluation scores in the above-mentioned sorting results are determined as a plurality of candidate sample data.

[0034] In a possible implementation, the generating of target labels corresponding to the sample data to be labeled according to the sample data to be labeled and the plurality of candidate sample data includes:

[0035] Determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data;

[0036] Determine a plurality of vocabulary coverage scores between the aforementioned sample data to be labeled and the aforementioned plurality of candidate sample data;

[0037] Determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores;

[0038] Determining target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores;

[0039] The true label corresponding to the above target sample data is determined as the target label corresponding to the above sample data to be labeled.

[0040] In a possible implementation, the determining of a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data includes:

[0041] Extracting the first vocabulary set included in the above-mentioned sample data to be annotated by using a preset part-of-speech recognition tool;

[0042] Extracting a second vocabulary set included in each candidate sample data in the plurality of candidate sample data by using the preset part-of-speech recognition tool;

[0043] According to the first vocabulary set and the second vocabulary set included in the candidate sample data, a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data are determined.

[0044] In a possible implementation, the method further includes:

[0045] Generate prompt information corresponding to the sample data to be labeled according to the target labels corresponding to the sample data to be labeled;

[0046] The preset large language model is trained based on the above sample data to be labeled and the above prompt information.

[0047] In a second aspect, an embodiment of the present application provides a label generation device, the device comprising:

[0048] An acquisition module is used to acquire a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled;

[0049] A determination module, used to determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data;

[0050] A screening module, used to screen the multiple initial sample data based on the multiple sample quality assessment scores to obtain multiple candidate sample data;

[0051] The first generating module is used to generate, for each of the plurality of sample data to be labeled, a target label corresponding to each of the sample data to be labeled according to the sample data to be labeled and the plurality of candidate sample data.

[0052] In a possible implementation, the determination module includes:

[0053] A first determining unit, configured to determine, for each of the plurality of initial sample data, an average similarity score between the initial sample data and the plurality of initial sample data;

[0054] A second determining unit, used to determine the influence score of the initial sample data in the labeled sample data set;

[0055] A third determining unit, used to determine the information increment score of the initial sample data in the labeled sample data set;

[0056] The fourth determination unit is used to determine the sample quality assessment score corresponding to the above-mentioned initial sample data according to the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score, and then obtain multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data.

[0057] In a possible implementation, the first determining unit includes:

[0058] A first determining subunit, used to determine a semantic vector corresponding to the initial sample data;

[0059] A generating subunit, configured to generate a plurality of first similarity scores between the initial sample data and the plurality of initial sample data according to the semantic vector corresponding to the initial sample data;

[0060] The second determining subunit is used to determine an average similarity score corresponding to the initial sample data according to the multiple first similarity scores.

[0061] In a possible implementation, the second determining unit includes:

[0062] A first prediction subunit is used to predict the labels of the labeled sample data set through a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0063] A removal subunit, used to remove the current initial sample data from the labeled sample data set to obtain a labeled sample data subset;

[0064] A second prediction subunit is used to predict the labels of the labeled sample data subset through the large language model to obtain a second label prediction probability of the true labels of all initial sample data in the labeled sample data subset output by the large language model;

[0065] The third determination subunit is used to determine the influence score of the initial sample data in the labeled sample data set according to the first label prediction probability and the second label prediction probability.

[0066] In a possible implementation, the third determining unit includes:

[0067] A third prediction subunit is used to predict the labels of the labeled sample data set through a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0068] A fourth determining subunit, configured to determine the self entropy data of the initial sample data based on the first label prediction probability;

[0069] A fifth determining subunit, configured to determine a plurality of cross entropy data between the initial sample data and the plurality of initial sample data according to the first label prediction probability;

[0070] The sixth determination subunit is used to determine the information increment score corresponding to the above-mentioned initial sample data based on the above-mentioned self-entropy data and the above-mentioned multiple cross-entropy data.

[0071] In a possible implementation manner, the fourth determining unit includes:

[0072] An acquisition subunit, used to acquire a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score;

[0073] The seventh determination subunit is used to determine the sample quality assessment score corresponding to the above-mentioned initial sample data based on the above-mentioned first weight, the above-mentioned second weight, the above-mentioned third weight, the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score.

[0074] In a possible implementation, the screening module includes:

[0075] A fifth determining unit, configured to sort the plurality of initial sample data according to the plurality of sample quality assessment scores to obtain a sorting result;

[0076] The sixth determining unit is used to determine a first preset number of initial sample data with the highest sample quality assessment scores in the above-mentioned sorting results as a plurality of candidate sample data.

[0077] In a possible implementation, the first generating module includes:

[0078] A seventh determination unit, configured to determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data;

[0079] An eighth determining unit, configured to determine a plurality of vocabulary coverage scores between the aforementioned sample data to be labeled and the aforementioned plurality of candidate sample data;

[0080] a ninth determining unit, configured to determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores;

[0081] a tenth determining unit, configured to determine target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores;

[0082] The eleventh determining unit is used to determine the real label corresponding to the above-mentioned target sample data as the target label corresponding to the above-mentioned sample data to be labeled.

[0083] In a possible implementation manner, the eighth determining unit includes:

[0084] A first extraction subunit is used to extract a first vocabulary set included in each of the sample data to be annotated by using a preset part-of-speech recognition tool;

[0085] A second extraction subunit is used to extract a second vocabulary set included in each candidate sample data of the plurality of candidate sample data by using the preset part-of-speech recognition tool;

[0086] The twelfth determining unit is used to determine a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data according to the first vocabulary set and the second vocabulary set included in the candidate sample data.

[0087] In a possible implementation, the above device further includes:

[0088] A second generating module is used to generate prompt information corresponding to the sample data to be labeled according to the target labels corresponding to the sample data to be labeled;

[0089] The training module is used to train the preset large language model based on the above-mentioned sample data to be labeled and the above-mentioned prompt information.

[0090] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory;

[0091] The above-mentioned memory stores a computer program, and the above-mentioned computer program is suitable for being loaded by the above-mentioned processor and executing the steps of the method provided by the first aspect of the embodiment of the present application or any possible implementation method of the first aspect.

[0092] In a fourth aspect, an embodiment of the present application provides a computer storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method provided by the first aspect of the embodiment of the present application or any possible implementation method of the first aspect.

[0093] The embodiment of the present application obtains a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes multiple initial sample data and the true label of each initial sample data, and the unlabeled sample data set includes multiple sample data to be labeled; determines multiple sample quality assessment scores corresponding to the multiple initial sample data; based on the multiple sample quality assessment scores, screens the multiple initial sample data to obtain multiple candidate sample data; and for each sample data to be labeled in the multiple sample data to be labeled, generates a target label corresponding to each sample data to be labeled according to the sample data to be labeled and the multiple candidate sample data. By using the labeled initial sample data to screen out the candidate sample data with higher quality, the model can refer to more effective samples when generating the target label, and then combines the characteristic information of the candidate sample data in the label generation process of the unlabeled sample data, so that the generated target label is closer to the true label, reduces the mislabeling rate, and improves the efficiency and accuracy of label generation, which helps to improve the training efficiency when training a large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0095] Figure 1 A structural schematic diagram of a label generation system provided by an exemplary embodiment of the present application;

[0096] Figure 2 A flowchart of a label generation method provided by an exemplary embodiment of the present application;

[0097] Figure 3 A flowchart of a method for determining multiple candidate sample data provided by an exemplary embodiment of the present application;

[0098] Figure 4 A flowchart of a method for determining a target tag provided by an exemplary embodiment of the present application;

[0099] Figure 5 A schematic diagram of the structure of a label generating device provided by an exemplary embodiment of the present application;

[0100] Figure 6 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0101] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0102] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0103] Please refer to the following Figure 1 , which exemplarily shows a structural diagram of a label generation system provided in an embodiment of the present application. Figure 1 As shown, the system includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 are connected via a network, such as a wired or wireless network connection.

[0104] Among them, the terminal device 110 can be used to display a graphical user interface, and the terminal device 110 is used to interact with the user through the graphical user interface, for example, downloading and installing the corresponding client and running it through the terminal device 110. In the embodiment of the present application, the terminal device 110 can also be used by relevant personnel to upload labeled sample data sets and unlabeled sample data sets. The terminal device 110 can send the labeled sample data sets and the unlabeled sample data sets to the server 120, so that the server 120 determines the multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data; and based on the above-mentioned multiple sample quality assessment scores, the above-mentioned multiple initial sample data are screened to obtain multiple candidate sample data; and then for each of the above-mentioned multiple sample data to be labeled, the target label corresponding to each of the above-mentioned sample data to be labeled is generated according to the above-mentioned sample data to be labeled and the above-mentioned multiple candidate sample data, and the obtained target label is fed back to the terminal device 110.

[0105] Optionally, the terminal device 110 itself can also directly determine multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data based on the labeled sample data sets and unlabeled sample data sets uploaded by relevant personnel; based on the above-mentioned multiple sample quality assessment scores, the above-mentioned multiple initial sample data are screened to obtain multiple candidate sample data; for each sample data to be labeled in the above-mentioned multiple sample data to be labeled, generate a target label corresponding to each sample data to be labeled based on the above-mentioned sample data to be labeled and the above-mentioned multiple candidate sample data.

[0106] Optionally, in the labeled sample data set, each initial sample data may correspond to one or more real labels, wherein the real labels may be manually labeled data or data generated by an automated or semi-automated labeling system.

[0107] Optionally, the labeled sample data set is a data set containing a smaller amount of data, and the unlabeled sample data set is a data set containing a larger amount of data. The labeled sample data set is not a subset of the unlabeled sample data set, and the labeled sample data set has all the features of the unlabeled sample data set, that is, the labeled sample data can represent the unlabeled sample data set.

[0108] Optionally, the sample quality assessment score corresponding to the initial sample data is used to indicate the quality or importance of the initial sample data. The sample quality assessment score can be calculated from a combination of multiple indicators that can reflect the value and information contribution of the initial sample data to model training.

[0109] Optionally, the plurality of candidate sample data are included in the plurality of initial sample data, that is, the plurality of candidate sample data are subsets screened from the labeled sample data set.

[0110] Optionally, the target label corresponding to each sample data to be labeled is a predicted label generated for each sample data to be labeled, and the target label indicates the category to which the corresponding sample data to be labeled belongs or the label to be assigned.

[0111] An exemplary embodiment of the present application provides a label generation method. The label generation method can be applied to the above terminal device. For details, please refer to Figure 2 , which exemplarily shows a flow chart of a label generation method provided in an embodiment of the present application. Figure 2 As shown, the tag generation method includes the following S21-S24:

[0112] S21. Obtain a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled.

[0113] In some embodiments, the initial sample data may be a sentence, and the true label of the initial sample data may represent the category or attribute of the sentence.

[0114] In some embodiments, in the labeled sample data set, each initial sample data may correspond to one or more real labels, wherein the real labels may be manually labeled data or data generated by an automated or semi-automated annotation system.

[0115] In some embodiments, the labeled sample data set is a data set containing a smaller amount of data, and the unlabeled sample data set is a data set containing a larger amount of data. The labeled sample data set is not a subset of the unlabeled sample data set, and the labeled sample data set has all the features of the unlabeled sample data set, that is, the labeled sample data can represent the unlabeled sample data set.

[0116] S22. Determine multiple sample quality assessment scores corresponding to the multiple initial sample data.

[0117] In some embodiments, the initial sample data and the sample quality assessment score correspond one to one.

[0118] The sample quality evaluation score corresponding to the initial sample data is used to indicate the quality or importance of the initial sample data. The sample quality evaluation score can be calculated from multiple indicators, which can reflect the value and information contribution of the initial sample data to the model training.

[0119] Optionally, the sample quality assessment score can be calculated comprehensively based on at least one of the following indicators: the average similarity score between the above-mentioned initial sample data and the above-mentioned multiple initial sample data, the influence score of the above-mentioned initial sample data in the above-mentioned labeled sample data set, and the information increment score corresponding to the above-mentioned initial sample data, etc.

[0120] S23. Based on the above-mentioned multiple sample quality assessment scores, the above-mentioned multiple initial sample data are screened to obtain multiple candidate sample data.

[0121] The above-mentioned multiple candidate sample data are included in the above-mentioned multiple initial sample data, that is, the multiple candidate sample data are subsets screened from the above-mentioned labeled sample data set.

[0122] S24: for each of the plurality of sample data to be labeled, generate a target label corresponding to each of the sample data to be labeled according to the sample data to be labeled and the plurality of candidate sample data.

[0123] The target label corresponding to each sample data to be labeled is the predicted label generated for each sample data to be labeled, and the target label indicates the category to which the corresponding sample data to be labeled belongs or the label to be assigned. Specifically, the target label is the result inferred by the model or algorithm based on the relationship between the sample data to be labeled and the candidate sample data (high-quality sample data with known true labels).

[0124] In S24, the process of generating target labels is actually to automatically label the unlabeled sample data to be labeled, and assign corresponding labels to each sample data to be labeled, so that it changes from unlabeled data to labeled data. These target labels can be used for subsequent large language model training, data analysis or other applications.

[0125] The embodiment of the present application obtains a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes multiple initial sample data and the true label of each initial sample data, and the unlabeled sample data set includes multiple sample data to be labeled; determines multiple sample quality assessment scores corresponding to the multiple initial sample data; based on the multiple sample quality assessment scores, screens the multiple initial sample data to obtain multiple candidate sample data; and for each sample data to be labeled in the multiple sample data to be labeled, generates a target label corresponding to each sample data to be labeled according to the sample data to be labeled and the multiple candidate sample data. By using the labeled initial sample data to screen out the candidate sample data with higher quality, the model can refer to more effective samples when generating the target label, and then combines the characteristic information of the candidate sample data in the label generation process of the unlabeled sample data, so that the generated target label is closer to the true label, reduces the mislabeling rate, and improves the efficiency and accuracy of label generation, which helps to improve the training efficiency when training a large language model.

[0126] In some embodiments, in S22, the determining of the plurality of sample quality assessment scores corresponding to the plurality of initial sample data includes S221-S224:

[0127] S221. For each initial sample data in the plurality of initial sample data, determine an average similarity score between the initial sample data and the plurality of initial sample data.

[0128] The average similarity score between the initial sample data and the multiple initial sample data is used to measure the similarity between the initial sample data and the labeled sample data set.

[0129] S222: Determine the influence score of the initial sample data in the labeled sample data set.

[0130] Among them, the influence score of the above-mentioned initial sample data in the above-mentioned labeled sample data set can be used to measure the contribution or influence of the initial sample data on model training and prediction.

[0131] S223: Determine the information increment score of the initial sample data in the labeled sample data set.

[0132] The above-mentioned information increment score indicates the amount of additional information provided by the initial sample data relative to other initial sample data. The higher the information increment score, the greater the difference in the predicted distribution between the current initial sample data and other initial sample data, which provides more new information and is more valuable for model training.

[0133] S224. Determine a sample quality assessment score corresponding to the initial sample data according to the average similarity score, the influence score and the information increment score, and then obtain multiple sample quality assessment scores corresponding to the multiple initial sample data.

[0134] In an embodiment of the present application, by calculating the average similarity score, influence score and information increment score of the samples, sample data that is valuable for model training can be effectively screened out, thereby reducing the interference of low-quality or duplicate information samples and improving the training effect and prediction accuracy of the model.

[0135] In some embodiments, in S221, for each of the multiple initial sample data, an average similarity score between the initial sample data and the multiple initial sample data is determined, including S2211-S2213:

[0136] S2211. Determine the semantic vector corresponding to the above initial sample data.

[0137] In some embodiments, the text content of the initial sample data can be converted into a semantic vector through a related semantic embedding model (such as a sentence embedding model SBERT). The semantic vector can effectively capture the features and semantic information of the initial sample data.

[0138] In addition, the semantic vector corresponding to the above initial sample data can also be determined by an all-distil Multi-Language Mini Language Model (all-MiniLM), such as all-MiniLM-L6-v2 (6th layer, version 2), all-MiniLM-L12-v2 (12th layer, version 2); or the semantic vector corresponding to the above initial sample data can be determined by an all-language distilled robust optimization method (first version) (All Distilled Robustly Optimized BERT Approach Version 1, all-distilroberta-v1), an all-language masked and permuted network (basic version 2) (All Maskedand Permuted Network Base Version 2, d allmpnet-base-v2), etc.

[0139] Optionally, the labeled sample dataset can be represented as D l , the initial sample data and its true label in the labeled sample data set can be expressed as D i , the labeled sample data set includes the number of initial sample data N, where i is any positive integer not less than N, representing the number of the initial sample data. Specifically, each initial sample data and its true label D i It can be expressed as D i =(x i ,y i ), where x i represents the initial sample data, y i Represents the initial sample data x i The real label.

[0140] Determine the semantic vector e corresponding to each initial sample data through the relevant semantic embedding model i , and then obtain the semantic vector set E={e 1 ,……e N}, where e 1 ~e N Represents the semantic vector corresponding to each initial sample data.

[0141] S2212: Generate a plurality of first similarity scores between the initial sample data and the plurality of initial sample data according to the semantic vector corresponding to the initial sample data.

[0142] In some embodiments, the plurality of first similarity scores are similarity scores between the current initial sample data and all initial sample data in the labeled sample data set (including the current initial sample data itself).

[0143] In some embodiments, a plurality of first similarity scores between the initial sample data and the plurality of initial sample data may be generated based on a cosine similarity formula according to the semantic vector corresponding to the initial sample data.

[0144] In other embodiments, multiple first similarity scores between the initial sample data and the multiple initial sample data may be calculated by using Euclidean distance, Manhattan distance, Minkowski distance, Jaccard similarity, Pearson correlation coefficient, etc.

[0145] Optionally, the current initial sample data x i Compared with the above labeled sample dataset D l Any initial sample data x among all the initial sample data (including the current initial sample data itself) j The similarity score can be expressed as cosine(e i ,e j ).

[0146] S2213: Determine an average similarity score corresponding to the initial sample data according to the multiple first similarity scores.

[0147] In some embodiments, in S2213, determining the average similarity score corresponding to the initial sample data according to the multiple first similarity scores includes: taking the average of the multiple first similarity scores as the average similarity score corresponding to the initial sample data.

[0148] Optionally, the average similarity score corresponding to the initial sample data is It can be expressed as:

[0149]

[0150] The lower the average similarity score corresponding to the initial sample data, the more it can reflect the difference between the initial sample data and other initial sample data in the entire labeled sample data set.

[0151] In the embodiment of the present application, the semantic vector can extract the deep semantic features of the text data, and can not only be limited to the surface vocabulary similarity during the processing, but also perform deeper semantic comparisons to more accurately measure the similarity between sample data and improve the quality of sample data screening.

[0152] In some embodiments, in S222, determining the influence score of the initial sample data in the labeled sample data set includes S2221-S2224:

[0153] S2221. Predict the labels of the labeled sample data set using a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model.

[0154] Among them, the above-mentioned preset large language model is a large language model that needs to be trained or optimized.

[0155] In some embodiments, the initial sample data of the labeled sample data set corresponds one-to-one to the first label prediction probability.

[0156] In some embodiments, the labeled sample data set (including the current initial sample data) can be input into a preset large language model to obtain the output result of the large language model, and then the prediction information of each initial sample data is extracted from the output result, and the prediction probability of the true label of each initial sample given by the large language model is recorded as the first label prediction probability. l Any initial sample data x j The true label y j The first label prediction probability can be expressed as P LM (y j |x j ,D i ), where y i is any initial sample data in the labeled sample dataset.

[0157] S2222: remove the current initial sample data from the labeled sample data set to obtain a labeled sample data subset.

[0158] Among them, the above-mentioned labeled sample data subset is a subset of the above-mentioned labeled sample data set.

[0159] S2223. Predict the labels of the labeled sample data subset using the large language model to obtain the second label prediction probability of the true labels of all initial sample data in the labeled sample data subset output by the large language model.

[0160] In some embodiments, the initial sample data in the labeled sample data subset corresponds one-to-one to the second label prediction probability.

[0161] In some embodiments, the above-mentioned labeled sample data subset (excluding the current initial sample data) can be input into a preset large language model to obtain the output result of the above-mentioned large language model, and then the prediction information of the initial sample data is extracted from the output result, and the predicted probability of the true label of each initial sample given by the large language model is recorded as the second label prediction probability.

[0162] Any initial sample data x in the labeled sample data subset output by the large language model j The true label y j The second label prediction probability can be expressed as P LM (y j |x j ), where y j is any initial sample data in the labeled sample data subset.

[0163] S2224. Determine the influence score of the initial sample data in the labeled sample data set based on the first label prediction probability and the second label prediction probability.

[0164] In some embodiments, in S2224, the influence score of the initial sample data in the labeled sample data set is determined based on the first label prediction probability and the second label prediction probability, including: determining the difference between the first label prediction probability and the second label prediction probability as the influence score of the initial sample data in the labeled sample data set.

[0165] Optionally, the influence score of the initial sample data in the labeled sample data set can be expressed as:

[0166]

[0167] The greater the influence score of the initial sample data in the above labeled sample data set, the greater the contribution of the initial sample to the final prediction output result of the large language model.

[0168] In the embodiment of the present application, by comparing the prediction results of the model with and without the current initial sample data, the initial sample data that has a greater impact on the model prediction effect can be identified. These initial sample data are more representative or contain unique information, which can help the model better learn classification rules and improve the model training efficiency.

[0169] In some embodiments, in S223, the step of determining the information increment score of the initial sample data in the labeled sample data set includes S2231-S2234:

[0170] S2231. Predict the labels of the labeled sample data set using a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model.

[0171] In some embodiments, the specific process of S2231 is consistent with the above-mentioned process of S2221 and will not be repeated here.

[0172] S2232. Determine the self-entropy data of the initial sample data based on the first label prediction probability.

[0173] The self-entropy data of the initial sample data can be used to measure the uncertainty of the large language model's prediction of the current initial sample data.

[0174] In some embodiments, the self entropy data H of the initial sample data can be determined by the following formula: θ (D i ):

[0175]

[0176] Among them, p i In S2231, the preset large language model is used for the current initial sample data x i The true label y i Predicted first label prediction probability.

[0177] For each initial sample data, the smaller the self-entropy data, the higher the certainty of the large language model's prediction of the initial sample data, and the larger the self-entropy data, the lower the certainty of the large language model's prediction of the initial sample data.

[0178] S2233. Determine multiple cross entropy data between the above-mentioned initial sample data and the above-mentioned multiple initial sample data according to the above-mentioned first label prediction probability.

[0179] In some embodiments, the above-mentioned initial sample data corresponds one-to-one to the above-mentioned multiple cross entropy data, that is, the cross entropy data between the current initial sample data and any initial sample data in the labeled sample data set is calculated to obtain multiple cross entropy data.

[0180] In some embodiments, the current initial sample data x can be calculated by the following formula i Any initial sample data x in the labeled sample data set j The cross entropy data H θ (x j |D i ):

[0181]

[0182] Among them, p j (x j ) is the initial sample data x through the preset large language model in S2231 j The first label prediction probability of the true label prediction. i (D i ) is the initial sample data x through the preset large language model in S2231 i The first label prediction probability of the true label prediction.

[0183] S2234. Determine the information increment score corresponding to the above-mentioned initial sample data based on the above-mentioned self-entropy data and the above-mentioned multiple cross-entropy data.

[0184] In some embodiments, in S2234, determining the information increment score corresponding to the initial sample data according to the self entropy data and the plurality of cross entropy data includes:

[0185] Calculate the difference between each cross entropy data in the above multiple cross entropy data and the above self entropy data, and then calculate the average value of the difference between each cross entropy data and the above self entropy data, and determine the average value as the information increment score corresponding to the current above initial sample data.

[0186] Specifically, the initial sample data x i The corresponding information increment score info(D i ) can be expressed as:

[0187]

[0188] In the embodiment of the present application, the information increment score reflects the additional information brought by a certain initial sample relative to other samples, helping the model to identify which samples provide unique semantic information, thereby increasing the effectiveness of model training; and by selecting samples with high information increment scores, the learning content of the model can be enriched, so that the model can better cover diverse inputs, thereby improving the generalization ability of the model. Especially in small sample scenarios, the information increment score helps to focus on high-value samples and improve the performance of the model in data-scarce situations.

[0189] In some embodiments, in S224, the sample quality assessment score corresponding to the initial sample data is determined according to the average similarity score, the influence score and the information increment score, including S2241-S2242:

[0190] S2241. Obtain a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score.

[0191] Among them, the first weight, the second weight and the third weight are pre-set, and the specific values ​​can be flexibly set according to actual conditions, and this application does not make specific limitations on this.

[0192] S2242. Determine a sample quality assessment score corresponding to the initial sample data based on the first weight, the second weight, the third weight, the average similarity score, the influence score and the information increment score.

[0193] In some embodiments, in S2242, determining the sample quality assessment score corresponding to the initial sample data according to the first weight, the second weight, the third weight, the average similarity score, the influence score, and the information increment score includes:

[0194] Calculate the first product of the first weight and the inverse of the average similarity score, calculate the second product of the second weight and the influence score, and calculate the third product of the third weight and the inverse of the information increment score; determine the sum of the first product, the second product and the third product as the sample quality assessment score corresponding to the initial sample data.

[0195] Optionally, the initial sample data x j The corresponding sample quality assessment score (D j ) can be expressed as:

[0196]

[0197] Among them, α 1 is the first weight, α 2 is the second weight, α 3 The third weight.

[0198] In the embodiment of the present application, the three dimensions of sample evaluation are combined by means of linear transformation, which can adapt to different scenarios and sample focus points by changing the weight coefficients of the dimensions. At the same time, the linear transformation has high interpretability, which effectively improves the comprehensiveness, accuracy and reliability of sample data evaluation.

[0199] In some embodiments, in S23, the multiple initial sample data are screened based on the multiple sample quality assessment scores to obtain multiple candidate sample data, including S231-S232:

[0200] S231. Sort the multiple initial sample data according to the multiple sample quality evaluation scores to obtain a sorting result.

[0201] In some embodiments, the plurality of initial sample data may be sorted according to the sample quality assessment scores from large to small or from small to large.

[0202] S232: Determine a first preset number of initial sample data with the highest sample quality assessment scores in the above-mentioned sorting results as a plurality of candidate sample data.

[0203] In some embodiments, the first preset number m<<N.

[0204] In some embodiments, the plurality of candidate sample data is a candidate context example set D m .

[0205] In the embodiment of the present application, by sorting the sample quality assessment scores, high-quality samples can be preferentially selected to participate in model training, thereby improving the prediction accuracy and overall performance of the model. The first preset number m is much smaller than the total number of samples N, indicating that only a small number of high-quality samples are selected for training or use. This can significantly reduce the amount of data while maintaining model performance, reduce training costs and consumption of computing resources, and then the model can complete training faster, thereby improving training efficiency.

[0206] The above is the first stage of the label generation method provided by this application. Figure 3 A flowchart of a method for determining multiple candidate sample data provided in an embodiment of the present application is shown in FIG. Figure 3 The method for determining the plurality of candidate sample data includes the following S301-S308:

[0207] S301: Obtain a labeled sample data set and an unlabeled sample data set.

[0208] S302: Determine an average similarity score between the initial sample data and multiple initial sample data.

[0209] S303: Determine the influence score of the initial sample data in the labeled sample data set.

[0210] S304: Determine the information increment score of the initial sample data in the labeled sample data set.

[0211] S305. Obtain a first weight corresponding to the average similarity score, a second weight corresponding to the above-mentioned influence score, and a third weight corresponding to the information increment score.

[0212] S306. Determine the sample quality assessment score corresponding to the initial sample data according to the first weight, the second weight, the third weight, the average similarity score, the influence score and the information increment score, and then obtain multiple sample quality assessment scores.

[0213] S307. Sort the multiple initial sample data according to the multiple sample quality evaluation scores to obtain a sorting result.

[0214] S308: Determine a first preset number of initial sample data with the highest sample quality evaluation scores in the sorting results as a plurality of candidate sample data.

[0215] The specific steps of S301-S308 are consistent with the above-mentioned S21-S23 and will not be repeated here.

[0216] In an embodiment of the present application, in the first stage, a candidate context example set (including multiple candidate sample data) is constructed to screen out candidate sample data with higher sample quality assessment scores to reduce the amount of data for subsequent data processing, thereby effectively saving computing resources.

[0217] Furthermore, the second stage of the label generation method provided by the present application is described below.

[0218] In some embodiments, in S24, generating target labels corresponding to the sample data to be labeled and the plurality of candidate sample data according to the sample data to be labeled includes S241-S245:

[0219] S241: Determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data.

[0220] In some embodiments, the candidate sample data and the second similarity scores correspond one to one.

[0221] In some embodiments, the semantic vector corresponding to each sample data to be labeled and the semantic vector corresponding to each candidate sample data can be determined by a relevant semantic embedding model, and then multiple second similarity scores between the sample data to be labeled and each candidate sample data are calculated by the cosine similarity formula to obtain multiple second similarity scores between each sample data to be labeled and the above-mentioned multiple candidate sample data.

[0222] S242: Determine a plurality of vocabulary coverage scores between each of the sample data to be labeled and the plurality of candidate sample data.

[0223] The candidate sample data and the vocabulary coverage scores correspond to each other. The vocabulary coverage scores between the sample data to be labeled and each candidate sample data are used to measure the degree of overlap between the sample data to be labeled and each candidate sample data at the vocabulary level.

[0224] In some embodiments, in S242, the determining of multiple vocabulary coverage scores between the above-mentioned sample data to be labeled and the above-mentioned multiple candidate sample data includes S2421-S2423:

[0225] S2421. Extracting a first vocabulary set included in each of the sample data to be annotated by using a preset part-of-speech recognition tool.

[0226] In some embodiments, the first vocabulary set may include at least one vocabulary related to the shallow representation of the sample, such as nouns, adjectives, and adverbs contained in the sample data to be labeled.

[0227] S2422, extracting a second vocabulary set included in each candidate sample data in the plurality of candidate sample data by using the preset part-of-speech recognition tool;

[0228] In some embodiments, the second vocabulary set may include at least one vocabulary related to the shallow representation of the sample, such as nouns, adjectives, and adverbs contained in the candidate sample data.

[0229] S2423: Determine a plurality of vocabulary coverage scores between the plurality of sample data to be labeled and the plurality of candidate sample data according to the first vocabulary set and the second vocabulary set included in the plurality of candidate sample data.

[0230] The unlabeled sample dataset can be represented as D u , each sample data to be labeled in the unlabeled sample data set can be expressed as d i The candidate sample data can be expressed as d j , j is the number of each candidate sample data, and j is a positive integer not less than the first preset number m.

[0231] In some embodiments, each sample data to be labeled d i With candidate sample data d j The vocabulary coverage score between cover(d i ,d j ) can be determined by the following formula:

[0232]

[0233] Where ∈ is the offset coefficient, which is used to prevent the vocabulary coverage score from being 0 and to adjust the size ratio of the vocabulary coverage score. i ,d j ) ranges from (∈, 1+∈], and the vocabulary coverage score can be used as a scaling factor for cosine similarity to avoid focusing only on the semantic information between samples when looking for similar samples. set(·) represents the vocabulary set extracted by the above-mentioned preset part-of-speech recognition tool.

[0234] In an embodiment of the present application, by identifying shallow vocabulary (such as nouns, adjectives, adverbs, etc.) in the sample, the shallow features of the sample to be labeled can be more comprehensively understood, helping the model to focus on the diversity of samples during the labeling process. This vocabulary coverage can capture the representational features of different samples and improve the breadth and accuracy of sample screening. And the vocabulary coverage score can be used as a scaling factor for cosine similarity, combining semantic information with the matching degree of shallow features when calculating sample similarity, avoiding focusing only on deep semantic information. As a result, the model can more comprehensively evaluate the similarity between samples and improve the accuracy of selecting candidate samples.

[0235] S243: Determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores.

[0236] Among them, the candidate sample data corresponds to the corrected similarity score one by one.

[0237] In some embodiments, in S243, multiple corrected similarity scores corresponding to the multiple candidate sample data are determined based on the multiple second similarity scores and the multiple vocabulary coverage scores, including: taking the product of the second similarity score and the vocabulary coverage score between each sample data to be labeled and each candidate sample data as the corrected similarity score corresponding to the candidate sample data, so as to obtain the multiple corrected similarity scores corresponding to the multiple candidate sample data.

[0238] In some downstream domain tasks, there are usually some professional vocabulary, which rarely exists in the pre-trained corpus of the sentence vector model. The professional vocabulary features are identified by the part-of-speech parser and added to the similarity calculation. As a weight, it will not change the contextual semantic information of the sentence itself. For example, the cosine similarity of the sample data "I love to eat tomatoes" and the sample data "I don't like to eat tomatoes" is less than 0, and the vocabulary coverage score of the two is 1+∈. After multiplication, the absolute value of the cosine similarity becomes larger but the direction remains unchanged. The vocabulary coverage score is designed to make the semantically similar and lexically similar samples more similar, and the coverage score of samples with dissimilar expressions but similar vocabulary can reduce their similarity to prevent them from being mistakenly selected into the context learning example.

[0239] S244: Determine target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores.

[0240] In some embodiments, a second preset number (k) of candidate sample data with the highest corresponding corrected similarity scores among the above-mentioned multiple candidate sample data can be determined as target sample data, and the second preset number is smaller than the first preset number.

[0241] Specifically, k target sample data may be determined from multiple candidate sample data using a K-Nearest Neighbor (KNN) algorithm.

[0242] S245: Determine the true label corresponding to the target sample data as the target label corresponding to the sample data to be labeled.

[0243] The samples to be labeled and their target labels are added to the labeled sample dataset for subsequent model training or analysis.

[0244] In the embodiment of the present application, the similarity of semantic vectors and vocabulary coverage scores are used to accurately match the most similar candidate sample data for the unlabeled sample data to be labeled, and its true label is assigned to the sample to be labeled, thereby realizing automatic labeling, saving labor costs and time costs, and effectively improving the efficiency of label generation. In addition, this process makes full use of the existing high-quality labeled sample data and improves the accuracy of label generation.

[0245] The above is the second stage of the label generation method provided by this application. Figure 4 A flow chart of a target tag determination method provided in an embodiment of the present application is shown as follows: Figure 4 The method for determining the plurality of candidate sample data includes the following S401-S407:

[0246] S401: Determine multiple second similarity scores between each sample data to be labeled and multiple candidate sample data.

[0247] S402: extracting a first vocabulary set included in each sample data to be annotated by using a preset part-of-speech recognition tool.

[0248] S403: extracting a second vocabulary set included in each candidate sample data from a plurality of candidate sample data by using a preset part-of-speech recognition tool.

[0249] S404: Determine a plurality of vocabulary coverage scores between each sample data to be labeled and a plurality of candidate sample data according to the first vocabulary set and the second vocabulary set included in each candidate sample data.

[0250] S405: Determine a plurality of modified similarity scores corresponding to a plurality of candidate sample data according to a plurality of second similarity scores and a plurality of vocabulary coverage scores.

[0251] S406: Determine target sample data from multiple candidate sample data according to multiple modified similarity scores.

[0252] S407: Determine the true label corresponding to the target sample data as the target label corresponding to the sample data to be labeled.

[0253] The specific steps of S401-S407 are consistent with the above-mentioned S24 and will not be repeated here.

[0254] In the second phase, this application focuses on the construction of an example recaller to recall the k samples that are most similar to the currently input sample to be labeled from the candidate context example set as context examples. By combining the semantic relationships of the samples to be labeled and the syntactic information based on shallow surface vocabulary such as entities, patterns, and specific attributes, the corresponding target labels are automatically generated for the samples to be labeled, thereby improving the generalization of the large language model across multiple fields.

[0255] Furthermore, in some embodiments, the above method further includes:

[0256] According to the target labels corresponding to the sample data to be labeled, prompt information corresponding to the sample data to be labeled is generated; and a preset large language model is trained based on the sample data to be labeled and the prompt information.

[0257] In some embodiments, the sample data to be labeled and the corresponding target label can be filled into a preset prompt template to generate specific prompt information. Then, prompt information containing the target label is created for each sample data to be labeled, which can guide the large language model to learn the correct output.

[0258] In the embodiment of the present application, the target label of the sample data to be labeled is used to generate corresponding prompt information, and the large language model is trained based on this, which can effectively improve the performance of the model on specific tasks. This process makes full use of the automatically generated labels, reduces the workload of manual labeling, and accelerates the iteration and application of the model.

[0259] Please refer to the following Figure 5 , which is a schematic diagram of the structure of a label generation device provided by an exemplary embodiment of the present application. Figure 5 As shown, the label generating device 500 includes:

[0260] An acquisition module 501 is used to acquire a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled;

[0261] A determination module 502 is used to determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data;

[0262] A screening module 503 is used to screen the multiple initial sample data based on the multiple sample quality evaluation scores to obtain multiple candidate sample data;

[0263] The first generating module 504 is used to generate a target label corresponding to each of the sample data to be labeled among the plurality of sample data to be labeled according to the sample data to be labeled and the plurality of candidate sample data.

[0264] In a possible implementation, the determining module 502 includes:

[0265] A first determining unit, configured to determine, for each of the plurality of initial sample data, an average similarity score between the initial sample data and the plurality of initial sample data;

[0266] A second determining unit, used to determine the influence score of the initial sample data in the labeled sample data set;

[0267] A third determining unit, used to determine the information increment score of the initial sample data in the labeled sample data set;

[0268] The fourth determination unit is used to determine the sample quality assessment score corresponding to the above-mentioned initial sample data according to the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score, and then obtain multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data.

[0269] In a possible implementation, the first determining unit includes:

[0270] A first determining subunit, used to determine a semantic vector corresponding to the initial sample data;

[0271] A generating subunit, configured to generate a plurality of first similarity scores between the initial sample data and the plurality of initial sample data according to the semantic vector corresponding to the initial sample data;

[0272] The second determining subunit is used to determine an average similarity score corresponding to the initial sample data according to the multiple first similarity scores.

[0273] In a possible implementation, the second determining unit includes:

[0274] A first prediction subunit is used to predict the labels of the labeled sample data set through a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0275] A removal subunit, used to remove the current initial sample data from the labeled sample data set to obtain a labeled sample data subset;

[0276] A second prediction subunit is used to predict the labels of the labeled sample data subset through the large language model to obtain a second label prediction probability of the true labels of all initial sample data in the labeled sample data subset output by the large language model;

[0277] The third determination subunit is used to determine the influence score of the initial sample data in the labeled sample data set according to the first label prediction probability and the second label prediction probability.

[0278] In a possible implementation, the third determining unit includes:

[0279] A third prediction subunit is used to predict the labels of the labeled sample data set through a preset large language model to obtain a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0280] A fourth determining subunit, configured to determine the self entropy data of the initial sample data based on the first label prediction probability;

[0281] A fifth determining subunit, configured to determine a plurality of cross entropy data between the initial sample data and the plurality of initial sample data according to the first label prediction probability;

[0282] The sixth determination subunit is used to determine the information increment score corresponding to the above-mentioned initial sample data based on the above-mentioned self-entropy data and the above-mentioned multiple cross-entropy data.

[0283] In a possible implementation manner, the fourth determining unit includes:

[0284] An acquisition subunit, used to acquire a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score;

[0285] The seventh determination subunit is used to determine the sample quality assessment score corresponding to the above-mentioned initial sample data based on the above-mentioned first weight, the above-mentioned second weight, the above-mentioned third weight, the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score.

[0286] In a possible implementation, the screening module 503 includes:

[0287] A fifth determining unit, configured to sort the plurality of initial sample data according to the plurality of sample quality assessment scores to obtain a sorting result;

[0288] The sixth determining unit is used to determine a first preset number of initial sample data with the highest sample quality assessment scores in the above-mentioned sorting results as a plurality of candidate sample data.

[0289] In a possible implementation, the first generating module 504 includes:

[0290] A seventh determination unit, configured to determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data;

[0291] An eighth determining unit, configured to determine a plurality of vocabulary coverage scores between the aforementioned sample data to be labeled and the aforementioned plurality of candidate sample data;

[0292] a ninth determining unit, configured to determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores;

[0293] a tenth determining unit, configured to determine target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores;

[0294] The eleventh determining unit is used to determine the real label corresponding to the above-mentioned target sample data as the target label corresponding to the above-mentioned sample data to be labeled.

[0295] In a possible implementation manner, the eighth determining unit includes:

[0296] A first extraction subunit is used to extract a first vocabulary set included in each of the sample data to be annotated by using a preset part-of-speech recognition tool;

[0297] A second extraction subunit is used to extract a second vocabulary set included in each candidate sample data of the plurality of candidate sample data by using the preset part-of-speech recognition tool;

[0298] The twelfth determining unit is used to determine a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data according to the first vocabulary set and the second vocabulary set included in the candidate sample data.

[0299] In a possible implementation, the apparatus 500 further includes:

[0300] A second generating module is used to generate prompt information corresponding to the sample data to be labeled according to the target labels corresponding to the sample data to be labeled;

[0301] The training module is used to train the preset large language model based on the above-mentioned sample data to be labeled and the above-mentioned prompt information.

[0302] The division of the modules in the above-mentioned label generation device 500 is only for illustration. In other embodiments, the label generation device can be divided into different modules as needed to complete all or part of the functions of the above-mentioned label generation device. The implementation of each module in the label generation device provided in the embodiment of this specification can be in the form of a computer program. The computer program can be run on a terminal or a server. The program modules constituted by the computer program can be stored in the memory of the terminal or the server. When the computer program is executed by the processor, all or part of the steps of the label generation method described in the embodiment of this specification are implemented.

[0303] See next Figure 6 , which is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Figure 6 As shown, the electronic device 600 may include: a processor 610 and a memory 620 , and may also include a user interface 630 , a network interface 640 and a communication bus 650 .

[0304] Among them, the processor 610 may include one or more processing cores. The processor 610 uses various interfaces and lines to connect various parts of the entire electronic device 600, and executes various functions and processes data of the electronic device 600 by running or executing instructions, programs, code sets or instruction sets stored in the memory 620, and calling data stored in the memory 620. Optionally, the processor 610 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 610 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 610, and it can be implemented separately through a chip.

[0305] The memory 620 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 620 includes a non-transitory computer-readable storage medium. The memory 620 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a receiving function, a control function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments, etc. The memory 620 may optionally be at least one storage device located away from the aforementioned processor 610. As Figure 6 As shown, the memory 620 as a computer storage medium may include an operating system, a network communication module, a user interface module, and program instructions.

[0306] Optionally, the communication bus 650 is used to realize the connection and communication between these components. The user interface 630 may include a display screen (Display), a camera (Camera), and may also include a standard wired interface and a wireless interface; the network interface 640 may optionally include a standard wired interface and a wireless interface (such as a WI FI interface).

[0307] exist Figure 6 In the electronic device 600 shown, the processor 610 may be used to call the program instructions stored in the memory 620 and specifically perform the following operations:

[0308] Obtain a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled;

[0309] Determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data;

[0310] Based on the multiple sample quality assessment scores, the multiple initial sample data are screened to obtain multiple candidate sample data;

[0311] For each of the plurality of sample data to be labeled, a target label corresponding to each of the sample data to be labeled is generated according to the sample data to be labeled and the plurality of candidate sample data.

[0312] In a possible implementation, the determining of the plurality of sample quality assessment scores corresponding to the plurality of initial sample data includes:

[0313] For each of the plurality of initial sample data, determining an average similarity score between the initial sample data and the plurality of initial sample data;

[0314] Determine the influence score of the initial sample data in the labeled sample data set;

[0315] Determine the information increment score of the initial sample data in the labeled sample data set;

[0316] The sample quality assessment score corresponding to the above-mentioned initial sample data is determined according to the above-mentioned average similarity score, the above-mentioned influence score and the above-mentioned information increment score, and then multiple sample quality assessment scores corresponding to the above-mentioned multiple initial sample data are obtained.

[0317] In a possible implementation, the determining of the average similarity score between the initial sample data and the plurality of initial sample data includes:

[0318] Determine the semantic vector corresponding to the initial sample data;

[0319] Generating a plurality of first similarity scores between the initial sample data and the plurality of initial sample data according to the semantic vector corresponding to the initial sample data;

[0320] An average similarity score corresponding to the initial sample data is determined according to the multiple first similarity scores.

[0321] In a possible implementation, determining the influence score of the initial sample data in the labeled sample data set includes:

[0322] Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining the first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0323] The current initial sample data is removed from the labeled sample data set to obtain a labeled sample data subset;

[0324] Predicting the labels of the labeled sample data subset by the large language model to obtain the second label prediction probability of the true labels of all the initial sample data in the labeled sample data subset output by the large language model;

[0325] The influence score of the initial sample data in the labeled sample data set is determined according to the first label prediction probability and the second label prediction probability.

[0326] In a possible implementation, the determining of the information increment score of the initial sample data in the labeled sample data set includes:

[0327] Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining the first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model;

[0328] Determine the self entropy data of the initial sample data based on the first label prediction probability;

[0329] Determine a plurality of cross entropy data between the initial sample data and the plurality of initial sample data according to the first label prediction probability;

[0330] The information increment score corresponding to the initial sample data is determined according to the self-entropy data and the multiple cross-entropy data.

[0331] In a possible implementation, determining the sample quality assessment score corresponding to the initial sample data according to the average similarity score, the influence score, and the information increment score includes:

[0332] Obtaining a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score;

[0333] The sample quality assessment score corresponding to the initial sample data is determined according to the first weight, the second weight, the third weight, the average similarity score, the influence score and the information increment score.

[0334] In a possible implementation, the initial sample data and the sample quality assessment scores correspond one to one, and the multiple initial sample data are screened based on the multiple sample quality assessment scores to obtain multiple candidate sample data, including:

[0335] Sorting the multiple initial sample data according to the multiple sample quality assessment scores to obtain a sorting result;

[0336] A first preset number of initial sample data with the highest sample quality evaluation scores in the above-mentioned sorting results are determined as a plurality of candidate sample data.

[0337] In a possible implementation, the generating of target labels corresponding to the sample data to be labeled according to the sample data to be labeled and the plurality of candidate sample data includes:

[0338] Determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data;

[0339] Determine a plurality of vocabulary coverage scores between the aforementioned sample data to be labeled and the aforementioned plurality of candidate sample data;

[0340] Determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores;

[0341] Determining target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores;

[0342] The true label corresponding to the above target sample data is determined as the target label corresponding to the above sample data to be labeled.

[0343] In a possible implementation, the determining of a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data includes:

[0344] Extracting the first vocabulary set included in the above-mentioned sample data to be annotated by using a preset part-of-speech recognition tool;

[0345] Extracting a second vocabulary set included in each candidate sample data in the plurality of candidate sample data by using the preset part-of-speech recognition tool;

[0346] According to the first vocabulary set and the second vocabulary set included in the candidate sample data, a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data are determined.

[0347] In a possible implementation, the method further includes:

[0348] Generate prompt information corresponding to the sample data to be labeled according to the target labels corresponding to the sample data to be labeled;

[0349] The preset large language model is trained based on the above sample data to be labeled and the above prompt information.

[0350] The embodiment of the present application also provides a computer-readable storage medium, which stores instructions, and when the instructions are executed on a computer or a processor, the computer or the processor executes one or more steps in the above embodiment. If the components of the above label generation device are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium.

[0351] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The above-mentioned computer program product includes one or more computer instructions. When the above-mentioned computer program instructions are loaded and executed on a computer, the above-mentioned process or function according to the embodiment of the present application is generated in whole or in part. The above-mentioned computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium or transmitted by the above-mentioned computer-readable storage medium. The above-mentioned computer instructions can be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The above-mentioned computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, a data center, etc. that contains one or more available media integrated. The above-mentioned available media can be magnetic media (for example, floppy disks, hard disks, tapes), optical media (for example, digital versatile discs (DVD)), or semiconductor media (for example, solid state disks (SSD)), etc.

[0352] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes. In the absence of conflict, the technical features in this embodiment and the implementation scheme can be combined arbitrarily.

[0353] The above-mentioned embodiments are merely preferred embodiments of the present application and are not intended to limit the scope of the present application. Without departing from the design spirit of the present application, various modifications and improvements made to the technical solutions of the present application by ordinary technicians in this field should fall within the protection scope determined by the claims of the present application.

Claims

1. A label generation method, characterized in that: include: Obtain a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled; Determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data; Based on the multiple sample quality assessment scores, the multiple initial sample data are screened to obtain multiple candidate sample data; For each sample data to be labeled among the plurality of sample data to be labeled, a target label corresponding to each sample data to be labeled is generated according to the sample data to be labeled and the plurality of candidate sample data.

2. The method according to claim 1, characterized in that The determining of a plurality of sample quality assessment scores corresponding to the plurality of initial sample data comprises: For each initial sample data in the plurality of initial sample data, determining an average similarity score between the initial sample data and the plurality of initial sample data; Determine the influence score of the initial sample data in the labeled sample data set; Determine an information increment score of the initial sample data in the labeled sample data set; The sample quality assessment score corresponding to the initial sample data is determined according to the average similarity score, the influence score and the information increment score, thereby obtaining multiple sample quality assessment scores corresponding to the multiple initial sample data.

3. The method according to claim 2, characterized in that The determining the influence score of the initial sample data in the labeled sample data set includes: Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model; Eliminate the current initial sample data from the labeled sample data set to obtain a labeled sample data subset; Predicting the labels of the labeled sample data subset by the large language model to obtain a second label prediction probability of the true labels of all initial sample data in the labeled sample data subset output by the large language model; An influence score of the initial sample data in the labeled sample data set is determined according to the first label prediction probability and the second label prediction probability.

4. The method according to claim 2, characterized in that The determining of the information increment score of the initial sample data in the labeled sample data set includes: Predicting the labels of the labeled sample data set by using a preset large language model, and obtaining a first label prediction probability of the true labels of all initial sample data in the labeled sample data set output by the large language model; Determine the self entropy data of the initial sample data based on the first label prediction probability; Determine a plurality of cross entropy data between the initial sample data and the plurality of initial sample data according to the first label prediction probability; An information increment score corresponding to the initial sample data is determined according to the self-entropy data and the multiple cross-entropy data.

5. The method according to claim 2, characterized in that The determining, according to the average similarity score, the influence score and the information increment score, a sample quality assessment score corresponding to the initial sample data includes: Obtaining a first weight corresponding to the average similarity score, a second weight corresponding to the influence score, and a third weight corresponding to the information increment score; A sample quality assessment score corresponding to the initial sample data is determined according to the first weight, the second weight, the third weight, the average similarity score, the influence score and the information increment score.

6. The method according to claim 1, characterized in that The generating a target label corresponding to each of the sample data to be labeled according to each of the sample data to be labeled and the plurality of candidate sample data includes: Determine a plurality of second similarity scores between each of the sample data to be labeled and the plurality of candidate sample data; Determining a plurality of vocabulary coverage scores between each of the sample data to be labeled and the plurality of candidate sample data; Determine a plurality of modified similarity scores corresponding to the plurality of candidate sample data according to the plurality of second similarity scores and the plurality of vocabulary coverage scores; Determining target sample data from the plurality of candidate sample data according to the plurality of modified similarity scores; The true label corresponding to the target sample data is determined as the target label corresponding to the sample data to be labeled.

7. The method according to claim 6, characterized in that The determining of a plurality of vocabulary coverage scores between each of the sample data to be labeled and the plurality of candidate sample data comprises: Extracting a first vocabulary set included in each sample data to be annotated by using a preset part-of-speech recognition tool; Extracting a second vocabulary set included in each candidate sample data of the plurality of candidate sample data by using the preset part-of-speech recognition tool; According to the first vocabulary set and the second vocabulary set included in the candidate sample data, a plurality of vocabulary coverage scores between the sample data to be labeled and the plurality of candidate sample data are determined.

8. A label generating device, characterized in that: include: An acquisition module is used to acquire a labeled sample data set and an unlabeled sample data set, wherein the labeled sample data set includes a plurality of initial sample data and a true label of each initial sample data, and the unlabeled sample data set includes a plurality of sample data to be labeled; A determination module, configured to determine a plurality of sample quality assessment scores corresponding to the plurality of initial sample data; A screening module, configured to screen the plurality of initial sample data based on the plurality of sample quality assessment scores to obtain a plurality of candidate sample data; The first generating module is used to generate, for each sample data to be labeled among the plurality of sample data to be labeled, a target label corresponding to each sample data to be labeled according to each sample data to be labeled and the plurality of candidate sample data.

9. An electronic device, characterized in that: include: Processor and memory; The memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sample label generation method, model training method, device and equipment

    CN116541702A

  • Sample label classification method and system and electronic equipment

    CN116580254A

  • Data labeling method and device, electronic equipment and storage medium

    CN117556104A

  • Data screening method, data screening device, storage medium and electronic equipment

    CN117786170A

  • Automatic method and system for data labeling quality optimization based on coal mine

    CN117786411A