Large Model Data Annotation Method, Device, Equipment, Medium and Product
By performing grammatical transformation of the input text and prompt words and inputting them into multiple pretrained language models for annotation, weighted confidence scores are calculated to determine the annotation result, the problem of low reliability of large language models in data annotation tasks is solved, and more accurate and reliable data annotation is achieved.
Patent Information
- Application Number
- CN202510332571.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Large language models are less reliable in data annotation tasks, and may give inaccurate annotations when faced with vague, ambiguity, or rare textual expressions.
By performing grammatical transformation of the input text and prompt words for classification tasks, a diverse large model input is generated, and the label consistency of the large model is examined from different angles. Enter multiple target input text and target prompt words into multiple pretrained language models for text annotation, and obtain multiple prediction tags output from multiple pretrained language models. Based on multiple prediction tags, the weighted confidence scores of target input text, and the weights of the pre-trained language model, the weights of multiple tags are calculated, and the target annotation results are determined based on the confidence score and probability threshold.
By generating diverse input text and labeling results of multiple pre-trained language models, the labeling consistency and understanding ability of the large model can be more accurately evaluated, the reliability of data labeling can be improved, and users' trust in the labeling results of the pre-trained language model.
Smart Images

Figure CN119848555B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and particularly to a large model data annotation method, device, equipment, medium and product. Background Art
[0002] In the field of natural language processing, large language models are widely used in classification tasks such as sentiment analysis, fake news detection, bias recognition, medical diagnosis, and legal document inference. High-quality data annotation is a key prerequisite for training accurate and reliable large models.
[0003] In related technologies, using a large language model for data annotation can theoretically improve the annotation efficiency and reduce the labor cost. However, in some fields, its reliability is relatively low. For example, when encountering fuzzy, ambiguous, or rare text expressions, the large language model may give inaccurate annotations.
[0004] Therefore, how to improve the reliability of large language models in data annotation tasks is an urgent problem to be solved currently. Summary of the Invention
[0005] This application provides a large model data annotation method, device, equipment, medium and product to at least solve the problem of relatively low reliability of pre-trained language models in data annotation tasks in related technologies.
[0006] This application provides a large model data annotation method, and the method includes:
[0007] Perform syntactic transformations on the input text and the prompt words of the classification task respectively to obtain multiple target input texts and target prompt words after syntactic transformation;
[0008] Input the multiple target input texts and target prompt words into multiple pre-trained language models for text annotation to obtain multiple prediction labels output by the multiple pre-trained language models;
[0009] Obtain the weights of multiple target input texts, the weights of multiple pre-trained language models, and the probability thresholds of multiple labels;
[0010] Based on the multiple prediction labels of the classification task, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models, calculate the weighted confidence scores of multiple labels;
[0011] Determine the target annotation result of the classification task according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels.
[0012] This application also provides a large model data annotation device, and the device includes:
[0013] A transformation module, configured to perform grammar transformations on the input text and the prompt words of the task to be classified respectively, and obtain multiple target input texts and target prompt words after grammar transformation;
[0014] A labeling module, configured to input the multiple target input texts and target prompt words into multiple pre-trained language models for text labeling, and obtain multiple predicted labels output by the multiple pre-trained language models;
[0015] An acquisition module, configured to acquire the weights of the multiple target input texts, the weights of the multiple pre-trained language models, and the probability thresholds of the multiple labels;
[0016] A calculation module, configured to calculate the weighted confidence scores of the multiple labels based on the multiple predicted labels of the task to be classified, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models;
[0017] A determination module, configured to determine the target annotation result of the task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels.
[0018] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above large model data annotation methods when executing the computer program.
[0019] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above large model data annotation methods are implemented.
[0020] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above large model data annotation methods are implemented.
[0021] Through this application, the input text and prompt words of the task to be classified are respectively subjected to grammatical transformations to obtain multiple target input texts and target prompt words after the grammatical transformations. By performing grammatical transformations on the input text and prompt words of the task to be classified, diverse large model inputs are generated, and the annotation consistency of the large model is examined from different perspectives, thereby more accurately evaluating the evaluation results of the model. The multiple target input texts and target prompt words are input into multiple pre-trained language models for text annotation to obtain multiple predicted labels output by the multiple pre-trained language models. By having multiple pre-trained language models annotate multiple input versions, the understanding and annotation capabilities of the large model for the input can be evaluated more comprehensively. The weights of the multiple target input texts, the weights of the multiple pre-trained language models, and the probability thresholds of the multiple labels are obtained. Based on the multiple predicted labels of the task to be classified, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models, the weighted confidence scores of the multiple labels are calculated. According to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels, the target annotation result of the task to be classified is determined. By obtaining the confidence score of each label, the credibility of different large models for different annotation results can be reflected. This way of quantifying indicators makes the decision-making process of the model more transparent, thereby improving the reliability of the pre-trained language model in the data annotation task and enhancing users' trust in the annotation results of the pre-trained language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0023] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of a large model data annotation method provided by an embodiment of the present application;
[0025] Figure 2 It is a schematic structural diagram of a large model data annotation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0027] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0028] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0029] Term Explanation:
[0030] Large Model: An artificial intelligence large model refers to a "large parameter" model trained using large-scale data and powerful computing capabilities. These models usually have high generality and generalization capabilities and can be applied to fields such as natural language processing, image recognition, speech recognition, etc. They can be classified into large language models, visual large models, multimodal large models, basic large models, etc.
[0031] Data Annotation: By processing, classifying, labeling, and annotating raw data (such as images, texts, audios, videos, etc.), it can be made understandable and usable by machine learning models. Among them, data annotation by a large model refers to the process of automatically annotating unannotated data using large-scale pre-trained language models (such as GPT, BERT, etc.). Data annotation by a large model utilizes its powerful understanding and generation capabilities to automatically add labels to data, improving efficiency, but attention should be paid to accuracy and domain adaptability.
[0032] Syntax Transformation: Syntax transformation refers to in natural language processing (NLP), by changing the expression of a sentence, generating new sentences or phrases while keeping the original meaning unchanged, and applying it to scenarios such as data augmentation and model robustness improvement. However, attention should be paid to semantic consistency and language fluency.
[0033] Pre-trained language model: Generally refers to a large-scale neural network algorithm structure and its parameters obtained by designing a language model training task based on a large corpus (including language training materials such as sentences, paragraphs, etc.) and training a large-scale neural network algorithm structure. For subsequent other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. Among them, the neural network algorithm structure for pre-training the language model can be CNN, RNN, LSTM, etc., or a model constructed by an attention network, such as transformer, BERT, GPT, Clip, etc. This application does not make any limitations here. An attention network refers to a network model trained using the attention mechanism. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence and enabling the model to obtain a more accurate output in the end. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain models adapted to other tasks. By pre-training on a large corpus, the neural language representation model can learn powerful language representation capabilities and extract rich syntactic and semantic information from the text. In the embodiments of the present disclosure, the pre-trained language model may include, but is not limited to, GPT, BERT, etc.
[0034] Since the reliability of large language models is relatively low when performing data annotation in some fields, when encountering fuzzy, ambiguous or rare text expressions, the large language model may give inaccurate annotations. For example, in medical diagnosis assistance, incorrect annotations may mislead doctors' judgments and pose a threat to patients' health; in legal text analysis, inaccurate annotations may affect judicial fairness.
[0035] In related technologies, indicators such as accuracy, confidence, and majority voting are usually used to evaluate the annotation results of large language models. The accuracy indicator evaluates the model performance by calculating the proportion of the number of correctly annotated samples to the total number of samples. For example, in a test set containing 100 text samples, if the model correctly annotates 80 samples, then the accuracy of the model is 80%. The confidence evaluation method directly asks the model about the uncertainty of its own output. Some large language models provide relevant interfaces that can return the confidence score of their own annotation results. The higher the score, the more confident the model is in this result. The majority voting method is to have multiple large language models annotate the same input, and then select the annotation result with the highest frequency as the final output. For example, in a sentiment analysis task, multiple large language models respectively perform sentiment classification on a piece of text. If most models consider the text to have a positive sentiment, then the final annotation is positive.
[0036] However, the above methods all have defects to varying degrees. For example, the accuracy metric is too simple to reflect the performance of the model in complex scenarios. It does not take into account the model's ability to label samples of different difficulties, as well as the types and distributions of labeling errors. The self-confidence evaluation method relies on the model's own judgment of uncertainty. However, large models are essentially still neural network models, and the confidence evaluation mechanism inside the model is not always accurate and reliable. The confidence scores it gives may not match the actual reliability of the labels. The majority voting method does not consider the confidence of individual models and the performance differences of different models in different scenarios. In practical applications, different large language models may perform very differently when dealing with specific domains or specific types of text. Simple majority voting may obscure the advantages of certain models in specific scenarios, resulting in a decrease in the overall labeling accuracy. All in all, there is still a lack of an evaluation method that can further improve the labeling quality of large models.
[0037] Based on the above problems, embodiments of the present application provide a large model data annotation method, and a detailed description of the method is given in combination with the execution process of the large model data annotation method.
[0038] Refer to Figure 1 As shown, the large model data annotation method provided by the embodiments of the present invention includes the following steps:
[0039] S11. Perform syntactic transformations on the input text and the prompt words of the classification task respectively to obtain multiple target input texts and target prompt words after the syntactic transformation.
[0040] Among them, the input text refers to the original text content provided by the user or the system to the pre-trained language model for a certain processing or analysis. It is usually the input data of the model and can be any type of natural language text, such as a sentence, a paragraph, or an entire document. The prompt word is used to guide the pre-trained language model to generate a specific output or perform a specific task. By providing the form of context or a question, it helps the pre-trained language model understand the user's needs.
[0041] Specifically, since syntactic transformation reorganizes the structure of a sentence or text while keeping its semantics unchanged or basically unchanged, multiple target input texts and target prompt words after the syntactic transformation are obtained by performing syntactic transformations on the input text and the prompt words of the classification task respectively.
[0042] Exemplarily, the input text of the classification task is: "Xiaoming is very happy today.", and the prompt word is "Please help me judge whether the above text expresses positive emotions?", and multiple input texts and multiple prompt words after the syntactic transformation are obtained through different semantic transformation methods.
[0043] In some embodiments, the above step S11 (determining the target annotation result of the task to be classified according to the weighted confidence scores of the multiple tags and the probability thresholds of the multiple tags) may be implemented by the following steps:
[0044] (1) Perform syntactic transformation on the input text based on a first preset method to obtain multiple target input texts after syntactic transformation.
[0045] Optionally, the above step (1) may be implemented by the following method:
[0046] Perform voice transformation on the input text to obtain a first input text after syntactic transformation;
[0047] Perform sentence pattern transformation on the input text to obtain a second input text after syntactic transformation;
[0048] Perform synonym replacement on the input text to obtain a third input text after syntactic transformation.
[0049] Among them, voice transformation includes switching the voice of the input text between active voice and passive voice. Sentence pattern transformation includes converting the affirmative phrase in the input text into two negative phrases while keeping the original meaning unchanged. Synonym replacement means replacing the keywords in the input text with synonyms to generate different expressions with the same semantics.
[0050] Specifically, obtain the first input text by converting the active voice of the input text into passive voice, or converting the passive voice of the input text into active voice; obtain the second input text by converting the affirmative phrase of the input text into two negative phrases, or converting the double negative phrase of the input text into an affirmative phrase; obtain the third input text by replacing the keywords in the input text with synonyms.
[0051] Exemplarily, "Xiaoming wrote a letter. He sent the letter to his friend." can be converted to "A letter was written by Xiaoming. This letter was sent to his friend." This kind of conversion can test whether the pre-trained language model can maintain the correct understanding of semantics under the change of sentence structure.
[0052] Exemplarily, "He will definitely come to the party." is converted to "It is impossible for him not to come to the party." This kind of syntactic transformation helps to test the sensitivity of the model to semantic changes.
[0053] Exemplarily, "She likes reading novels very much, especially those works full of suspense and thrilling plots." becomes "She especially loves reading stories, especially those works full of suspense and exciting plots." In this way, it is examined whether the model pays attention to semantic content rather than surface lexical differences.
[0054] Further, the above step (1) can also be implemented in the following way:
[0055] Adjust the word order of the input text to obtain the fourth input text after grammatical transformation;
[0056] Expand or simplify the input text to obtain the fifth input text after grammatical transformation.
[0057] Specifically, by changing the order of words or phrases in the input text, sentences with similar semantics but different grammatical structures are generated. The complexity of the sentence is changed by adding or deleting some modifying components while maintaining the core semantics.
[0058] Through different grammatical transformations, diverse input texts can be generated, or multiple question forms can be generated to help the pre-trained language model better understand the user's question.
[0059] (2) Perform grammatical transformation on the prompt word based on the second preset method to obtain the target prompt word after grammatical transformation.
[0060] Optionally, the above step (2) can be implemented in the following way:
[0061] Perform sentence pattern transformation on the prompt word to obtain the target prompt word after grammatical transformation.
[0062] Specifically, by performing double-negation conversion on the affirmative phrase in the prompt word, the target prompt word after grammatical transformation can be obtained.
[0063] Exemplarily, "Please help me determine whether the following text expresses positive emotions?" can be converted to "Please help me determine whether the following text is not expressing negative emotions". This kind of grammatical transformation helps to test the model's sensitivity to semantic transformation of reasoning prompts.
[0064] By performing grammatical transformation on the input text and prompt word of the classification task to generate diverse large model inputs, the annotation consistency of the large model is examined from different perspectives, so as to more accurately evaluate the evaluation results of the model.
[0065] S12. Input the multiple target input texts and the target prompt word into multiple pre-trained language models for text annotation to obtain multiple prediction labels output by the multiple pre-trained language models.
[0066] Specifically, input the original input text and the original prompt word, as well as the multiple target input texts and the target prompt word after grammatical transformation into multiple pre-trained language models for annotation to obtain multiple prediction labels output by the multiple pre-trained language models.
[0067] Exemplarily, taking the sentiment classification task as an example, assume that the possible annotation results (predicted labels) include three categories: positive emotion, negative emotion, and neutral emotion. If among the annotation results of a pre-trained language model for multiple inputs (taking 7 inputs as an example), 5 are positive emotions and 2 are negative emotions, then the probability distribution of the label type being positive emotion is 5 / 7, and the probability distribution of the label type being negative emotion is 2 / 7.
[0068] S13. Obtain the weights of multiple target input texts, the weights of multiple pre-trained language models, and the probability thresholds of multiple labels.
[0069] Specifically, in the annotation quality assessment process, three types of parameters are involved, namely Iweight, Mweight, and kL. Among them, Iweight represents the weight assigned to a given input text, Mweight represents the weight assigned to a given pre-trained language model M, and kL represents the probability threshold for each label. For a specific task, users with historical experience can set the weights of the three types of parameters according to experience. However, generally, since the subjective determination of weights by humans will inevitably be affected by personal biases, limited experience, and cognitive limitations. Different people may give very different weights, and these weights are not necessarily based on the true characteristics of the data and the requirements of the model. Moreover, different classification tasks and datasets have unique characteristics, and different weight configurations are required to achieve the best performance. In addition, there are complex interactions between the different weight factors of the three types of parameters. For example, the weight coefficients of different syntactic transformations have different effects on different large models, and different large models also perform differently on different tasks. Therefore, a method is needed to determine that the weight values of the above three parameters are within a reasonable range.
[0070] For classification tasks in most scenarios, users have a certain number of training samples with correct annotation results. Based on this specific task and this specific training sample, the optimal values of the above three types of weight coefficients can be determined through heuristic learning methods.
[0071] Optionally, the above step S13 can be implemented in the following manner:
[0072] Initialize the candidate solution population.
[0073] Among them, the candidate solution population includes the initial weights of multiple target input texts, the initial weights of multiple pre-trained language models, and the initial probability thresholds of multiple labels;
[0074] Evaluate the fitness of each candidate solution based on the accuracy function;
[0075] Perform crossover operations and mutation operations based on the fitness of each candidate solution to generate a new candidate solution population;
[0076] Through multiple iterations, gradually optimize the candidate solution population until the preset number of iterations is reached, and determine the weights of multiple target input texts, the weights of multiple pre-trained language models, and the probability thresholds of multiple labels.
[0077] Specifically, initialize the candidate solution population. By randomly assigning values multiple times, a candidate solution population containing Iweight, Mweight, and kL can be generated, and the ranges of these values are all between (0, 1). Then, use the accuracy function to evaluate the fitness of each candidate solution, and carry out the iteration of the population mutation and crossover process based on the distribution of the initial population fitness. During crossover, generate new candidate solutions by combining and randomly exchanging partial values of the candidate solutions; during mutation, generate new candidate solutions by randomly changing partial values of the candidate solutions. If the optimal fitness of the new candidate solution is better than the optimal fitness of the current population, replace the current optimal fitness. This process continues for several iterations or until convergence, and finally return the optimized parameters. During the optimization process, the parameters need to meet specific constraint conditions to ensure that the weights of the pre-trained language models, the weights of the input texts, and the probability thresholds are within a reasonable range.
[0078] In addition, for large data samples, after obtaining the weights of multiple target input texts, the weights of multiple pre-trained language models, and the probability thresholds of multiple labels, the remaining samples can be labeled by combining these weight coefficients and probability thresholds.
[0079] Among them, the accuracy function is expressed as: . This accuracy function is used to maximize the accuracy by minimizing the loss function, where the loss function is expressed by the following formula:
[0080] .
[0081] Among them, represents the loss value between the predicted label and the true label, n represents the number of samples of the labeled true labels, y pred is the predicted label, and y true is the labeled true label.
[0082] Through the heuristic algorithm, the weight adjustment of the parameters is carried out based on objective data and established optimization strategies, without being interfered by subjective factors, and can provide more objective and reliable weights, thus ensuring the stability and reliability of the evaluation performance.
[0083] S14. Based on the multiple predicted labels of the to-be-classified task, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models, calculate the weighted confidence scores of the multiple labels.
[0084] In some embodiments, the above step S14 (calculating the weighted confidence scores of multiple labels based on the multiple predicted labels of the classification task to be classified, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models) can be implemented through the following steps:
[0085] A. Based on the multiple predicted labels of the classification task to be classified and the weights of the multiple target input texts, calculate the confidence score of the target pre-trained language model for the target label.
[0086] Optionally, the above step A can be implemented in the following manner:
[0087] A1. Obtain the preset weight of the target predicted label.
[0088] Obtain the target predicted label assigned by the target pre-trained language model to the target input text;
[0089] Determine whether the target predicted label is consistent with the target label;
[0090] If the target predicted label is consistent with the target label, determine that the preset weight of the target predicted label is the first value;
[0091] If the target predicted label is inconsistent with the target label, determine that the preset weight of the target predicted label is the second value.
[0092] Wherein, the target pre-trained language model is any one of the multiple pre-trained language models. The target label is any one of the labels that may be included when the pre-trained language model annotates the target input text. For example, taking the sentiment classification task as an example, the input text is: "Xiaoming is very happy today.", and the prompt is "Please help me determine whether the above text expresses positive emotions?", in this case, the target labels may include: negative emotion, positive emotion, and neutral emotion.
[0093] Specifically, obtain the target predicted label assigned by the target pre-trained language model to the target input text, determine whether the target predicted label is consistent with the target label. If the target predicted label is consistent with the target label, determine that the preset weight of the target predicted label is 1; if the target predicted label is inconsistent with the target label, determine that the preset weight of the target predicted label is 0.
[0094] A2. Calculate the confidence score of the target pre-trained language model for the target label according to the preset weight of the target predicted label and the weights of the multiple target input texts.
[0095] Specifically, calculate the confidence score of the target pre-trained language model for the target label according to the preset weight of the target predicted label and the weights of the multiple target input texts.
[0096] Optionally, step A2 above can be calculated by the following formula:
[0097] .
[0098] Wherein, represents the confidence score of the target pre-trained language model M for the target label L; represents the weight assigned to the target input text, and its value ranges from 0 to 1, and the sum of the weights of multiple target input texts is 1; represents the target pre-trained language model M's assignment of labels to the target input text l ; represents the preset weight of the label , when is equal to L, has a value of 1; when is not equal to L, has a value of 0.
[0099] B. Calculate the weighted confidence scores of the multiple pre-trained language models for the target label based on the confidence score of the target pre-trained language model for the target label and the weights of the multiple pre-trained language models.
[0100] Optionally, step B above can be implemented in the following manner:
[0101] B1. Calculate the confidence scores of each pre-trained language model for each label respectively;
[0102] B2. Calculate the weighted confidence scores of the multiple pre-trained language models for the target label according to the weights of each pre-trained language model and the confidence scores of each pre-trained language model for each label.
[0103] Specifically, calculate the confidence scores of each pre-trained language model for each label respectively, and then calculate the weighted confidence scores of the multiple pre-trained language models for the target label according to the weights of each pre-trained language model and the confidence scores of each pre-trained language model for each label.
[0104] Optionally, step B above can be calculated by the following formula:
[0105] .
[0106] Wherein, represents the weighted confidence score of the multiple pre-trained language models (i.e., all pre-trained language models in the embodiments of the present disclosure) for the target label L; represents the sum of the weights of the multiple pre-trained language models, and this value is 1; represents the weight assigned to the target pre-trained language model M, and its value ranges from 0 to 1.
[0107] S15. Determine the target annotation result of the classification task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels.
[0108] In some embodiments, if the classification task to be classified is a single-classification task, the above step S15 (determining the target annotation result of the classification task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels) can be implemented in the following manner:
[0109] Determine whether the weighted confidence score of each label is greater than the probability threshold of each label;
[0110] If the weighted confidence score of the target label among the multiple labels is greater than the probability threshold of the target label, determine that the target label is the target annotation result of the classification task to be classified.
[0111] Among them, the probability threshold corresponding to each label is different.
[0112] Specifically, judge the weighted confidence score of each label according to the set probability threshold to determine the final annotation label. That is, determine whether the weighted confidence score of each label is greater than the probability threshold of each label. If one of the multiple labels meets the corresponding threshold condition, determine that the target label is the target annotation result of the classification task to be classified.
[0113] If the weighted confidence scores of at least two labels among the multiple labels are greater than the probability thresholds of at least two labels, calculate the relative scores of the at least two labels;
[0114] Sort the relative scores of the at least two labels, and determine that the label with the highest relative score among the at least two labels is the target annotation result of the classification task to be classified.
[0115] Specifically, if at least two labels among the multiple labels meet the corresponding threshold conditions, calculate the relative scores of the at least two labels according to the following formula:
[0116] .
[0117] Among them, represents the relative score of any label L; represents the probability threshold of any label L.
[0118] In some embodiments, if the classification task to be classified is a multi-classification task, the above step S15 (determining the target annotation result of the classification task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels) can be implemented in the following manner:
[0119] Determine whether the weighted confidence scores of each label are greater than the probability thresholds of each label;
[0120] If the weighted confidence scores of at least two labels among multiple labels are greater than the probability thresholds of at least two labels, calculate the relative scores of the at least two labels;
[0121] Sort the relative scores of the at least two labels in descending order, and determine the top N labels with the highest relative scores among the at least two labels as the target annotation results of the classification task to be classified.
[0122] Wherein, N is greater than or equal to 2.
[0123] Specifically, for a multi-classification task, if multiple labels meet the threshold condition, the labels corresponding to multiple classification tasks are determined according to the score ranking of the relative scores.
[0124] This method uses a multi-task scenario, which is not only applicable to a single annotation task of a single pre-trained language model, but also can be carried out in multiple annotation task scenarios of multiple pre-trained language models. By reasonably allocating the weights between labels and comprehensively considering the outputs of multiple models, the weighted confidence scores of multiple labels are calculated to determine the final annotation label, which can fully utilize the evaluation index (i.e., the probability threshold) to improve the accuracy of the annotation result.
[0125] The large model data annotation method provided by the embodiments of the present disclosure performs syntactic transformations on the input text and the prompt words of the classification task to be classified respectively, and obtains multiple target input texts and target prompt words after the syntactic transformation. By performing syntactic transformations on the input text and the prompt words of the classification task to be classified, diverse large model inputs are generated, and the annotation consistency of the large model is investigated from different perspectives, so as to more accurately evaluate the evaluation result of the model. Input the multiple target input texts and target prompt words into multiple pre-trained language models for text annotation, and obtain multiple predicted labels output by the multiple pre-trained language models. By annotating multiple input versions with multiple pre-trained language models, the understanding and annotation capabilities of the large model for the input can be more comprehensively evaluated. Obtain the weights of multiple target input texts, the weights of multiple pre-trained language models, and the probability thresholds of multiple labels. Based on the multiple predicted labels of the classification task to be classified, the weights of multiple target input texts, and the weights of multiple pre-trained language models, calculate the weighted confidence scores of multiple labels. According to the weighted confidence scores of multiple labels and the probability thresholds of multiple labels, determine the target annotation results of the classification task to be classified. By obtaining the confidence score of each label, the credibility of different large models for different annotation results can be reflected. This way of quantifying indicators makes the decision-making process of the model more transparent, thereby improving the reliability of the pre-trained language model in the data annotation task and enhancing the user's trust in the annotation results of the pre-trained language model.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0127] An embodiment of the present application also provides a large model data annotation device. The content of the omitted virtual device claims should be described in detail in the specification and corresponding to the method claims one by one.
[0128] Figure 2 FIG. is a schematic structural diagram of a large model data annotation device 200 provided by the present disclosure, as Figure 2 shown, the device of this embodiment includes: a transformation module 210, an annotation module 220, an acquisition module 230, a calculation module 240, and a determination module 250, wherein,
[0129] The transformation module 210 is configured to perform syntax transformation on the input text and the prompt words of the classification task respectively, and obtain multiple target input texts and target prompt words after syntax transformation;
[0130] The annotation module 220 is configured to input the multiple target input texts and target prompt words into multiple pre-trained language models for text annotation, and obtain multiple prediction labels output by the multiple pre-trained language models;
[0131] The acquisition module 230 is configured to acquire the weights of the multiple target input texts, the weights of the multiple pre-trained language models, and the probability thresholds of the multiple labels;
[0132] The calculation module 240 is configured to calculate the weighted confidence scores of the multiple labels based on the multiple prediction labels of the classification task, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models;
[0133] The determination module 250 is configured to determine the target annotation result of the classification task according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels.
[0134] As an optional implementation manner of an embodiment of the present disclosure, the transformation module 210 includes:
[0135] The first syntax transformation unit is configured to perform syntax transformation on the input text based on a first preset method to obtain multiple target input texts after syntax transformation;
[0136] The second syntax transformation unit is configured to perform syntax transformation on the prompt word based on a second preset method to obtain the target prompt word after syntax transformation.
[0137] As an optional implementation manner of an embodiment of the present disclosure, the first syntax transformation unit is specifically configured to:
[0138] Perform voice transformation on the input text to obtain the first input text after grammar transformation;
[0139] Perform sentence pattern transformation on the input text to obtain the second input text after grammar transformation;
[0140] Perform synonym replacement on the input text to obtain the third input text after grammar transformation.
[0141] As an optional implementation manner of the embodiment of the present disclosure, the second grammar transformation unit is specifically configured to:
[0142] Perform sentence pattern transformation on the prompt word to obtain the target prompt word after grammar transformation.
[0143] As an optional implementation manner of the embodiment of the present disclosure, the computing module 240 includes:
[0144] A confidence score calculation unit, configured to calculate the confidence score of the target pre-trained language model for the target label based on multiple predicted labels of the task to be classified and the weights of the multiple target input texts;
[0145] A weighted confidence score calculation unit, configured to calculate the weighted confidence score of the multiple pre-trained language models for the target label based on the confidence score of the target pre-trained language model for the target label and the weights of the multiple pre-trained language models.
[0146] As an optional implementation manner of the embodiment of the present disclosure, the confidence score calculation unit includes:
[0147] An acquisition subunit, configured to acquire the preset weight of the target predicted label;
[0148] A calculation subunit, configured to calculate the confidence score of the target pre-trained language model for the target label according to the preset weight of the target predicted label and the weights of the multiple target input texts.
[0149] As an optional implementation manner of the embodiment of the present disclosure, the acquisition subunit is specifically configured to:
[0150] Acquire the target predicted label assigned by the target pre-trained language model to the target input text;
[0151] Determine whether the target predicted label is consistent with the target label;
[0152] If the target predicted label is consistent with the target label, determine that the preset weight of the target predicted label is the first value;
[0153] If the target prediction label is inconsistent with the target label, determine that the preset weight of the target prediction label is a second value.
[0154] As an optional implementation manner of the embodiments of the present disclosure, the weighted confidence score calculation unit is specifically configured to:
[0155] Calculate the confidence scores of each pre-trained language model for each label respectively;
[0156] Calculate the weighted confidence scores of multiple pre-trained language models for the target label according to the weights of each pre-trained language model and the confidence scores of each pre-trained language model for each label.
[0157] As an optional implementation manner of the embodiments of the present disclosure, if the classification task to be classified is a single classification task, the determination module 250 includes:
[0158] A judgment unit, configured to judge whether the weighted confidence scores of each label are greater than the probability thresholds of each label;
[0159] A comparison unit, configured to determine the target label as the target annotation result of the classification task to be classified if the weighted confidence score of the target label among multiple labels is greater than the probability threshold of the target label.
[0160] As an optional implementation manner of the embodiments of the present disclosure, the judgment unit is further specifically configured to:
[0161] If the weighted confidence scores of at least two labels among multiple labels are greater than the probability thresholds of at least two labels, calculate the relative scores of the at least two labels;
[0162] Sort the relative scores of the at least two labels, and determine the label with the highest relative score among the at least two labels as the target annotation result of the classification task to be classified.
[0163] For the description of the features in the corresponding embodiments of the large model data annotation device 200, reference may be made to the relevant descriptions in the corresponding embodiments of the large model data annotation method, which will not be elaborated here one by one.
[0164] The large model data annotation device provided by the embodiments of the present disclosure performs grammatical transformations on the input text and the prompt words of the classification task respectively to obtain multiple target input texts and target prompt words after grammatical transformation. By performing grammatical transformations on the input text and the prompt words of the classification task, diverse large model inputs are generated, and the annotation consistency of the large model is examined from different perspectives, so as to more accurately evaluate the evaluation results of the model. The multiple target input texts and target prompt words are input into multiple pre-trained language models for text annotation to obtain multiple predicted labels output by the multiple pre-trained language models. By annotating multiple input versions through multiple pre-trained language models, the understanding and annotation capabilities of the large model for the input can be more comprehensively evaluated. The weights of the multiple target input texts, the weights of the multiple pre-trained language models, and the probability thresholds of the multiple labels are obtained. Based on the multiple predicted labels of the classification task, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models, the weighted confidence scores of the multiple labels are calculated. According to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels, the target annotation result of the classification task is determined. By obtaining the confidence score of each label, the credibility of different large models for different annotation results can be reflected. This way of quantifying indicators makes the decision-making process of the model more transparent, thereby improving the reliability of the pre-trained language model in the data annotation task and enhancing the user's trust in the annotation results of the pre-trained language model.
[0165] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the large model data annotation method.
[0166] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the large model data annotation method when running.
[0167] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0168] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the large model data annotation method.
[0169] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the large model data annotation method.
[0170] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0171] The above has introduced in detail a large model data annotation method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A large model data annotation method, characterized in that: The method comprises: Performing grammatical transformation on the input text and prompt words of the classification task respectively, and obtaining a plurality of target input texts and target prompt words after the grammatical transformation; Inputting the multiple target input texts and target prompt words into multiple pre-trained language models for text annotation, and obtaining multiple predicted labels output by the multiple pre-trained language models; Obtain weights of multiple target input texts, weights of multiple pre-trained language models, and probability thresholds of multiple labels; Calculating weighted confidence scores of the multiple labels based on the multiple predicted labels of the task to be classified, the weights of the multiple target input texts, and the weights of the multiple pre-trained language models; Determining a target labeling result of the task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels; The step of obtaining weights of multiple target input texts, weights of multiple pre-trained language models, and probability thresholds of multiple labels includes: Initializing a candidate solution population; the candidate solution population includes initial weights of multiple target input texts, initial weights of multiple pre-trained language models, and initial probability thresholds of multiple labels; Evaluate the fitness of each candidate solution based on the accuracy function; Perform crossover and mutation operations based on the fitness of each candidate solution to generate a new candidate solution population; Through multiple iterations, the candidate solution population is gradually optimized until a preset number of iterations is reached, and weights of multiple target input texts, weights of multiple pre-trained language models, and probability thresholds of multiple labels are determined.
2. The large model data annotation method according to claim 1, characterized in that: The step of performing grammatical transformation on the input text and prompt words of the classification task to be processed, and obtaining a plurality of target input texts and target prompt words after the grammatical transformation, comprises: Performing grammatical transformation on the input text based on a first preset manner to obtain a plurality of target input texts after grammatical transformation; The prompt word is grammatically transformed based on a second preset manner to obtain a target prompt word after the grammatical transformation.
3. The large model data annotation method according to claim 2, characterized in that: The step of performing grammatical transformation on the input text based on a first preset manner to obtain a plurality of target input texts after grammatical transformation includes: Performing voice transformation on the input text to obtain a first input text after grammatical transformation; Performing sentence transformation on the input text to obtain a second input text after grammatical transformation; Synonym replacement is performed on the input text to obtain a third input text after grammatical transformation.
4. The large model data annotation method according to claim 2, characterized in that: The step of performing grammatical transformation on the prompt word based on the second preset method to obtain a target prompt word after grammatical transformation includes: The prompt word is subjected to sentence transformation to obtain a target prompt word after grammatical transformation.
5. The large model data annotation method according to claim 1, characterized in that: The step of calculating weighted confidence scores of the plurality of labels based on the plurality of predicted labels of the task to be classified, the weights of the plurality of target input texts, and the weights of the plurality of pre-trained language models comprises: Calculating a confidence score of a target pre-trained language model for a target label based on the multiple predicted labels of the task to be classified and the weights of the multiple target input texts; Based on the confidence score of the target pre-trained language model for the target label and the weights of the multiple pre-trained language models, weighted confidence scores of the multiple pre-trained language models for the target label are calculated.
6. The large model data annotation method according to claim 5, characterized in that: The step of calculating the confidence score of the target pre-trained language model for the target label based on the weights of the multiple predicted labels of the task to be classified and the multiple target input texts includes: Get the preset weight of the target prediction label; According to the preset weight of the target predicted label and the weights of the multiple target input texts, the confidence score of the target pre-trained language model for the target label is calculated.
7. The large model data annotation method according to claim 6, characterized in that: The step of obtaining a preset weight of a target prediction label includes: Get the target prediction label assigned by the target pre-trained language model to the target input text; Determine whether the target prediction label is consistent with the target label; If the target prediction label is consistent with the target label, determining the preset weight of the target prediction label to be a first value; If the target prediction label is inconsistent with the target label, the preset weight of the target prediction label is determined to be a second value.
8. The large model data annotation method according to claim 5, characterized in that: The calculating, based on the confidence score of the target pre-trained language model for the target label and the weights of the multiple pre-trained language models, weighted confidence scores of the multiple pre-trained language models for the target label comprises: Calculate the confidence scores of each pre-trained language model for each label respectively; According to the weights of each pre-trained language model and the confidence scores of each pre-trained language model for each label, the weighted confidence scores of multiple pre-trained language models for the target label are calculated.
9. The large model data annotation method according to claim 1, characterized in that: If the task to be classified is a single classification task, determining the target labeling result of the task to be classified according to the weighted confidence scores of the multiple labels and the probability thresholds of the multiple labels includes: Determine whether the weighted confidence score of each label is greater than the probability threshold of each label; If the weighted confidence score of the target label among the multiple labels is greater than the probability threshold of the target label, the target label is determined to be the target labeling result of the task to be classified.
10. The large model data labeling method according to claim 9, characterized in that: The step of judging whether the weighted confidence score of each predicted label is greater than the probability threshold of each label further includes: If the weighted confidence scores of at least two tags among the multiple tags are greater than the probability threshold of the at least two tags, then calculating the relative scores of the at least two tags; The relative scores of the at least two tags are sorted, and the tag with the highest relative score among the at least two tags is determined as the target labeling result of the task to be classified.
11. A large model data annotation device, characterized in that: The device comprises: A transformation module, used to perform grammatical transformation on the input text and prompt words of the classification task, and obtain multiple target input texts and target prompt words after grammatical transformation; A labeling module, used for inputting the multiple target input texts and target prompt words into multiple pre-trained language models for text labeling, and obtaining multiple predicted labels output by the multiple pre-trained language models; An acquisition module, used to obtain weights of multiple target input texts, weights of multiple pre-trained language models, and probability thresholds of multiple tags; A calculation module, configured to calculate weighted confidence scores of a plurality of labels based on a plurality of predicted labels of the task to be classified, weights of the plurality of target input texts, and weights of the plurality of pre-trained language models; A determination module, configured to determine a target labeling result of the task to be classified according to weighted confidence scores of the multiple labels and probability thresholds of the multiple labels; The acquisition module is specifically used for: Initializing a candidate solution population; the candidate solution population includes initial weights of multiple target input texts, initial weights of multiple pre-trained language models, and initial probability thresholds of multiple labels; Evaluate the fitness of each candidate solution based on the accuracy function; Perform crossover and mutation operations based on the fitness of each candidate solution to generate a new candidate solution population; Through multiple iterations, the candidate solution population is gradually optimized until a preset number of iterations is reached, and weights of multiple target input texts, weights of multiple pre-trained language models, and probability thresholds of multiple labels are determined.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the large model data labeling method as claimed in any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the large model data labeling method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the large model data labeling method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Text generation method and device, computer equipment, storage medium and program product
CN118520854A
Text rewriting cue word generation method and device, electronic equipment and storage medium
CN119476214A