Data annotation method, annotation model training and capability evaluation method and system
By training the labeling model and using confidence intervals to screen highly accurate labeling results, the problems of time-consuming and inflexible manual data labeling are solved, automated data labeling is achieved, labeling efficiency and accuracy are improved, and efficient training and evaluation of machine learning models are supported.
Patent Information
- Application Number
- CN202510851493.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
AI Technical Summary
Manual data labeling is time-consuming, inflexible, inconsistently accurate, and costly, making it difficult to meet the training and evaluation needs of machine learning models.
By training the labeling model, using the confidence interval to screen the labeling results with high accuracy, and combining the scene recognition model to determine the credibility of the labeling results, automated data labeling is achieved.
It improves the efficiency and accuracy of data labeling, reduces dependence on manual labeling, and improves the training and evaluation efficiency of machine learning models.
Smart Images

Figure CN120745875A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a data annotation method, annotation model training and capability assessment method and system. Background Art
[0002] Data labeling is a common requirement in the field of artificial intelligence. For example, the training process of machine learning models requires a large amount of training data, which often needs to be pre-labeled. Another example is that when evaluating machine learning models, the output data of the machine learning models often needs to be labeled. This labeled data can be used to judge the effectiveness of model inference or iteration.
[0003] In practical applications, data labeling is often done manually. Manual labeling faces numerous challenges. First, manual labeling is often time-consuming, often requiring several days to complete. Furthermore, manual labeling offers limited flexibility, typically confined to fixed work hours, which limits its efficiency. Individual differences between labelers can also lead to inconsistent judgment criteria, increasing the uncertainty of the labeling results. Furthermore, manual labeling comes with high training costs. For new tasks, training labelers and standardizing judgment criteria further adds an additional burden.
[0004] The content of the background technology section is merely information known to the inventor personally, and does not mean that the above information has entered the public domain before the application date of this disclosure, nor does it mean that it can become the prior art of the present disclosure. Summary of the Invention
[0005] This specification provides a data annotation method, a method and system for training and evaluating the capabilities of an annotation model. This data annotation method can automatically annotate data, thereby improving annotation efficiency, reducing reliance on manual annotation, and improving data annotation accuracy.
[0006] In the first aspect, the present specification provides a data labeling method, comprising: obtaining target data to be labeled, and determining a first item to be labeled of the target data; obtaining a trained first labeling model, and labeling capability information of the first labeling model for the first item, the labeling capability information representing a confidence interval, and when the confidence of the labeling result output by the first labeling model for the first item falls within the confidence interval, the labeling accuracy of the first labeling model is greater than or equal to the target accuracy; using the first labeling model to label the target data on the first item, and obtaining a first labeling result and a confidence corresponding to the first labeling result; and determining whether the first labeling result is a credible labeling result based on the labeling capability information and the confidence, and if so, using the first labeling result as the labeling result of the target data on the first item.
[0007] In some embodiments, obtaining the trained first labeling model and the labeling capability information of the first labeling model for the first project includes: obtaining the trained first labeling model and the labeling capability information of the first labeling model for the first project from a database, wherein the database stores labeling models corresponding to multiple preset projects and the labeling capability information of each labeling model for its corresponding preset project, the first project is one of the multiple preset projects, and the first labeling model is the labeling model corresponding to the first project.
[0008] In some embodiments, the first item corresponds to multiple candidate annotation contents, and the first annotation result is one of the multiple candidate annotation contents; the annotation capability information includes: sub-capability information corresponding to each of the multiple candidate annotation contents, wherein the sub-capability information corresponding to each candidate annotation content represents the confidence interval applicable to the candidate annotation content.
[0009] In some embodiments, the determining whether the first annotation result is a credible annotation result based on the annotation capability information and the confidence level includes: if the confidence level is within the confidence level interval applicable to the first annotation result, determining that the first annotation result is a credible annotation result; or if the confidence level is outside the confidence level interval applicable to the first annotation result, determining that the first annotation result is not a credible annotation result.
[0010] In some embodiments, the sub-capability information corresponding to each candidate annotation content includes confidence intervals corresponding to multiple scenarios, and determining whether the first annotation result is a credible annotation result based on the annotation capability information and the confidence includes: determining the target scenario to which the target data belongs, and obtaining the target confidence interval corresponding to the target scenario from the sub-capability information corresponding to the first annotation result; and determining whether the first annotation result is a credible annotation result based on the confidence and the target confidence interval.
[0011] In some embodiments, determining the target scene to which the target data belongs includes: obtaining a trained scene recognition model; and using the scene recognition model to perform scene recognition processing on the target data to obtain the target scene.
[0012] In some embodiments, obtaining the trained scene recognition model includes: obtaining the trained scene recognition model from a database, wherein the database stores annotation models corresponding to multiple preset items, and annotation capability information of each annotation model for its corresponding preset item, the first item is one of the multiple preset items, the first annotation model is the annotation model corresponding to the first item, the scene recognition model is the second annotation model corresponding to the second item, and the second item is other items among the multiple preset items except the first item.
[0013] In some embodiments, determining whether the first annotation result is a credible annotation result based on the confidence level and the target confidence level interval includes: if the confidence level is within the target confidence level interval, determining that the first annotation result is a credible annotation result; or if the confidence level is outside the target confidence level interval, determining that the first annotation result is not a credible annotation result.
[0014] In some embodiments, the confidence of the first annotation result is obtained in the following manner: obtaining the generation probability of each word in the first annotation result by the first annotation model, where the generation probability is predicted by the first annotation model in the process of generating word units; and taking the average value of the sum of the generation probabilities of each word unit in the first annotation result as the confidence of the first annotation result.
[0015] In some embodiments, the method further includes: if the first annotation result is not a credible annotation result, sending the target data to an annotation person to annotate the first project to obtain a second annotation result, and using the second annotation result as the annotation result of the target data on the first project.
[0016] In some embodiments, the target data includes: question information, and answer information generated by the target model in response to the question information; the labeling result of the target data on the first project is used to evaluate the question-answering ability of the target model.
[0017] In some embodiments, using the first annotation model to annotate the target data on the first project includes: generating guidance instructions based on the first project and the target data; and inputting the guidance instructions into the first annotation model to guide the first annotation model to annotate the target data on the first project.
[0018] In the second aspect, the present specification also provides a method for training and evaluating the capabilities of a labeling model, including: obtaining a first data set and a second data set corresponding to a first project, the first data set including multiple first sample data and sample labeling results of each first sample data on the first project, and the second data set including multiple second sample data and sample labeling results of each second sample data on the first project; training a first labeling model based on the first data set so that the first labeling model has labeling capabilities for the first project; labeling each second sample data in the second data set through the first labeling model to obtain predicted labeling results of each second sample data on the first project, and confidence levels corresponding to the predicted labeling results; determining the labeling capability information of the first labeling model for the first project based on the predicted labeling results corresponding to each second sample data and their confidence levels, as well as the sample labeling results corresponding to each second sample data, wherein the labeling capability information represents a confidence interval, and when the confidence level of the labeling results output by the first labeling model for the first project falls within the confidence interval, the labeling accuracy of the first labeling model is greater than or equal to the target accuracy.
[0019] As can be seen from the above technical solutions, the data labeling method and system provided in this specification use the labeling model to label the target data on the labeling project to obtain the labeling results and the confidence of the labeling results, and then determine whether the labeling result is a credible labeling result based on the labeling capability information and the confidence. If so, the labeling result is used as the labeling result of the target data on the labeling project. The data labeling method provided by the embodiments in this specification can realize automatic labeling based on the labeling model, which can partially or completely replace manual labeling, thereby improving labeling efficiency. And by retaining only the labeling results with high confidence in the labeling model, it is possible to improve the accuracy of the labeling results while improving the labeling efficiency.
[0020] Other features of the data annotation methods, annotation model training, and capability assessment methods and systems provided in this specification are partially outlined in the following description. The inventive aspects of the data annotation methods, annotation model training, and capability assessment methods and systems provided in this specification can be fully explained through practice or use of the methods, apparatus, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1A A schematic diagram illustrating an application of a data annotation method provided in accordance with an embodiment of this specification is shown;
[0023] Figure 1B A schematic diagram of an application scenario of a data annotation method provided according to an embodiment of this specification is shown;
[0024] Figure 2 shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0025] Figure 3 A flow chart of a method for training and evaluating the capabilities of a labeling model provided in accordance with an embodiment of this specification is shown;
[0026] Figure 4 A schematic diagram showing the relationship between the length of a confidence interval and the accuracy rate provided according to an embodiment of this specification;
[0027] Figure 5 A schematic diagram of dividing a second data set into subsets according to an embodiment of this specification is shown;
[0028] Figure 6 A schematic diagram illustrating dividing an episode into scene subsets according to an embodiment of this specification is shown; and
[0029] Figure 7 A flow chart of a data annotation method provided according to an embodiment of this specification is shown. DETAILED DESCRIPTION
[0030] The following description provides specific application scenarios and requirements for this specification, with the goal of enabling those skilled in the art to make and use the contents of this specification. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but is intended to be accorded the broadest scope consistent with the claims.
[0031] The terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. For example, as used herein, the singular forms "a," "an," and "the" may also include the plural forms unless the context clearly indicates otherwise. When used in this specification, the terms "comprise," "include," and / or "contain" are intended to refer to the presence of the associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.
[0032] These and other features of this specification, as well as the operation and function of the associated elements of the structure, and the economical assembly and manufacture of the components, can be significantly improved with consideration of the following description. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0033] The flowcharts used in this specification illustrate operations implemented by systems according to some embodiments of the present specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. Rather, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0034] For the convenience of description, the terms that will appear in the following text of this specification are first explained.
[0035] Large models: Refers to machine learning models with a large number of parameters and complex computational structures, typically constructed using deep neural networks (such as Transformer networks). The number of parameters in a large model can reach billions or even trillions. By training on massive amounts of data, it can learn complex patterns and features, demonstrating strong generalization capabilities. The training method for large models typically includes two stages: pre-training and fine-tuning. First, the large model is pre-trained using unlabeled data. This stage typically uses unsupervised learning methods to enable the large model to master general knowledge and patterns. Next, fine-tuning is performed using task-specific datasets. This is typically a supervised learning process, such as Supervised Fine-Tuning (SFT). The goal of the fine-tuning stage is to optimize the large model for a specific task, thereby improving its task adaptability. With its large number of parameters and diverse training data, large models are able to handle multimodal tasks such as text generation, speech recognition, and image generation.
[0036] The solutions provided herein can be used to automatically label data. In some embodiments, the labeled data can be training samples, which are then used to train a model. Through the data labeling methods described in the embodiments of this specification, the system can automatically label training samples and quickly generate labeled training data, thereby significantly improving data labeling efficiency and reducing reliance on manual intervention during the data labeling process.
[0037] In some embodiments, the labeled data can be question-and-answer data from a large model, used to evaluate the model's question-and-answer capabilities. Based on the context of the assessment, the system automatically labels the question-and-answer data, effectively identifying whether the answers generated by the model meet the requirements and assessing whether their performance in real-world applications is consistent with expectations. The system can also further identify potential errors in model predictions based on the labeling results, providing data support for model optimization and enabling continuous monitoring and optimization of the launched model.
[0038] In some embodiments, the annotated data can also be question-and-answer data of multiple large models, and the annotated data is used to evaluate the question-and-answer capabilities of different models. After each iteration of the large model, a quick evaluation is required to evaluate the improvement of the iterative version compared to the previous version. Through the data annotation method in the embodiments of this specification, the system can generate annotation results by comparing the output results of different iterative versions, and promptly discover deficiencies in model optimization based on the annotation results, thereby quickly adjusting the research and development direction.
[0039] Based on the data annotation method in the embodiments of this specification, rapid evaluation of the model in the above embodiments can be achieved. Figure 1A, the answers of model A and model B can be automatically evaluated to generate annotation results, and their performance can be compared. It is also possible to evaluate the answers of model A and model B separately to generate annotation results. It is also possible to annotate the content of user questions for other analysis. The data annotation method of this embodiment can achieve both fast annotation and high-accuracy annotation, thereby improving the efficiency and accuracy of model evaluation. In addition, by annotating different annotation items, it is also possible to achieve automated multi-dimensional evaluation.
[0040] It should be noted that this embodiment does not limit the object type of data annotation and is applicable to various data types and application scenarios such as text, images, videos, and audio.
[0041] It should be noted that the above-mentioned application scenarios for data annotation scenarios are only some examples of multiple usage scenarios. The data annotation method provided in this specification can be applied not only to the scenarios listed above, but also to other scenarios requiring data annotation. Those skilled in the art should understand that when the data annotation method provided in this specification is applied to other usage scenarios, its implementation method and technical effects are similar.
[0042] See Figure 1B , Figure 1B FIG1 shows a schematic diagram of an application scenario of data annotation provided according to an embodiment of this specification. Figure 1B As shown, the application scenario 100 may include a training and capability evaluation system 110 for an annotation model (hereinafter referred to as the training evaluation system 110 ) and an annotation system 120 .
[0043] See also Figure 1B The application scenario 100 may involve two phases, namely a training phase and an inference phase. The following describes the two phases using the first annotation model corresponding to the first project as an example.
[0044] During the training phase, the training evaluation system 110 can train the base model based on the first sample data in the first dataset to obtain multiple trained candidate models. The first dataset can be considered a training set. All of the trained candidate models have the ability to label the first project. The training evaluation system 110 can evaluate the multiple candidate models obtained after training based on the second sample data in the second dataset and select a labeled model from the multiple candidate models. Specifically, the second dataset can be used to evaluate the accuracy of the multiple candidate models, and the candidate model with the highest accuracy among the multiple candidate models is selected as the first labeled model. The second dataset can be considered a validation set. The training phase also includes evaluating the labeling capability information. The labeling capability information of the labeling model is determined based on the second sample data in the second dataset. The labeling capability information is then tested using the third sample data in the third dataset to ensure that appropriate labeling capability information is found. The third dataset can be considered a test set. Finally, the trained labeling model and the corresponding labeling capability information are deployed in the labeling system 120.
[0045] During the inference phase, the annotation system 120 has pre-deployed an annotation model. When target data to be annotated is required to be annotated for a first project, the annotation system 120 can annotate the target data for the first project, obtain a first annotation result and a confidence level corresponding to the first annotation result, and determine whether the first annotation result is a credible annotation result based on the annotation capability information and the confidence level. If so, the first annotation result is used as the annotation result for the target data for the first project.
[0046] In some embodiments, the training and evaluation system 110 may store data and instructions for implementing the training and capability evaluation methods for the annotation model provided herein, and may execute or be used to execute the data and instructions. In some embodiments, the training and evaluation system 110 may include hardware devices with data information processing capabilities and the necessary programs to drive the hardware devices.
[0047] In some embodiments, the annotation system 120 may store data and instructions for implementing the data annotation methods provided herein, and may execute or be used to execute the data and instructions. In some embodiments, the annotation system 120 may include hardware devices with data information processing capabilities and the necessary programs required to drive the hardware devices.
[0048] It is understandable that the training evaluation system 110 and the annotation system 120 may correspond to the same system or to different systems, and this specification does not impose any limitation on this.
[0049] It should be noted that the training and evaluation system 110 can correspond to a single device or a cluster of devices, and this specification does not impose any restrictions on this. When the training and evaluation system 110 corresponds to a single device, the training and evaluation methods of the annotation model can be executed entirely on that device. When the training and evaluation system 110 corresponds to a cluster of devices, the training and evaluation methods of the annotation model can be executed in coordination on multiple devices corresponding to the cluster, and this specification does not impose any restrictions on this.
[0050] The annotation system 120 may correspond to a single device or a cluster of devices, and this specification does not impose any restrictions on this. When the annotation system 120 corresponds to a single device, the data annotation method may be executed entirely on that device. When the annotation system 120 corresponds to a cluster of devices, the data annotation method may be executed collaboratively on multiple devices corresponding to the cluster, and this specification does not impose any restrictions on this.
[0051] It should be noted that the user data obtained in this manual has been authorized by the user and does not involve user privacy.
[0052] Figure 2 FIG2 shows a hardware structure diagram of a computing system 200 provided according to an embodiment of this specification. The computing system 200 can be used as Figure 1B The training and evaluation system 110 in the embodiment of the present invention executes the training and capability evaluation method of the annotation model described in this specification. The computing system 200 can also be used as Figure 1B The annotation system 120 in the embodiment executes the data annotation method described in this specification.
[0053] like Figure 2 As shown, computing system 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, computing system 200 may further include communication port 250 and internal communication bus 210. Computing system 200 may further include I / O component 260.
[0054] The internal communication bus 210 can connect various system components, such as the storage medium 230 , the processor 220 , the communication port 250 , and the I / O component 260 .
[0055] I / O components 260 support input / output between computing system 200 and other components.
[0056] Communication port 250 is used for data communication between computing system 200 and the outside world. For example, communication port 250 can be used for data communication between computing system 200 and network 140. Communication port 250 can be a wired communication port or a wireless communication port.
[0057] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, which may include a program, routine, object, component, data structure, procedure, module, etc.
[0058] At least one processor 220 can be in communication with at least one storage medium 230. When the computing system 200 is running, the at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the data labeling method provided in this specification. The processor 220 can execute the steps included in the data labeling method. The processor 220 can be in the form of one or more processors. In some embodiments, the processor 220 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0059] For illustrative purposes only, the computing system 200 shown in the accompanying drawings only shows one processor 220. However, it should be noted that the computing system 200 described herein may also include multiple processors. Therefore, the operations and / or method steps disclosed herein may be performed by a single processor or jointly by multiple processors. For example, if the computing system 200 is described herein as performing steps A and B by the processor 220, it should be understood that steps A and B may also be performed jointly or separately by two different processors 220 (e.g., the first processor performs step A and the second processor performs step B, or the first and second processors jointly perform steps A and B).
[0060] Figure 31 shows a flow chart of a method P300 for training and evaluating the capabilities of a labeling model provided in accordance with an embodiment of the present specification. As before, the training evaluation system 110 can execute the method P300 for training and evaluating the capabilities of a labeling model in accordance with the present specification. Specifically, the processor in the training evaluation system 110 can read the instruction set stored in its local storage medium, and then execute the method P300 for training and evaluating the capabilities of a labeling model in accordance with the provisions of the instruction set. Figure 3 As shown, method P300 may include steps S310-S340:
[0061] S310: Obtain a first data set and a second data set corresponding to the first project, where the first data set includes a plurality of first sample data and a sample labeling result of each first sample data on the first project, and the second data set includes a plurality of second sample data and a sample labeling result of each second sample data on the first project.
[0062] During the training and capability evaluation of the annotation model, it is necessary to first obtain sample data for training and evaluation. Data sources may include annotation data specifically used for training the annotation model, historical annotation data, etc. In the embodiments of this specification, sample data can be obtained from a designated platform, such as a dedicated annotation tool platform. On the platform, manpower can be organized to annotate the data, and manually annotated data can be obtained from the annotation tool platform as sample data. The type of sample data here can be at least one of text, voice, picture or video. Those skilled in the art can train models for different data types based on the method provided in the embodiments of this specification to obtain different types of annotation models.
[0063] When obtaining sample data from the annotation tool platform, you need to first configure the annotation project task template ID, annotation project, and annotation project dependency fields of the annotation project based on the annotation project configuration platform to pull different sample data for different annotation projects.
[0064] The embodiments of this specification also provide examples of various annotation projects. Based on different application scenarios, users can select different annotation projects. The annotation projects here can also be regarded as atomic annotation units, which are independent and minimum annotation units. Each annotation project represents a specific goal in the data annotation task. By splitting the task into independent annotation units, the data can be processed more carefully and accurately, thereby improving the efficiency and accuracy of annotation.
[0065] The following uses the evaluation annotation project as an example to provide some examples of annotation problems and their corresponding candidate annotation contents:
[0066] 1. Known content: <User's question: xxx>, annotation item <What is the intention of the user's question?>, candidate annotation content:
Intention 1, Intention Figure 2 , Intention Figure 3 …
[0067] 2. Known content: <User's question: xxx>, annotation item <The emotion of the user's question?>, candidate annotation content:
Positive, Neutral, Negative
[0068] 3. Known content: <User's question: xxx, User's answer: xxx>, annotation item <Does the answer solve the user's problem?>, candidate annotation content:
Solve, Unsolved
[0069] 4. Known content: <Answer A: xxx, Answer B: xxx>, annotation item <Which answer of A and B is better?>, candidate annotation content:
A, B
[0070] 5. Known content: <User's question: xxx, User's answer: xxx>, annotation item <The professionalism of the answer?>, candidate annotation content: 【0, 1, 2, 3, 4, 5】;
[0071] 6. Known content: <User's question: xxx, User's answer: xxx>, annotation item <The relevance of the answer?>, candidate annotation content:
Relevant, Irrelevant
[0072] 7. Known content: <User's question: xxx, User's answer: xxx>, annotation item <The timeliness of the answer?>, candidate annotation content:
Good, Bad
[0073] 8. Known content: <Original text: xxx, Rewritten text: xxx>, annotation item <Does the rewrite pass?>, candidate annotation content:
Good, Bad, Neutral
[0074] In some embodiments, different annotation models can also be trained for different annotation items, that is, one annotation model only has the corresponding annotation ability for one annotation item. In some embodiments, the same annotation model can also be trained for different annotation items, that is, the same annotation model has the ability to annotate multiple annotation items. Those skilled in the art can understand that in different usage scenarios, the annotation model can be trained differently to meet the needs of annotating different annotation items.
[0075] After configuring the content based on different annotation items, the training and evaluation system 110 can automatically obtain sample data from the annotation data return table of the annotation tool platform.
[0076] After obtaining the sample data, it is also necessary to clean and preprocess the sample data, remove the noise and irrelevant information in the sample data, and standardize the format of the sample data to ensure the consistency and high quality of the sample data. The training and evaluation system 110 then assembles the cleaned sample data into corresponding prompt-label, and uses the prompt to guide the annotation model to annotate the sample data. Among them, the prompt in the prompt-label contains the prompt words input to the annotation model, and the prompt words contain the aforementioned annotation items and candidate annotation contents. The label in the prompt-label is the corresponding sample annotation result in the sample data (which can also be regarded as the true result of the annotation item). Taking the above annotation item 3 as an example, please refer to the following prompt-label assembled for annotation item 3.
[0077] Prompt: You are a judge, and you need to judge whether the current answer solves the user's problem. Your answer can only be one of [solved, not solved].
[0078] <Current problem>{query}< / Current problem>
[0079] <Current answer>{answer}< / Current answer>
[0080] According to the evaluation criteria, your judgment result is:
[0081] Label: Not solved.
[0082] In some embodiments, it is also possible to remove duplicates and disambiguate the sample data after assembling it into prompt-label, avoid problems of data duplication and inconsistency, and improve the training quality. The sample data can be further divided into a first data set (training set), a second data set (validation set), and a third data set (test set), which can be split according to a ratio (such as 0.9, 0.07, 0.03 respectively). Taking the first item as the annotation item as an example, the data set can be divided into a first data set and a second data set corresponding to the first item. It should be noted that the first data set (training set), the second data set (validation set), and the third data set (test set) can be sample data based on the same data source, or sample data based on different data sources.
[0083] S320: Train a first annotation model based on the first data set, so that the first annotation model has the annotation ability for the first item.
[0084] In some embodiments, the training evaluation system 110 can first obtain a base model, train the base model based on the first data set to obtain multiple candidate models; and use the second data set to evaluate the accuracy of the multiple candidate models, and use the candidate model with the highest accuracy among the multiple candidate models as the first labeled model.
[0085] The base model can be a common large model such as qwen2-0.5B. SFT fine-tuning training is performed on the base model using the first dataset to obtain several trained candidate models. The candidate model with the highest accuracy is selected from the second dataset as the first annotation model. Fine-tuning the large language model using the first dataset (including manual annotation results) can better meet the standards of manual annotation.
[0086] After training and selection, the first annotation model is capable of annotating the first item. However, the accuracy of its annotation results remains difficult to guarantee. Manual annotation accuracy is approximately 95%, so ensuring that the final annotation results meet or exceed manual accuracy becomes a key challenge for machine annotation.
[0087] To this end, the embodiments of this specification also propose an embodiment for screening high-accuracy annotation results based on model annotation capability information. The confidence interval is represented by the annotation capability information. When the confidence of the annotation result output by the first annotation model for the first project falls within the confidence interval, it indicates that the annotation accuracy of the model is greater than or equal to the target accuracy. In the inference stage, if the confidence of the annotation result falls within the confidence interval, the annotation result is considered to be credible; if the confidence does not fall within the interval, the annotation result is considered to be unreliable.
[0088] During the inference process, the training and evaluation system 110 retains only those annotation results whose confidence levels fall within the confidence interval, thereby ensuring the overall accuracy of the annotation results. This effectively improves the reliability of machine annotation and ensures that the final results meet the expected standards. Therefore, determining and evaluating annotation capability information becomes a key step in the training and evaluation phase of the annotation model. See steps S330 and S340.
[0089] S330: Label each second sample data in the second data set using the first labeling model to obtain a predicted labeling result of each second sample data on the first item and a confidence level corresponding to the predicted labeling result.
[0090] The aforementioned guidance instructions are applicable to each process of training, evaluation, and inference. Therefore, when the first annotation model annotates each second sample data in the second dataset, the training and evaluation system 110 can generate guidance instructions based on the first project and the second sample data; and input the guidance instructions into the first annotation model to guide the first annotation model to annotate each second sample data.
[0091] The process of generating output in a large model is to predict the next word (token) based on the previous content. Each token can be regarded as a character or word. The model will select the token with the highest probability from multiple predicted tokens as the current output, then concatenate this output token with the previous content and continue to predict the next token until it encounters a terminator. The probability of each generated token is added and averaged to obtain a probability value that represents the probability of output generation. The generation probability is positively correlated with the degree of certainty of the model. That is, the greater the generation probability, the more confident the model is about the result, and the more likely the true result is to be correct. See Figure 4 This specification verifies the relationship between accuracy and the confidence interval represented by the generation probability. It can be seen that when the upper limit of the confidence interval is fixed at 1, the shorter the confidence interval (meaning the closer the lower limit of the confidence interval is to the upper limit of 1), the greater the proportion of correct prediction results. As low-confidence results are added, the accuracy rate will also decrease. In some embodiments, the confidence of the predicted annotation result can be expressed as the average of the sum of the generation probabilities of each word.
[0092] S340: Based on the predicted labeling results corresponding to each second sample data and their confidence levels, as well as the sample labeling results corresponding to each second sample data, determine the labeling capability information of the first labeling model for the first project, wherein the labeling capability information represents a confidence interval. When the confidence level of the labeling results output by the first labeling model for the first project falls within the confidence interval, the labeling accuracy of the first labeling model is greater than or equal to the target accuracy.
[0093] In some embodiments, the training evaluation system 110 can search for the annotation capability information using a grid search method. Specifically, an initial confidence interval is first determined, where the lower limit of the initial confidence interval is a preset first confidence level and the upper limit is a preset second confidence level, where the second confidence level is greater than the first confidence level. The lower limit of the initial confidence interval is increased by a specified step size to obtain a target confidence interval, until the target confidence interval is determined to meet a preset condition, wherein the preset condition includes: the accuracy of the annotation results of each second sample data falling within the target confidence interval is greater than or equal to the target accuracy rate; and the annotation capability information is generated based on the target confidence interval.
[0094] In some embodiments, when determining the initial confidence interval, a lower preset first confidence value can be selected, and the lower limit of the confidence interval can be gradually increased through a grid search algorithm (for example, increasing by 0.01 each time). Until the accuracy of the annotation results of each second sample data falling within the target confidence interval is greater than or equal to the target accuracy. At this point, it can be considered that the annotation capability information generated by the confidence interval has met the conditions. In this process, the upper limit of the confidence interval can remain unchanged, for example, the interval [0,1] is selected, where the upper limit "1" is fixed, and the lower limit starts from 0.1 and gradually increases according to the grid search method until a suitable lower limit confidence threshold is found, thereby determining the target confidence interval and finally obtaining the annotation capability information. Among them, the target accuracy can be set to the accuracy that can be achieved by manual annotation (95%), or it can be selected based on the requirements of different scenarios, and this specification does not limit it.
[0095] In some embodiments, it is also possible to confirm in the third data set whether the labeling capability information meets the preset conditions to further verify the labeling capability information. If the model accuracy of the part of each third sample data in the third data set that adopts the confidence interval is also greater than or equal to the target accuracy, it is considered that the labeling capability information meets the conditions and the grid search is terminated. At this time, the confidence interval is considered to be the labeling capability information. Specifically, the training evaluation system 110 first obtains the third data set corresponding to the first project, and the third data set includes a plurality of third sample data and the sample labeling results of each third sample data on the first project. The third sample data in the third data set is labeled by the first labeling model to obtain the predicted labeling results of each third sample data on the first project, and the confidence corresponding to the predicted labeling result; the preset conditions also include: the accuracy of the labeling results of each third sample data in the third data set that falls into the target confidence interval is greater than or equal to the target accuracy.
[0096] The annotation model, trained and evaluated using the above process, retains high-confidence predictions during inference, typically achieving high discrimination accuracy. Low-confidence predictions are then handed over to human judgment, which further improves accuracy by appropriately sacrificing some recall. This approach effectively optimizes the balance between precision and recall, ensuring high-quality final annotation results.
[0097] However, in some embodiments, searching for annotation capability information directly on the entire data set may produce undesirable results in certain annotation projects. For example, consider the first project, annotation project 3, where the question is "Does the answer solve the user's problem?" and the candidate annotation content is [solved, unsolved]. If, in the first dataset (training set), the data labeled "solved" accounts for a high proportion, while the data labeled "unsolved" is relatively rare, then the annotation model will perform better when processing data labeled "solved" because the annotation model has been exposed to more data of this type during training and has a higher confidence in generating data labeled "solved." However, the annotation model will not perform as well when processing data labeled "unsolved" because the annotation model has been exposed to less data of this type and has a lower confidence in generating data labeled "unsolved." If the second dataset (validation set) is used to search for annotation capability information on all second sample data as a whole, since the model has processed fewer data labeled "unsolved," the model's confidence in predicting "unsolved" will be lower, resulting in poor reliability of the annotation results for this portion. This makes it difficult to effectively calculate the annotation capability for the "unsolved" label, thereby affecting the overall accuracy. Since confidence cannot distinguish the annotation ability of the “unresolved” label, the accuracy of the model is affected by the “unresolved” part, which ultimately leads to a decrease in the accuracy of the overall prediction results (including “resolved” and “unresolved”).
[0098] Therefore, for annotation items for which no annotation capability information can be found, some embodiments of this specification further propose determining corresponding sub-capability information based on the candidate annotation content of the annotation item. Specifically, a first item corresponds to multiple candidate annotation contents, and the predicted annotation result is one of the multiple candidate annotation contents; the annotation capability information includes: sub-capability information corresponding to each of the multiple candidate annotation contents, where the sub-capability information corresponding to each candidate annotation content represents the confidence interval applicable to the candidate annotation content.
[0099] Taking the above-mentioned annotation item 3 as an example, the candidate annotation contents include [solved, unsolved]. The corresponding sub-capability information is searched for the "solved" and "unsolved" labels respectively. When there are two candidate annotation contents, there will be two corresponding sub-capability information. For the case where the label is "solved", the annotation model has a higher degree of confidence in the generated result, so the confidence interval of the "solved" sub-capability information representation determined by the training evaluation system 110 may be larger. The annotation model can screen out more high-accuracy prediction results of the "solved" class based on the sub-capability information of "solved". For the label "unsolved", the label model has a lower degree of confidence in the "unsolved" class data, so the confidence interval of the "unsolved" sub-capability information representation determined by the training evaluation system 110 may be smaller. The annotation model will also screen out fewer high-accuracy results of the "unsolved" class.
[0100] In some embodiments, the training evaluation system 110 divides the second data set into multiple sets based on the predicted labeling results corresponding to each second sample data, each set including multiple second sample data, and the predicted labeling results corresponding to each second sample data in the same set correspond to the same candidate labeling content; for each set, the predicted labeling results corresponding to the set are used as the target candidate labeling content, and based on the predicted labeling results corresponding to each second sample data in the set and their confidence levels, as well as the sample labeling results corresponding to each second sample data in the set, the sub-capability information corresponding to the target candidate labeling content is determined.
[0101] See Figure 5 , taking the above-mentioned annotation item 3 as an example, the predicted annotation results include two labels: "solved" and "unsolved". Based on the predicted annotation results, the second data set is divided into two corresponding sets - Set A: "solved set" and Set B: "unsolved set". Among them, the "solved set" contains all sample data whose predicted annotation results are "solved", and the "unsolved set" contains all sample data whose predicted annotation results are "unsolved". Next, the grid search algorithm is applied to the "solved set" and the "unsolved set" respectively to gradually find the confidence interval. Based on the above embodiment, the sub-capability information corresponding to the "solved" label and the "unsolved" label can be determined respectively, thereby generating accurate annotation capability information for each label.
[0102] Through the above embodiment, searching for the corresponding sub-capability information for candidate annotation content can maximize the selection of highly accurate prediction results, thereby improving the overall annotation quality of the model. Judging the annotation results based on the sub-capability information of each tag not only avoids deviations in the overall annotation, but also refines the model's judgment criteria, ensuring that the prediction results of each tag are fully verified and optimized, and ensuring the accuracy of the overall annotation.
[0103] Furthermore, in some embodiments, if suitable sub-capability information cannot be found or the prediction results obtained using the sub-capability information are still unsatisfactory, this may indicate that the sub-capability's discrimination is still insufficient, or if further accuracy is desired, further refinement may be considered. To this end, embodiments of this specification also propose an embodiment for further refining the sub-capability information to achieve optimization.
[0104] Specifically, the sub-capability information corresponding to each candidate annotation content includes confidence intervals corresponding to multiple scenes. The training evaluation system 110 determines the scene corresponding to the predicted annotation result corresponding to each second sample data in the collection; divides the collection into multiple scene subsets, each scene subset includes multiple second sample data in the collection, and the predicted annotation results corresponding to each second sample data in the same scene subset correspond to the same scene; for each scene subset, the predicted annotation result corresponding to the scene subset is used as the predicted annotation result corresponding to the target scene, and based on the predicted annotation results corresponding to each second sample data in the scene subset and its confidence, as well as the sample annotation results corresponding to each second sample data in the scene subset, the confidence interval corresponding to the target scene is determined.
[0105] It is worth noting that the "scenario" described in the embodiments of this specification is a broad term, referring to any scenario in which sample data can be segmented based on the predicted annotation results. For example, it can refer to user behavior scenarios, user intent scenarios, user emotion scenarios, or time scenarios. In other words, further subdivision of sub-capabilities can be called scenario segmentation. In the same scenario, the more similar the data performance is, the better the discriminability of the confidence interval will be. In other words, the finer the segmentation scenario, the stronger the screening ability will be.
[0106] See Figure 6 Continuing with the example of annotation project 3, we will further describe the process of segmenting the intent scenarios in annotation project 1. Assume there are four intent scenarios: Scenario 1, Scenario 2, Scenario 3, and Scenario 4. The combination of annotation project 3 and annotation project 1 yields 2*4=8 scenario subsets. Confidence intervals must be calculated for each scenario subset. Within the "solved" and "unsolved" subsets, each subset is further subdivided into Scenario 1, Scenario 2, Scenario 3, and Scenario 4 based on the intent scenarios corresponding to the predicted annotation results. For example, within the "solved" subset, the subsets can be further subdivided into the "solved - Scenario 1" subset, the "solved - Scenario 2" subset, and so on; within the "unsolved" subset, the subsets can be further subdivided into the "unsolved - Scenario 1" subset, the "unsolved - Scenario 2" subset, and so on. Finally, a grid search algorithm is applied to each scenario subset to gradually find the confidence threshold for the confidence interval. The confidence threshold refers to the lower limit of the confidence interval. Through this process, we can determine the confidence thresholds for specific scenarios, such as "Solution - Scenario 1," "Solution - Scenario 2," and so on. By dividing the data corresponding to sub-capabilities into more specific and detailed scenarios, the confidence thresholds become more discriminative. The annotation model can calculate independent confidence thresholds for each specific scenario, thereby identifying more high-confidence prediction results.
[0107] During the above process, the training evaluation system 110 inputs each second sample data set into a trained scene recognition model to determine the scene corresponding to the predicted labeling result for each second sample data set. It is worth noting that the scene recognition model can be a separately trained model for scene recognition, or it can be a labeling model corresponding to other labeling items selected from a database.
[0108] In some embodiments, the training evaluation system 110 stores the first labeling model and the labeling capability information of the first labeling model for the first project in a database, wherein the database is used to store labeling models corresponding to multiple preset projects, and the labeling capability information of each labeling model for its corresponding preset project. The labeling capability information can be stored in the database in a tabular form, or in other feasible ways. Please refer to Table 1, which is used to store the labeling capability information of the labeling model for the labeling project. Taking labeling project 3 as an example, Table 1 includes the labeling project, labeling model, scene subset, and the confidence interval corresponding to each scene subset.
[0109] Table 1
[0110]
[0111] In some embodiments, during the training and evaluation phase, different labeling models corresponding to different preset items can be trained based on different sample data, and the labeling capability information of each labeling model for its corresponding preset item can be obtained by evaluation. Since the accuracy of the high confidence part is greater than or equal to the target accuracy on both the second data set (validation set) and the third data set (test set), it can be considered that the labeling model and labeling capability information after training and evaluation can be applied to other data. The above-mentioned labeling models and corresponding labeling capability information are stored in the database. During the evaluation process, the labeling models stored in the database can also be used for scene recognition of other labeling items.
[0112] Figure 7 1 shows a flow chart of a data annotation method P400 provided according to an embodiment of the present specification. As before, the annotation system 120 can execute the data annotation method P400 of the present specification. Specifically, the processor 220 in the annotation system 120 can read the instruction set stored in its local storage medium, and then execute the data annotation method P400 of the present specification according to the provisions of the instruction set. Figure 7 As shown, method P400 may include steps S410-S440:
[0113] S410: Obtain target data to be labeled, and determine a first item of the target data to be labeled.
[0114] The target data may have multiple annotation items to be annotated, and the execution process of each annotation item is similar. Here, the first item is used as an example. The type of the target data here can be at least one of text, voice, picture or video.
[0115] In a scenario where the annotation model is used to evaluate model performance, the target data may include: question information, and answer information generated by the target model in response to the question information; the annotation results of the target data on the first item are used to evaluate the question-answering capability of the target model.
[0116] S420: Obtain a trained first labeling model and labeling capability information of the first labeling model for the first project, wherein the labeling capability information represents a confidence interval. When the confidence of the labeling result output by the first labeling model for the first project falls within the confidence interval, the labeling accuracy of the first labeling model is greater than or equal to the target accuracy.
[0117] In some embodiments, the labeling system 120 obtains a trained first labeling model and labeling capability information of the first labeling model for a first project from a database, wherein the database stores labeling models corresponding to multiple preset projects and labeling capability information of each labeling model for its corresponding preset project, the first project is one of the multiple preset projects, and the first labeling model is the labeling model corresponding to the first project.
[0118] If the target data has multiple items to be annotated, the corresponding annotation models and corresponding annotation capability information can be selected from the database. Multiple annotation models can be processed in parallel to annotate multiple items of the target data at the same time, which can shorten the annotation time. Before use, the information in the database can be deployed to the annotation system 120. When using it, the user enters the data and the name and configuration of the corresponding annotation project into the annotation system 120, and the annotation system 120 will pull the corresponding annotation model and corresponding annotation capability information for prediction.
[0119] S430: Using the first labeling model to label the target data on the first item, to obtain a first labeling result and a confidence level corresponding to the first labeling result.
[0120] During the inference phase, the labeling system 120 generates guidance instructions based on the first project and the target data; and inputs the guidance instructions into the first labeling model to guide the first labeling model to label the target data on the first project.
[0121] Unlike the training and evaluation phases, the guidance instructions in the inference phase do not include true labels. The rest of the guidance instructions are similar to those in the inference phase. Both phases use the guidance instructions to guide the first labeling model to label the target data on the first item, and we will not elaborate on this here.
[0122] As mentioned above, during the inference phase, the confidence level of the first annotation result is determined by obtaining the generation probability of each token in the first annotation result from the first annotation model. The generation probability is the prediction obtained by the first annotation model during the token generation process. The confidence level of the first annotation result is then calculated as the average of the sum of the generation probabilities for each token in the first annotation result. This process will not be further elaborated here.
[0123] S440: Determine whether the first annotation result is a credible annotation result based on the annotation capability information and the confidence level, and if so, use the first annotation result as the annotation result of the target data on the first project.
[0124] Assuming the confidence interval is [a, b], annotation results with confidence levels within [a, b] are considered high-confidence results, and their accuracy meets the target accuracy. If the confidence levels do not fall within [a, b], the accuracy of these annotation results is considered to be below the target accuracy, and the first annotation result is considered unreliable. Here, "a" is the confidence threshold.
[0125] In some embodiments, if the first annotation result is not a credible annotation result, the annotation system 120 may send the target data to an annotator for annotation of the first project to obtain a second annotation result, and use the second annotation result as the annotation result of the target data on the first project. In the embodiments of this specification, low-confidence results are manually processed, combining the advantages of model annotation and manual annotation, thereby improving annotation efficiency and ensuring annotation accuracy. This ensures the accuracy of the overall annotation and achieves high-quality and efficient data processing.
[0126] As mentioned above, during the training and evaluation phase, if the first project uses a labeling capability information to determine that the discrimination of the credible labeling result is not high, then when there are multiple candidate labeling contents corresponding to the first project, and the first labeling result is one of the multiple candidate labeling contents, the labeling capability information includes: the sub-capability information corresponding to each of the multiple candidate labeling contents, wherein the sub-capability information corresponding to each candidate labeling content represents the confidence interval applicable to the candidate labeling content.
[0127] If the confidence level is within the confidence interval applicable to the first annotation result, the first annotation result is determined to be a credible annotation result; or if the confidence level is outside the confidence interval applicable to the first annotation result, the first annotation result is determined not to be a credible annotation result. As mentioned above, a confidence threshold can also be used to determine whether the confidence level falls within the confidence interval. Here, the confidence threshold is the lower limit of the confidence interval. If the confidence level is greater than the confidence threshold, it can be considered that the confidence level falls within the confidence interval. The accuracy of this part of the annotation result can meet the expected target accuracy, and the first annotation result is a credible annotation result.
[0128] Taking the example of labeled item 3 in the first project, assume that during the training and evaluation phases, the confidence interval for "solved" is determined to be [0.5, 1], with a confidence threshold of 0.5 (i.e., the lower limit of the confidence interval), and the confidence interval for "unsolved" is determined to be [0.8, 1], with a confidence threshold of 0.8. If the first labeling result is "solved" and its confidence is 0.6, which is greater than 0.5, the confidence of the first labeling result is considered to fall within the confidence interval and is therefore credible. Ultimately, "solved" is used as the labeling result for the target data in the first project. If the first labeling result is "unsolved" and its confidence is 0.6, which is less than 0.8, the labeling result is deemed untrustworthy. In this case, the target data can be labeled by human annotators to ensure the accuracy of the results. By retaining only the high-confidence labeling results in the corresponding sub-capability information, the accuracy of the final labeling result can be effectively guaranteed.
[0129] As mentioned above, during the training and evaluation phases, if the credible annotation results determined based on the sub-capability information still cannot fully distinguish the accuracy of the data, it is necessary to further divide the sub-capability information into scenarios. That is to say, the sub-capability information corresponding to each candidate annotation content includes confidence intervals corresponding to multiple scenarios, which more finely divides the sub-capability information of the annotation results. Then in the reasoning phase, the annotation system 120 will first determine the target scenario to which the target data belongs, and then obtain the target confidence interval corresponding to the target scenario from the sub-capability information corresponding to the first annotation result; based on the confidence and the target confidence interval, it determines whether the first annotation result is a credible annotation result.
[0130] As mentioned above, the target scene here is a wide range of scenes, and the annotation system 120 can obtain a trained scene recognition model; using the scene recognition model, the target data is subjected to scene recognition processing to obtain the target scene. The scene recognition model can be a separately trained model for identifying the scene of the target data, or it can be another annotation model in the database. In the case where the scene recognition model is another annotation model, the annotation system 120 obtains the trained scene recognition model from the aforementioned database, and the scene recognition model is the second annotation model corresponding to the second project, and the second project is the other project in the multiple preset projects except the first project.
[0131] Similarly, in the scenario of sub-capability information corresponding to the first annotation result, if the confidence level is within the target confidence level interval, the first annotation result is determined to be a credible annotation result; or if the confidence level is outside the target confidence level interval, the first annotation result is determined to be not a credible annotation result.
[0132] Taking the example of annotation item 3 in the first project, assume that during the training and evaluation phases, the confidence interval for the "Solution - Scenario 1" scenario subset is determined to be [0.5, 1], with a confidence threshold of 0.5 (i.e., the lower limit of the confidence interval), and the confidence interval for the "Solution - Scenario 2" scenario subset is determined to be [0.7, 1], with a confidence threshold of 0.7. If the first annotation result, after being recognized by the first annotation model and the scene recognition model, is "Solution - Scenario 1," and its confidence is 0.6, which is greater than 0.5, the confidence of the first annotation result is considered to fall within the confidence interval and is therefore credible. Ultimately, "Solution" is used as the annotation result for the target data in the first project. If the first annotation result is "Solution - Scenario 2," and its confidence is 0.6, which is less than 0.7, then the annotation result is deemed unreliable. In this case, the target data will be labeled by human annotators to ensure the accuracy of the results.
[0133] In some embodiments, if there are also labeling results determined by expert experience rules for the first item labeling of the target data, or labeling results predicted by other labeling models trained based on other base models, the labeling system 120 can combine the labeling results of the above-mentioned first labeling model and other results for voting, and take the result with the most votes from multiple parties as the final result, and use a posteriori voting method to further improve the accuracy of the final result.
[0134] The computing system 200 provided in the embodiments of this specification can be connected to or integrated into a labeling tool platform for an automated data labeling tool for such a platform. The computing system 200 can also serve as the underlying labeling engine of the platform to provide data labeling tasks. The computing system 200 can also serve as an atomic evaluation capability (based on each labeling project) and access other systems through an interface. Other systems can call the computing system 200 through scheduling, and the computing system 200 pulls the corresponding labeling model from the database, and then returns the labeling results to other systems based on the above embodiment.
[0135] In summary, in the data labeling method and system provided in this specification, the labeling model is used to label the target data on the labeling project to obtain the labeling result and the confidence of the labeling result, and then based on the labeling capability information and the confidence, it is determined whether the labeling result is a credible labeling result. If so, the labeling result is used as the labeling result of the target data on the labeling project. The data labeling method provided by the embodiments in this specification can realize automatic labeling based on the labeling model, which can partially or completely replace the actual labeling, and ultimately achieve the effect of reducing or even replacing manual labeling. And by retaining only the labeling results with high confidence in the labeling model, the accuracy of the labeling results can be improved while ensuring efficiency.
[0136] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for data labeling. When the at least one instruction set is executed by a processor, the at least one instruction set instructs the processor to implement the steps of the labeling model training and capability assessment method P300 or the data labeling method P400 described in this specification. In some possible implementations, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product is run on a computing system 200, the program code is used to cause the computing system 200 to perform the steps of the labeling model training and capability assessment method P300 or the data labeling method P400 described in this specification. The program product for implementing the above method can use a portable compact disk read-only memory (CD-ROM) to include program code and can be run on the computing system 200. However, the program product of this specification is not limited to this. In this specification, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system. The program product can use any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media include: a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. The computer-readable storage medium may include a data signal transmitted in baseband as part of a carrier wave, which carries readable program code. This transmitted data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable storage medium may also be any readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof. Program code for performing the operations of this specification may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages.The program code may execute entirely on computing system 200, partly on computing system 200, as a stand-alone software package, partly on computing system 200 and partly on a remote computing system, or entirely on the remote computing system.
[0137] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0138] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0139] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0140] It should be understood that in the foregoing descriptions of the embodiments of this specification, to facilitate understanding of a feature and to simplify this specification, various features are combined in a single embodiment, figure, or description thereof. However, this does not necessarily mean that these features are combined. When reading this specification, a person skilled in the art may label some of the devices as separate embodiments. In other words, the embodiments of this specification can also be understood as the integration of multiple sub-embodiments. The content of each sub-embodiment is also valid even when it includes fewer than all the features of a single previously disclosed embodiment.
[0141] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, articles, and the like, cited herein, except to the extent that it is inconsistent or conflicting with this document or that it has a limiting effect on the broadest scope of the claims, is hereby incorporated by reference for all purposes now or hereafter connected with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of a term in any material and the description, definition, and / or use of a term in this document, the term in this document shall control.
[0142] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A data annotation method, comprising: Obtaining target data to be labeled, and determining a first item of the target data to be labeled; Obtaining a trained first labeling model and labeling capability information of the first labeling model for the first project, wherein the labeling capability information represents a confidence interval, and when a confidence level of a labeling result output by the first labeling model for the first project falls within the confidence interval, a labeling accuracy of the first labeling model is greater than or equal to a target accuracy rate; Using the first labeling model to label the target data on the first item, obtaining a first labeling result and a confidence level corresponding to the first labeling result; as well as Based on the labeling capability information and the confidence level, it is determined whether the first labeling result is a credible labeling result, and if so, the first labeling result is used as the labeling result of the target data on the first project.
2. The method according to claim 1, wherein The obtaining of the trained first labeling model and labeling capability information of the first labeling model for the first project includes: The trained first annotation model and the annotation capability information of the first annotation model for the first project are obtained from the database, wherein The database stores annotation models corresponding to multiple preset items and annotation capability information of each annotation model for its corresponding preset item. The first item is one of the multiple preset items, and the first annotation model is the annotation model corresponding to the first item.
3. The method according to claim 1, wherein The first item corresponds to multiple candidate annotation contents. The first annotation result is one of the multiple candidate annotation contents; The annotation capability information includes: sub-capability information corresponding to each of the plurality of candidate annotation contents, wherein the sub-capability information corresponding to each candidate annotation content represents a confidence interval applicable to the candidate annotation content.
4. The method according to claim 3, wherein: The determining, based on the annotation capability information and the confidence level, whether the first annotation result is a credible annotation result includes: If the confidence level is within the confidence level interval applicable to the first annotation result, then determining that the first annotation result is a credible annotation result; or If the confidence level is outside the confidence level interval applicable to the first annotation result, it is determined that the first annotation result is not a credible annotation result.
5. The method according to claim 3, wherein: The sub-capability information corresponding to each candidate annotation content includes confidence intervals corresponding to multiple scenarios, and determining whether the first annotation result is a credible annotation result based on the annotation capability information and the confidence intervals includes: Determine the target scenario to which the target data belongs, Obtaining a target confidence interval corresponding to the target scenario from the sub-capability information corresponding to the first annotation result; Based on the confidence level and the target confidence interval, it is determined whether the first annotation result is a credible annotation result.
6. The method according to claim 5, wherein: Determining the target scenario to which the target data belongs includes: Obtain a trained scene recognition model; The target data is subjected to scene recognition processing using the scene recognition model to obtain the target scene.
7. The method according to claim 6, wherein: The obtaining of the trained scene recognition model comprises: The trained scene recognition model is obtained from the database, wherein The database stores annotation models corresponding to multiple preset items, and annotation capability information of each annotation model for its corresponding preset item, the first item is one of the multiple preset items, the first annotation model is the annotation model corresponding to the first item, the scene recognition model is the second annotation model corresponding to the second item, and the second item is other items among the multiple preset items except the first item.
8. The method according to claim 5, wherein The determining, based on the confidence level and the target confidence level interval, whether the first annotation result is a credible annotation result includes: If the confidence level is within the target confidence level interval, determining that the first annotation result is a credible annotation result; or If the confidence level is outside the target confidence level interval, it is determined that the first annotation result is not a credible annotation result.
9. The method according to claim 1, wherein The confidence level of the first annotation result is obtained in the following manner: Obtaining a generation probability of each word in the first annotation result by the first annotation model, where the generation probability is predicted by the first annotation model in a process of generating word words; as well as The average value of the sum of the generation probabilities of the word-grams in the first tagging result is used as the confidence level of the first tagging result.
10. The method according to claim 1, wherein The method further comprises: If the first annotation result is not a credible annotation result, the target data is sent to an annotation person for annotation of the first project to obtain a second annotation result, and the second annotation result is used as the annotation result of the target data on the first project.
11. The method according to claim 1, wherein The target data includes: question information, and answer information generated by the target model in response to the question information; The labeling result of the target data on the first project is used to evaluate the question-answering capability of the target model.
12. The method according to claim 1, wherein The using the first annotation model to annotate the target data on the first item includes: generating guidance instructions based on the first item and the target data; and The guiding instruction is input into the first labeling model to guide the first labeling model to label the target data on the first item.
13. A method for training and evaluating the capabilities of a labeling model, comprising: Obtaining a first data set and a second data set corresponding to a first project, wherein the first data set includes a plurality of first sample data and a sample labeling result of each first sample data on the first project, and the second data set includes a plurality of second sample data and a sample labeling result of each second sample data on the first project; training a first labeling model based on the first data set, so that the first labeling model has labeling capabilities for the first item; Annotating each second sample data in the second data set using the first annotation model to obtain a predicted annotation result of each second sample data on the first item and a confidence level corresponding to the predicted annotation result; Based on the predicted labeling results corresponding to each second sample data and their confidence levels, as well as the sample labeling results corresponding to each second sample data, the labeling capability information of the first labeling model for the first project is determined, wherein the labeling capability information represents a confidence interval. When the confidence level of the labeling results output by the first labeling model for the first project falls within the confidence interval, the labeling accuracy of the first labeling model is greater than or equal to the target accuracy.
14. The method according to claim 13, wherein The step of training a first annotation model based on the first data set includes: Get the base model; Training the base model based on the first data set to obtain multiple candidate models; and The accuracy of the multiple candidate models is evaluated using the second data set, and the candidate model with the highest accuracy among the multiple candidate models is used as the first labeled model.
15. The method according to claim 13, wherein The first item corresponds to multiple candidate annotation contents. The predicted annotation result is one of the multiple candidate annotation contents; The annotation capability information includes: sub-capability information corresponding to each of the plurality of candidate annotation contents, wherein the sub-capability information corresponding to each candidate annotation content represents a confidence interval applicable to the candidate annotation content.
16. The method according to claim 15, wherein The determining, based on the predicted labeling results and confidence levels corresponding to the respective second sample data and the sample labeling results corresponding to the respective second sample data, the labeling capability information of the first labeling model for the first project includes: Based on the predicted labeling results corresponding to the respective second sample data, the second data set is divided into a plurality of subsets, each subset including a plurality of second sample data, and the predicted labeling results corresponding to the respective second sample data in the same subset correspond to the same candidate labeling content; For each episode, the predicted labeling result corresponding to the episode is used as the target candidate labeling content, and based on the predicted labeling result corresponding to each second sample data in the episode and its confidence, as well as the sample labeling result corresponding to each second sample data in the episode, the sub-capability information corresponding to the target candidate labeling content is determined.
17. The method according to claim 16, wherein The sub-capability information corresponding to each candidate annotation content includes confidence intervals corresponding to multiple scenarios. The determining of the sub-capability information corresponding to the target candidate annotation content based on the predicted annotation results and confidence intervals corresponding to each second sample data in the subset, and the sample annotation results corresponding to each second sample data in the subset, includes: Determining a scene corresponding to each second sample data in the set; Dividing the set into a plurality of scene subsets, each scene subset including a plurality of second sample data in the set, and each second sample data in the same scene subset corresponds to the same scene; For each scene subset, the scene corresponding to the scene subset is taken as the target scene, and the confidence interval corresponding to the target scene is determined based on the predicted labeling results and their confidence corresponding to each second sample data in the scene subset, and the sample labeling results corresponding to each second sample data in the scene subset.
18. The method according to claim 13, wherein Determining the labeling capability information of the first labeling model for the first project based on the predicted labeling results and confidence levels corresponding to the respective second sample data and the sample labeling results corresponding to the respective second sample data includes: Determining an initial confidence interval, where a lower limit of the initial confidence interval is a preset first confidence level and an upper limit is a preset second confidence level, and the second confidence level is greater than the first confidence level; The lower limit of the initial confidence interval is increased according to a specified step size to obtain a target confidence interval until it is determined that the target confidence interval meets a preset condition, wherein the preset condition includes: the accuracy of the labeling result of each second sample data falling within the target confidence interval is greater than or equal to the target accuracy: The annotation capability information is generated based on the target confidence interval.
19. The method according to claim 18, wherein The method further comprises: Obtaining a third data set corresponding to the first project, the third data set including a plurality of third sample data and sample labeling results of each third sample data on the first project, Annotating each third sample data in the third data set using the first annotation model to obtain a predicted annotation result of each third sample data on the first item and a confidence level corresponding to the predicted annotation result; The preset condition also includes: the accuracy of the labeling results of each third sample data in the third data set that falls within the target confidence interval is greater than or equal to the target accuracy.
20. A data annotation system comprising: at least one storage medium storing at least one instruction set for data labeling; as well as At least one processor is communicatively connected to the at least one storage medium, wherein when the data labeling system is running, the at least one processor reads the at least one instruction set and executes the data labeling method according to any one of claims 1-12 according to the instructions of the at least one instruction set.
21. A system for training and evaluating the capabilities of a labeling model, comprising: At least one storage medium storing at least one instruction set for training and capability evaluation of a labeling model; as well as At least one processor is communicatively connected to the at least one storage medium, wherein when the training and capability evaluation system of the annotation model is running, the at least one processor reads the at least one instruction set and executes the training and capability evaluation method of the annotation model as described in any one of claims 13 to 19 according to the instructions of the at least one instruction set.