Auxiliary labeling system and method based on large model
Through the auxiliary labeling system based on large models, combined with large model pre-labeling and manual labeling, the data labeling process is optimized, and the problems of high labor costs, low efficiency, low marking quality and inconsistent standards are solved, and efficient and high-quality data labeling is achieved.
Patent Information
- Application Number
- CN202510264092.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
There are problems in the current data labeling process with high labor costs, low efficiency, low marking quality and inconsistent standards, especially after pre-labeling of large models, it still requires a lot of manual review and correction.
The auxiliary labeling system based on large models is adopted, including corpus module, labeling instructor operation module, large model pre-labeling module, standardized document module, large model self-correction module and ordinary labeling operator operation module. The data labeling process is optimized through a combination of large model pre-labeling and manual labeling.
The data labeling process is optimized, the manual labeling cost is reduced, the labeling efficiency and quality is improved, and the uniformity and high quality of the labeling data are ensured.
Smart Images

Figure CN120196947A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data annotation, and particularly relates to an auxiliary annotation system and method based on a large model. Background Art
[0002] Data annotation is the basic work for machine learning and artificial intelligence model training, and its quality directly affects the model performance. However, the current data annotation process has the following problems: 1. High labor cost and low efficiency: Traditional data annotation mainly relies on manual annotators, and the annotation process is time-consuming and costly. 2. Low annotation quality and inconsistent standards: Due to the lack of unified annotation standards, the annotation results of the same data by different annotators may vary, resulting in uneven annotation quality. 3. Limitations of large model pre-annotation: Although large model pre-annotation can improve the annotation efficiency, a large amount of manpower is still required in the subsequent manual review and correction process, and the problems of high cost and low efficiency are not completely solved. Summary of the Invention
[0003] The present invention aims to provide an auxiliary annotation system and method based on a large model to optimize the data annotation process through the auxiliary annotation system based on the large model, reduce costs and increase efficiency of manual annotation, improve the efficiency of data annotation, and ensure the unity and high quality of the annotated data.
[0004] To achieve the above objectives, the present invention adopts the following technical solutions: Provide an auxiliary annotation system based on a large model, including: A corpus module for storing the original data to be annotated, the rough annotation corpus during the annotation process, and the refined annotation corpus obtained after the annotation is completed; An annotation instructor operation module for allowing the annotation instructor to design prompt words and optimize based on relevant data, perform manual quality control on the annotation results, and classify and summarize the problems generated during the annotation process; A large model pre-annotation module for calling the large model and pre-annotating the original data based on the prompt words designed by the annotation instructor to obtain the first rough annotation corpus; A standardized document module for the annotation instructor to formulate a unified standardized document for training ordinary annotators according to the results of classifying and summarizing the problems; A large model self-correction module for calling the large model and self-correcting the first rough annotation corpus based on the self-correction prompt words designed by the annotation instructor to obtain the second rough annotation corpus; Ordinary annotator operation module, which is used for ordinary annotators to carry out standardized training, manually pre-annotate the first rough corpus, and manually annotate the second rough corpus to obtain the refined corpus.
[0005] Preferably, the corpus module includes an original data unit, a first rough corpus unit, a second rough corpus unit, and a refined corpus unit. The original data unit is used to input and store the original data to be annotated. The first rough corpus unit is used to store the first rough corpus formed after pre-annotation by the large model for calling. The second rough corpus unit is used to store the second rough corpus formed after self-correction by the large model for calling. The refined corpus unit is used to store the refined corpus after manual annotation and passing the quality control.
[0006] Preferably, the annotation instructor operation module includes a prompt word design unit, a first manual quality control unit, and a first problem classification and induction unit. The prompt word design unit is used for the annotation instructor to design prompt words. The first manual quality control unit is used for the annotation instructor to conduct manual quality control on the designed prompt words. The first problem classification and induction unit is used for the annotation instructor to classify and summarize the problems in the first rough corpus after being annotated by the large model.
[0007] Preferably, the annotation instructor operation module further includes a second manual quality control unit, a third manual quality control unit, and a second problem classification and induction unit. The second manual quality control unit is used for the annotation instructor to conduct manual quality control on the first rough corpus after being annotated by the large model. The third manual quality control unit is used for the annotation instructor to conduct manual quality control on the corpus after manual pre-annotation. The second problem classification and induction unit is used for the annotation instructor to classify and summarize the problems in the corpus after passing the quality control of the third manual quality control unit.
[0008] Preferably, the annotation instructor operation module further includes a fourth manual quality control unit, which is used for the annotation instructor to conduct manual quality control on the corpus after manual annotation.
[0009] Preferably, the standardized document module includes a first standardized document unit and a second standardized document unit. The first standardized document unit is used to store the first annotation document formed after being processed by the first problem classification and induction unit for calling. The second standardized document unit is used to store the second standardized document formed after being processed by the second problem classification and induction unit for calling.
[0010] Preferably, the general annotator operation module includes a standardization training unit, a manual pre-annotation unit, and a manual annotation unit. The standardization training unit is used for general annotators to conduct standardization training according to the first standardized document. The manual pre-annotation unit is used for general annotators after standardization training to conduct manual pre-annotation on the first rough slogan corpus quality-controlled by the second manual quality control unit. The manual annotation unit is used for general annotators after standardization training to conduct manual annotation on the second rough slogan corpus.
[0011] The present invention also provides an auxiliary annotation method based on a large model. This method uses the above-mentioned auxiliary annotation system based on a large model. This method includes: Prompt design and optimization: The annotation instructor designs prompts based on the actual needs of the training corpus according to the original data, and optimizes the prompts according to the quality control results in the pre-annotation stage of the large model; Large model pre-annotation: Call the large model to pre-annotate the original data based on the designed prompts to form the first rough slogan corpus; Annotation result quality control and problem analysis: In the pre-annotation stage of the large model, the annotation instructor conducts quality control on the results of the first rough slogan corpus generated by the large model pre-annotation, and summarizes and classifies the problems; Standardized document formulation and general annotator training: The annotation instructor formulates a unified first standardized document according to the results of problem classification and induction, and trains general annotators based on this first standardized document; Manual pre-annotation stage: General annotators select and match the first rough slogan corpus formed by the large model pre-annotation according to the first standardized document; The annotation instructor further classifies and summarizes the preliminary results formed by the manual pre-annotation to form a second standardized document including error types; Large model self-correction: The annotation instructor assembles the problem types in the original data, the first rough slogan corpus, and the second standardized document into self-correction prompts; Call the large model to self-correct the first rough slogan corpus to form the second rough slogan corpus; Manual annotation stage: General annotators conduct fine annotation on the second rough slogan corpus according to the first standardized document and the second standardized document; The annotation instructor conducts quality control on the fine annotation results of general annotators, and after passing, forms the final fine slogan corpus.
[0012] Preferably, in the prompt design and optimization, the large model is standardized to output a JSON table structure and field descriptions by designing prompts and optimizing.
[0013] Preferably, in the annotation result quality control and problem analysis, when summarizing and classifying problems, by sampling the fine slogan corpus, the accurate value range and common error types of each field are formed, and a problem classification and induction document is generated.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: The large model-based auxiliary annotation system includes a corpus module, an annotation instructor operation module, a large model pre-annotation module, a standardized document module, a large model self-correction module, and a general annotator operation module. The corpus module is used to store the original data to be annotated, the rough annotation corpus during the annotation process, and the refined annotation corpus obtained after the annotation ends. The annotation instructor operation module is used for the annotation instructor to design prompt words and optimize based on relevant data, perform manual quality control on the annotation results, and classify and summarize the problems generated during the annotation process. The large model pre-annotation module is used to call the large model and pre-annotate the original data based on the prompt words designed by the annotation instructor to obtain the first rough annotation corpus. The standardized document module is used for the annotation instructor to formulate a unified standardized document for training general annotators according to the results of the classification and summary of the problems. The large model self-correction module is used to call the large model and self-correct the first rough annotation corpus based on the self-correction prompt words designed by the annotation instructor to obtain the second rough annotation corpus. The general annotator operation module is used for general annotators to receive standardized training, perform manual pre-annotation on the first rough annotation corpus, and perform manual annotation on the second rough annotation corpus to obtain the refined annotation corpus. The large model-based auxiliary annotation system is used for data annotation in the fields of machine learning, artificial intelligence, model training, etc. By introducing the combination of large model pre-annotation and manual annotation, the data annotation process is optimized, the annotation efficiency is improved, the cost is reduced, and at the same time, the high quality of the annotated data is ensured. This system is applicable to data annotation tasks in the fields of machine learning, artificial intelligence, etc., and has broad application prospects. Description of the Drawings
[0015] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is the overall architecture diagram of an embodiment of the large model-based auxiliary annotation system of the present invention.
[0016] Figure 2 It is the architecture diagram of the corpus module in an embodiment of the large model-based auxiliary annotation system of the present invention.
[0017] Figure 3 It is the architecture diagram of the annotation instructor operation module in an embodiment of the large model-based auxiliary annotation system of the present invention.
[0018] Figure 4 It is the architecture diagram of the standardized document module in an embodiment of the large model-based auxiliary annotation system of the present invention.
[0019] Figure 5This is the architecture diagram of the ordinary annotator operation module in an embodiment of the large model-based auxiliary annotation system of the present invention.
[0020] Figure 6 This is the flowchart of an embodiment of the large model-based auxiliary annotation method of the present invention. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] In one embodiment, a large model-based auxiliary annotation system is provided, as Figure 1 shown. The large model-based auxiliary annotation system includes a corpus module 100, an annotation instructor operation module 200, a large model pre-annotation module 300, a standardized document module 400, a large model self-correction module 500, and an ordinary annotator operation module 600. Among them, the corpus module 100 in the large model-based auxiliary annotation system is used to store the original data to be annotated, the rough annotation corpus during the annotation process, and the refined annotation corpus obtained after the annotation ends; the annotation instructor operation module 200 is used for the annotation instructor to design prompt words and optimize based on relevant data, perform manual quality control on the annotation results, and classify and summarize the problems generated during the annotation process; the large model pre-annotation module 300 is used to call the large model and perform pre-annotation on the original data based on the prompt words designed by the annotation instructor to obtain the first rough annotation corpus; the standardized document module 400 is used for the annotation instructor to formulate a unified standardized document to train ordinary annotators according to the results of classifying and summarizing the problems; the large model self-correction module 500 is used to call the large model and perform self-correction on the first rough annotation corpus based on the self-correction prompt words designed by the annotation instructor to obtain the second rough annotation corpus; the ordinary annotator operation module 600 is used for ordinary annotators to receive standardized training, perform manual pre-annotation on the first rough annotation corpus, and perform manual annotation on the second rough annotation corpus to obtain the refined annotation corpus.
[0023] In this large model-based auxiliary annotation system, the corpus module 100 can be implemented by a database to classify and store the original data to be annotated, the rough annotation corpus during the annotation process, and the refined annotation corpus obtained after the annotation ends, so that during the entire data annotation process, the management of different types of data is more standardized and clear.
[0024] The annotation instructor operation module 200 is used for the annotation instructor to perform relevant operations. The responsibilities of the annotation instructor include annotation data structure design, prompt design and optimization, quality control and problem classification, standard document formulation, and training of ordinary annotators. Among them, the annotation data structure design is to design the JSON table structure and field description of the standardized output result according to the actual needs of the training corpus, based on professional knowledge and regulatory documents, and realize the output of the JSON table structure and field description through prompts; the prompt design and optimization is to optimize the large model prompts according to the quality control results in the pre-annotation stage of the large model; the quality control and problem classification is to perform quality control on the roughly annotated data generated by the large model. For example, 20% of the roughly labeled corpus can be selected for manual fine annotation to form the accurate value range and common error types of each field, and generate an annotation problem classification and induction document. The accurate value range of each field here is the set of values that the field may contain, and the common error type is the common error values that each field may have; the standard document formulation is to form a standard document based on the pre-annotation results of the large model to provide a basis for ordinary annotators; the annotator training is to train ordinary annotators according to the standard document to ensure that ordinary annotators clearly understand the value range and definition of each field in the document.
[0025] The large model called by the large model pre-annotation module 300 is a large language model, and the first roughly labeled corpus is the standardized data structure output after pre-annotation by the large model, such as the JSON table structure and field description.
[0026] The responsibilities of ordinary annotators include learning the standard document and manual annotation. Among them, the learning of the standard document is to understand and master the value range and definition of each field in the document through the training of the annotation instructor. The annotation instructor trains ordinary annotators through the unified standard document in the standard document module 400; the manual annotation is to perform manual annotation on the first roughly labeled corpus pre-annotated by the large model and the second roughly labeled corpus after self-correction by the large model according to the standard document. Ordinary annotators perform relevant operations such as standard training, manual pre-annotation, and manual annotation through the ordinary annotator operation module 600. By calling the large model through the large model self-correction module 500 and performing self-correction on the first roughly labeled corpus according to the self-correction prompt words, the data correction efficiency can be improved, the annotation workload of ordinary annotators can be reduced, the quality inspection workload of the annotation instructor can be reduced, and the annotation error rate can be reduced, improving the annotation quality.
[0027] As can be seen from the above embodiments, the large model-based auxiliary annotation system optimizes the data annotation process by introducing the combination of large model pre-annotation and manual annotation, realizes the improvement of annotation efficiency and the reduction of costs, and at the same time ensures the high quality of the annotated data. This system is applicable to data annotation tasks in fields such as machine learning and artificial intelligence and has a wide range of application prospects.
[0028] Optionally, in combination withFigure 2 As shown in the figure, the previous corpus module 100 includes a raw data unit 101, a first rough-labeled corpus unit 102, a second rough-labeled corpus unit 103, and a refined-labeled corpus unit 104. The raw data unit 101 is used to input and store the raw data to be labeled. The first rough-labeled corpus unit 102 is used to store the first rough-labeled corpus formed after pre-labeling by the large model for subsequent call. The second rough-labeled corpus unit 103 is used to store the second rough-labeled corpus formed after self-correction by the large model for subsequent call. The refined-labeled corpus unit 104 is used to store the refined-labeled corpus that has passed the manual labeling and quality control. Storing different data at various stages of data labeling through the raw data unit 101, the first rough-labeled corpus unit 102, the second rough-labeled corpus unit 103, and the refined-labeled corpus unit 104 is conducive to the accurate classification of different data, facilitating viewing and calling.
[0029] Optionally, in combination with Figure 3 As shown in the figure, the previous annotation instructor operation module 200 includes a prompt word design unit 201, a first manual quality control unit 202, and a first problem classification and induction unit 203. The prompt word design unit 201 is used for the annotation instructor to design prompt words. The first manual quality control unit 202 is used for the annotation instructor to perform manual quality control on the designed prompt words. The first problem classification and induction unit 203 is used for the annotation instructor to classify and summarize the problems in the first rough-labeled corpus after being labeled by the large model.
[0030] On the one hand, the prompt word design unit 201 allows the annotation instructor to design prompt words based on the raw data. On the other hand, it enables the annotation instructor to assemble the problem types in the raw data, the first rough-labeled corpus, and the standardized document into self-correction prompt words. After the prompt words are designed, another annotation instructor performs manual quality control on the designed prompt words through the first manual quality control unit 202 to improve the consistency between the data results generated by the large model according to the prompt words and the target data results. The annotation instructor forms the accurate value range and common error types of each field in the form of sampling and refined annotation through the first problem classification and induction unit 203, and generates a problem classification and induction document to obtain the first standardized document. Sampling and refined annotation means sampling and refining the first rough-labeled corpus to summarize the common errors in the large model's annotation. The first standardized document records the errors that are likely to occur during annotation, such as the common incorrect annotation values that appear.
[0031] Optionally, in combination with Figure 3As shown, the aforementioned annotation instructor operation module 200 further includes a second manual quality control unit 204, a third manual quality control unit 205, and a second problem classification and summarization unit 206. The second manual quality control unit 204 is used for the annotation instructor to perform manual quality control on the first rough labeled corpus after being labeled by the large model. The third manual quality control unit 205 is used for the annotation instructor to perform manual quality control on the corpus after manual pre-annotation. The second problem classification and summarization unit 206 is used for the annotation instructor to classify and summarize the problems in the corpus after being quality controlled by the third manual quality control unit 205.
[0032] The first rough labeled corpus is pre-labeled by the large model and there may be labeling errors. The annotation instructor retrieves the first rough labeled corpus through the second manual quality control unit 204 to inspect the labeling quality and find the errors in the labeling. The corpus after manual pre-annotation is obtained by ordinary annotators selecting and matching the first rough labeled corpus after annotation training, and there may also be some errors, which require the annotation instructor to perform manual quality control through the third manual quality control unit 205. For the problems found by the annotation instructor during the manual quality control process through the third manual quality control unit 205, they can be classified and summarized through the second problem classification and summarization unit 206 to obtain the second standardized document.
[0033] Optionally, as shown in Figure 3 the previous annotation instructor operation module further includes a fourth manual quality control unit 207, which is used for the annotation instructor to perform manual quality control on the corpus after manual annotation. Performing manual quality control on the corpus after manual annotation is beneficial to improving the quality of the final refined labeled corpus.
[0034] Optionally, as shown in Figure 4 the previous standardized document module 400 includes a first standardized document unit 401 and a second standardized document unit 402. The first standardized document unit 401 is used to store the first annotation document formed after being processed by the first problem classification and summarization unit 203 for calling. The second standardized document unit 402 is used to store the second standardized document formed after being processed by the second problem classification and summarization unit 206 for calling.
[0035] The annotation instructor summarizes and induces multiple different first standardized documents and second standardized documents for different data annotation structures. These standard documents record different error types and record the value ranges and definitions of each field in the data annotation output format. Through the standardization of common errors and field value ranges, they can be used for unified and standard learning by different ordinary annotators. The first standardized document is for learning, and the second standardized document is used for the annotation instructor to design error correction prompt words and for ordinary annotators to perform refined labeling on the second rough labeled corpus.
[0036] Optionally, in combination with Figure 5 As shown, the previous general annotator operation module 600 includes a standardization training unit 601, a manual pre-annotation unit 602, and a manual annotation unit 603. The standardization training unit 601 is used for general annotators to conduct standardization training according to the first standardization document. The manual pre-annotation unit 602 is used for general annotators who have undergone standardization training to perform manual pre-annotation on the first rough slogan data after being quality controlled by the second manual quality control unit 204. The manual annotation unit 603 is used for general annotators who have undergone standardization training to perform manual annotation on the second rough slogan data.
[0037] Under the guidance of the annotation instructor, general annotators retrieve the first standardization document through the standardization training unit 601 for training. The manual pre-annotation unit 602 allows general annotators to perform pre-annotation to summarize the problems that are likely to occur in manual annotation. The manual annotation unit 603 is the final stage of refined annotation, that is, general annotators perform manual annotation on the second rough slogan data to obtain refined slogan data.
[0038] In another embodiment, a large model-based assisted annotation method is provided. This method uses the large model-based assisted annotation system in the previous embodiment, in combination with Figure 6 As shown, this method includes: (1) Prompt design and optimization: The annotation instructor designs prompts based on the actual requirements of the training data and optimizes the prompts according to the quality control results in the pre-annotation stage of the large model.
[0039] (2) Large model pre-annotation: The large model is called to perform pre-annotation on the original data based on the designed prompts to form the first rough slogan data.
[0040] (3) Quality control of annotation results and problem analysis: In the pre-annotation stage of the large model, the annotation instructor performs quality control on the results of the first rough slogan data generated by the large model pre-annotation and summarizes and classifies the problems.
[0041] (4) Standardization document formulation and general annotator training: The annotation instructor formulates a unified first standardization document based on the results of problem classification and induction, and trains general annotators according to this first standardization document.
[0042] (5) Manual pre-annotation stage: General annotators select and match the first rough slogan data formed by the large model pre-annotation according to the first standardization document. The annotation instructor further classifies and summarizes the problems based on the preliminary results formed by the manual pre-annotation to form a second standardization document containing error types.
[0043] (6) Self-correction of the large model: The annotation instructor assembles the problem types in the original data, the first rough labeled corpus, and the second standardized document into self-correction prompt words; calls the large model to perform self-correction on the first rough labeled corpus to form the second rough labeled corpus.
[0044] (7) Manual annotation stage: Ordinary annotators perform fine annotation on the second rough labeled corpus according to the first standardized document and the second standardized document; the annotation instructor performs quality control on the fine annotation results of the ordinary annotators, and after passing the quality control, the final fine labeled corpus is formed.
[0045] In this method, the responsibilities of the annotation instructor include designing prompt words, quality control analysis, classifying and summarizing problems, and formulating standardized documents. The responsibilities of ordinary annotators include learning and following the standardized documents, and completing the final annotation in combination with the rough annotation results. The difference between ordinary annotators and traditional annotators is that they do not rely on personal understanding, but select and match the rough annotation results according to the unified standardized documents, which is conducive to the consistency of data annotation.
[0046] Based on the above, this large model-based assisted annotation method optimizes the process and role division, reduces the repetitive labor of manual annotation, and reduces the overall annotation cost. The annotation instructor is responsible for formulating unified annotation standards to ensure the consistency and high quality of the annotation results, and avoid quality problems caused by inconsistent standards in traditional annotation. Therefore, this method can significantly improve the data annotation efficiency and reduce the annotation cost, because the large model pre-annotation reduces the workload of manual annotation, and the division of labor between the annotation instructor and ordinary annotators further improves the annotation efficiency.
[0047] Optionally, as shown in Figure 6 In the above prompt word design and optimization, the large model is standardized to output the JSON table structure and field descriptions by designing prompt words and optimization. The JSON table structure is a lightweight data exchange format, based on two basic structures of objects and arrays, with the advantages of high flexibility, good compatibility, and high transmission efficiency. Optionally, as shown in Figure 6 In the above quality control and problem analysis of the annotation results, when summarizing the problem classification and induction, by sampling the fine labeled corpus, the accurate value range and common error types of each field are formed, and a problem classification and induction document is generated. By sampling the fine labeled corpus, common error types can be discovered to a considerable extent under the premise of controllable workload.
[0048] Based on the above embodiments, this large model-based assisted annotation method optimizes the data annotation process by introducing the combination of large model pre-annotation and manual annotation, realizes the improvement of annotation efficiency and cost reduction, and at the same time ensures the high quality of the annotated data. This method is applicable to data annotation tasks in fields such as machine learning and artificial intelligence, and has a wide range of application prospects.
[0049] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0050] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An auxiliary annotation system based on a large model, characterized in that: include: The corpus module is used to store the original data to be annotated, the rough marked corpus in the annotation process, and the fine marked corpus obtained after the annotation is completed; A labeling instructor operation module, which is used for labeling instructors to design prompt words and adjust them based on relevant data, perform manual quality control on labeling results, and classify and summarize problems generated during the labeling process; A large model pre-annotation module, which is used to call the large model and pre-annotate the original data based on the prompt words designed by the annotation instructor to obtain the first rough labeled corpus; Standardized document module: This module is used by annotation instructors to develop unified standardized documents to train ordinary annotators based on the results of problem classification and induction; A large model self-correction module, the large model self-correction module is used to call the large model and perform self-correction on the first rough label material based on the self-correction prompt words designed by the annotation instructor to obtain a second rough label material; The ordinary annotator operation module is used for ordinary annotators to perform standardized training, manually pre-annotate the first coarse labeled corpus, and manually annotate the second coarse labeled corpus to obtain fine labeled corpus.
2. The large model-based auxiliary annotation system according to claim 1, characterized in that: The corpus module includes an original data unit, a first coarse label corpus unit, a second coarse label corpus unit, and a refined label corpus unit. The original data unit is used to input and store the original data to be annotated. The first coarse label corpus unit is used to store the first coarse label corpus formed after pre-annotation by the large model for call. The second coarse label corpus unit is used to store the second coarse label corpus formed after self-correction of the large model for call. The refined label corpus unit is used to store the refined label corpus after manual annotation and quality control.
3. The large model-based auxiliary annotation system according to claim 2, characterized in that: The annotation instructor operation module includes a prompt word design unit, a first manual quality control unit, and a first problem classification and summarization unit. The prompt word design unit is used for the annotation instructor to design prompt words, the first manual quality control unit is used for the annotation instructor to perform manual quality control on the designed prompt words, and the first problem classification and summarization unit is used for the annotation instructor to perform problem classification and summarization on the first rough annotation corpus after the large model is annotated.
4. The large model-based auxiliary annotation system according to claim 3, characterized in that: The labeling instructor operation module also includes a second manual quality control unit, a third manual quality control unit, and a second problem classification and summarization unit. The second manual quality control unit is used for the labeling instructor to perform manual quality control on the first rough label corpus after the large model is labeled, the third manual quality control unit is used for the labeling instructor to perform manual quality control on the corpus after manual pre-labeling, and the second problem classification and summarization unit is used for the labeling instructor to perform problem classification and summarization on the corpus after quality control by the third manual quality control unit.
5. The large model-based auxiliary annotation system according to claim 4, characterized in that: The annotation instructor operation module further includes a fourth manual quality control unit, and the fourth manual quality control unit is used for the annotation instructor to perform manual quality control on the manually annotated corpus.
6. The large model-based auxiliary annotation system according to claim 5, characterized in that: The standardized document module includes a first standardized document unit and a second standardized document unit. The first standardized document unit is used to store a first annotated document formed after being processed by the first problem classification and induction unit for call, and the second standardized document unit is used to store a second standardized document formed after being processed by the second problem classification and induction unit for call.
7. The large model-based auxiliary annotation system according to claim 6, characterized in that: The general labeler operation module includes a standardization training unit, a manual pre-labeling unit, and a manual labeling unit. The standardization training unit is used for general labelers to perform standardization training according to a first standardized document. The manual pre-labeling unit is used for general labelers who have received standardization training to perform manual pre-labeling on the first rough label material after quality control by the second manual quality control unit. The manual labeling unit is used for general labelers who have received standardization training to perform manual labeling on the second rough label material.
8. An auxiliary annotation method based on a large model, characterized in that: The method adopts the large model-based auxiliary annotation system described in any one of 1 to 7 above, and the method includes: Prompt word design and optimization: The annotation instructor designs prompt words based on the original data according to the actual needs of the training corpus, and optimizes the prompt words based on the quality control results of the large model pre-annotation stage; Large model pre-annotation: The large model is called to pre-annotate the original data based on the designed prompt words to form the first rough label corpus; Quality control of annotation results and problem analysis: During the pre-annotation stage of the large model, the annotation instructor conducts quality control on the first rough annotation corpus generated by the pre-annotation of the large model, and summarizes and classifies the problems; Standardized document preparation and general labeler training: Labeling instructors will prepare a unified first standardized document based on the problem classification and summarize the results, and train general labelers based on the first standardized document; Manual pre-labeling stage: ordinary labelers select and match the first rough label corpus formed by the large model pre-labeling according to the first standardized document; the labeling instructor further classifies and summarizes the problems based on the preliminary results formed by manual pre-labeling to form a second standardized document containing error types; Big model self-correction: The annotation instructor assembles the original data, the first rough label corpus, and the problem types in the second standardized document into self-correction prompt words; calls the big model to self-correct the first rough label corpus to form the second rough label corpus; Manual labeling stage: ordinary labelers perform fine labeling on the second rough labeled corpus based on the first standardized document and the second standardized document; labeling instructors perform quality control on the fine labeling results of ordinary labelers, and form the final fine labeled corpus after passing the quality control.
9. The large model-based auxiliary annotation method according to claim 8, characterized in that: In the prompt word design and tuning, prompt words are designed and tuned to standardize the output JSON table structure and field description of the large model.
10. The large model-based auxiliary annotation method according to claim 8, characterized in that: In the quality control and problem analysis of the annotation results, when summarizing the problem classification, the accurate value range and common error types of each field are formed by sampling the finely marked corpus, and a problem classification summary document is generated.
Citation Information
Cited By
Dynamic equilibrium labeling method and system for large visual model
CN121121744A
A Dynamic Equilibrium Annotation Method and System for Large Visual Models
CN121121744B