Data annotation method and system based on large language model

By adopting a data labeling method based on large language models in the field of natural language processing, the data labeling process is automated, and the problems of low manual labeling efficiency and insufficient accuracy are solved, efficient and accurate data labeling is achieved, and cost is reduced.

CN119961605APending Publication Date: 2025-05-09INSPUR SOFTWARE CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510046402.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The manual data labeling process in the existing natural language processing field is inefficient, long cycles, and high cost. Relying on manpower leads to inconsistent labeling results and insufficient accuracy.

Method used

The data annotation method based on a large language model is adopted to realize automatic annotation and optimization of data through steps such as data preparation, model selection and loading, automatic annotation, manual verification and correction, integrated output of the annotation result and model optimization iteration.

Benefits of technology

It improves the efficiency and accuracy of data labeling, reduces labor costs, ensures consistency and high quality of labeling results, and can quickly adjust and adapt to specific scenario needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961605A_ABST
    Figure CN119961605A_ABST
Patent Text Reader

Abstract

The invention discloses a data annotation method and system based on a large language model, and belongs to the technical field of deep learning natural language processing, and the method comprises the following steps: data preparation: collecting and preprocessing original data; model selection and loading: selecting and loading a suitable large model, wherein the large model is used for carrying out preliminary labeling on the preprocessed data; automatic annotation: automatically annotating the preprocessed data by using the loaded large model, and filtering and optimizing a preliminary annotation result; manual verification and correction are carried out, wherein the result of automatic labeling is verified, and wrong labeling is corrected; annotation result integration and output, wherein the corrected annotation results are integrated, and an annotation data set is output in a proper format; and model optimization and iteration: performing optimization and iteration on the large model according to an annotation result. According to the method, automatic labeling of the data can be realized, so that time and labor cost are saved; and meanwhile, the accuracy and reliability of data annotation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning natural language processing, and specifically to a data annotation method and system based on a large language model. Background Art

[0002] In the field of natural language processing (NLP), manual data annotation is a key step in achieving machine learning and deep learning model training. The existing implementation scheme mainly includes the following steps. First, data collection is the starting point of manual data annotation. Collect data related to the target task from various text resources, such as social media text, news articles, user comments, etc. Secondly, data preprocessing is the key to ensuring data quality. This includes steps such as text cleaning (removing HTML tags, special characters, meaningless words, etc.), word segmentation, and part-of-speech tagging. The next step is manual annotation. For specific NLP tasks, such as named entity recognition (NER), sentiment analysis, and text classification, professional annotators perform detailed annotations on the text. For example, in the NER task, the annotator needs to identify and annotate entities such as names of people, places, and organization names in the text; in the sentiment analysis task, the annotator needs to judge the emotional tendency expressed by the text, such as positive, negative, or neutral. Finally, data verification and quality control are important steps to ensure the accuracy of the annotated data. Through cross-verification, expert review, etc., the quality of the annotated data is evaluated to ensure that the data can meet the needs of model training. Taking named entity recognition as an example, in order to improve the performance of search and recommendation systems, an e-commerce platform collected a large amount of user comment data and hired professional annotators to annotate entities such as names, product names, and brand names in the comments. After preprocessing, annotation, and verification, the data was used to train the NER model, thereby improving the accuracy and efficiency of the search and recommendation systems. However, the entire process is highly dependent on manpower and time, resulting in problems such as low efficiency, long cycle, and high cost. Summary of the invention

[0003] The technical task of the present invention is to address the above shortcomings and provide a data annotation method and system based on a large language model, which can realize automatic data annotation, thereby saving time and labor costs; at the same time, it improves the accuracy and reliability of data annotation.

[0004] The technical solution adopted by the present invention to solve its technical problem is:

[0005] A data annotation method based on a large language model, the implementation of the method includes the following steps:

[0006] 1) Data preparation: collect raw data and preprocess them;

[0007] 2) Model selection and loading: Select and load a suitable large model, which is used to perform preliminary annotation on the preprocessed data;

[0008] 3) Automatic labeling: Use the loaded large model to automatically label the preprocessed data, and filter and optimize the preliminary labeling results;

[0009] 4) Manual verification and correction: Verify the results of automatic annotation and correct incorrect annotations;

[0010] 5) Integration and output of annotation results: Integrate the corrected annotation results and output the annotation dataset in an appropriate format;

[0011] 6) Model optimization and iteration: If the annotation results still do not meet the requirements, the large model is optimized and iterated according to the annotation results to improve its annotation accuracy and efficiency; then steps 1)-5) are repeated until satisfactory annotation results are obtained.

[0012] Furthermore, the data preparation specifically includes:

[0013] Raw data collection, including text, images, audio, video or point cloud data;

[0014] Preprocessing involves necessary cleaning, format conversion, and enhancement of raw data to ensure data quality and consistency.

[0015] Furthermore, the model selection loading specifically includes:

[0016] Select a suitable large model: According to the requirements of the data labeling task, select a suitable pre-trained large model, including ernie, GPT series, image large model, etc.;

[0017] Load pre-trained model: Load the selected pre-trained model into the annotation system to take advantage of its powerful feature extraction and generalization capabilities.

[0018] Furthermore, the automatic labeling specifically includes:

[0019] Based on the characteristics of the task, write prompt instructions to direct the large model and use the large model's strong understanding ability to perform preliminary batch labeling;

[0020] Filtering and optimization of annotation results: Filter and optimize the preliminary annotation results to remove obviously erroneous annotations and improve the accuracy of annotations.

[0021] Furthermore, the automatic annotation outputs the result in json format for easy extraction.

[0022] Furthermore, the manual verification and correction specifically includes:

[0023] Manual verification: Professional annotators verify the results of automatic annotation to ensure the accuracy and consistency of annotation;

[0024] Correcting incorrect annotations: Errors or inaccurate annotations that occur during automatic annotation are manually corrected by the annotator.

[0025] Furthermore, the annotation results are integrated and output, specifically including:

[0026] Integrate annotation results: Integrate the manually verified and corrected annotation results to form the final annotation data set;

[0027] Output labeled data: Output the labeled dataset in an appropriate format for subsequent model training or application.

[0028] The present invention also claims protection for a data annotation system based on a large language model, comprising:

[0029] Data preparation module, used to collect raw data and perform preprocessing;

[0030] A model selection and loading module is used to select and load a suitable large model, and the large model is used to perform preliminary annotation on the preprocessed data;

[0031] The automatic labeling module uses the loaded large model to automatically label the preprocessed data and filter and optimize the preliminary labeling results;

[0032] The manual verification and correction module is used to verify the results of automatic annotation and correct the incorrect annotations;

[0033] The annotation result integration and output module is used to integrate the corrected annotation results and output the annotation data set in an appropriate format;

[0034] The model optimization and iteration module is used to optimize and iterate the large model according to the annotation results when the annotation results do not meet the requirements until a satisfactory annotation result is obtained;

[0035] The system specifically implements data labeling through the above-mentioned data labeling method based on the large language model.

[0036] The present invention also claims a data annotation device based on a large language model, comprising at least one memory and a processor;

[0037] The at least one memory is used to store a machine-readable program;

[0038] The at least one processor is used to call the machine-readable program to implement the above method.

[0039] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which implement the above method when executed by a processor.

[0040] Compared with the prior art, the data annotation method and system based on a large language model of the present invention has the following beneficial effects:

[0041] First, with its powerful generalization ability and knowledge reserve, the large model can achieve in-depth understanding and analysis of various data contents such as audio, text, images, and point clouds. This enables the data annotation process to more efficiently and accurately identify key information in the data, and then perform high-precision annotation.

[0042] Secondly, big model-assisted data labeling improves labeling efficiency. Traditional data labeling mainly relies on manual work, which has problems such as low efficiency, long cycle, and high cost. With the help of the automatic labeling capability of big models, data can be processed quickly, which can significantly improve labeling efficiency and reduce labor costs. In addition, big model-assisted data labeling also ensures the consistency and accuracy of labeling results. In the process of data labeling, inconsistent or incorrect labeling results often occur due to subjective bias of labelers or improper data processing. Big models are based on deep learning and advanced algorithm technology, which can reduce the impact of these subjective factors and ensure the consistency and accuracy of labeling results.

[0043] Finally, large models can also assist in data annotation and can be quickly optimized for specific scenarios. For example, in the fields of autonomous driving and medical imaging, large models can be quickly optimized according to actual needs to achieve accurate annotation of specific targets or scenarios, thus meeting the needs of actual applications.

[0044] In summary, large-scale model-assisted data annotation has produced significant technical effects in the field of artificial intelligence, including improving annotation efficiency, ensuring the accuracy and consistency of annotation results, and quickly optimizing for specific scenarios. The realization of these technical effects provides strong support for the further development of artificial intelligence technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a schematic diagram showing the implementation principle of the data annotation method based on the large language model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The embodiment of the present invention provides a data annotation method based on a large language model.

[0047] A data annotation method based on a large language model, the implementation of the method includes the following steps:

[0048] 1. Data preparation: Collect raw data and preprocess them. Specifically include:

[0049] Raw data collection, including text, images, audio, video or point cloud data;

[0050] Preprocessing involves necessary cleaning, format conversion, and enhancement of raw data to ensure data quality and consistency.

[0051] 2. Model selection and loading: Select and load a suitable large model, which is used to perform preliminary annotation on the preprocessed data. Specifically, it includes:

[0052] Choose a suitable large model: According to the requirements of the data labeling task, choose a suitable pre-trained large model, such as ernie, GPT series, image large model, etc.;

[0053] Load pre-trained model: Load the selected pre-trained model into the annotation system to take advantage of its powerful feature extraction and generalization capabilities.

[0054] 3. Automatic labeling: Use the loaded large model to automatically label the preprocessed data, and filter and optimize the preliminary labeling results. Specifically include:

[0055] Based on the characteristics of the task, write prompt instructions to direct the large model and use the large model's strong understanding ability to perform preliminary batch labeling;

[0056] Output the results in json format for easy extraction;

[0057] Filtering and optimization of annotation results: Filter and optimize the preliminary annotation results to remove obviously erroneous annotations and improve the accuracy of annotations.

[0058] 4. Manual verification and correction: Verify the results of automatic marking and correct incorrect markings.

[0059] Specifically include:

[0060] Manual verification: Professional annotators verify the results of automatic annotation to ensure the accuracy and consistency of annotation;

[0061] Correcting incorrect annotations: Errors or inaccurate annotations that occur during automatic annotation are manually corrected by the annotator.

[0062] 5. Integration and output of annotation results: Integrate the corrected annotation results and output the annotation dataset in an appropriate format. Specifically include:

[0063] Integrate annotation results: Integrate the manually verified and corrected annotation results to form the final annotation data set;

[0064] Output labeled data: Output the labeled dataset in an appropriate format for subsequent model training or application.

[0065] 6. Model optimization and iteration: If the annotation results still do not meet the requirements, the large model can be optimized and iterated based on the annotation results to improve its annotation accuracy and efficiency; then repeat steps 1-5 until a satisfactory annotation result is obtained.

[0066] This method significantly improves the efficiency and accuracy of data labeling through large models. Traditional data labeling relies on manual labeling by human labelers, which is inefficient and has a long cycle and cannot meet the needs of high-dimensional data. Large models such as ChatGPT and other pre-trained large models can realize automatic labeling of data, thereby saving time and labor costs. For example, the application of Manfu Technology's pre-trained large model in the automatic labeling algorithm of autonomous driving AI can improve efficiency by several to dozens of times compared with manual labeling, while greatly reducing data production costs. Secondly, the large model produces more targeted data for small models through methods such as automatic labeling for small models to learn, reducing the requirements for downstream task data labeling costs, and further reducing development and iteration costs. This not only speeds up the iteration speed of artificial intelligence technology, but also promotes its application and development in various industries. In addition, large models can also fuse multimodal data, effectively integrate source data such as NLP, vision, and speech, achieve the effect of 1+1>2, and further improve the knowledge completeness of AI models. This multimodal data fusion capability enables data labeling to more comprehensively reflect the characteristics and attributes of data, and improves the accuracy and reliability of data labeling. In summary, large-model assisted data labeling solves several key problems in the field of data labeling by improving labeling efficiency, reducing labeling costs, and enhancing model knowledge completeness, providing strong support for the further development of artificial intelligence technology.

[0067] An embodiment of the present invention further provides a data annotation system based on a large language model, and the system specifically implements data annotation through the data annotation method based on a large language model described in the above embodiment.

[0068] The system includes:

[0069] 1. Data preparation module, used to collect raw data and perform preprocessing:

[0070] Raw data collection, including text, images, audio, video or point cloud data;

[0071] Preprocessing involves necessary cleaning, format conversion, and enhancement of raw data to ensure data quality and consistency.

[0072] 2. Model selection and loading module, used to select and load a suitable large model, which is used to perform preliminary annotation on the preprocessed data:

[0073] Choose a suitable large model: According to the requirements of the data labeling task, choose a suitable pre-trained large model, such as ernie, GPT series, image large model, etc.;

[0074] Load pre-trained model: Load the selected pre-trained model into the annotation system to take advantage of its powerful feature extraction and generalization capabilities.

[0075] 3. Automatic labeling module, which uses the loaded large model to automatically label the preprocessed data, and filters and optimizes the preliminary labeling results:

[0076] Based on the characteristics of the task, write prompt instructions to direct the large model and use the large model's strong understanding ability to perform preliminary batch labeling;

[0077] Output the results in json format for easy extraction;

[0078] Filtering and optimization of annotation results: Filter and optimize the preliminary annotation results to remove obviously erroneous annotations and improve the accuracy of annotations.

[0079] 4. Manual verification and correction module, used to verify the results of automatic annotation and correct incorrect annotations:

[0080] Manual verification: Professional annotators verify the results of automatic annotation to ensure the accuracy and consistency of annotation;

[0081] Correcting incorrect annotations: Errors or inaccurate annotations that occur during automatic annotation are manually corrected by the annotator.

[0082] 5. The annotation result integration output module is used to integrate the corrected annotation results and output the annotation dataset in an appropriate format:

[0083] Integrate annotation results: Integrate the manually verified and corrected annotation results to form the final annotation data set;

[0084] Output labeled data: Output the labeled dataset in an appropriate format for subsequent model training or application.

[0085] 6. Model optimization and iteration module: When the annotation results do not meet the requirements, the large model is optimized and iterated according to the annotation results until a satisfactory annotation result is obtained.

[0086] An embodiment of the present invention further provides a data annotation device based on a large language model, comprising at least one memory and a processor;

[0087] The at least one memory is used to store a machine-readable program;

[0088] The at least one processor is used to call the machine-readable program to implement the data labeling method based on the large language model described in the above embodiment.

[0089] The embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor executes the data labeling method based on the large language model described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0090] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0091] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.

[0092] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0093] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0094] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A data annotation method based on a large language model, characterized in that: The implementation of this method includes the following steps: 1) Data preparation: collect raw data and preprocess them; 2) Model selection and loading: Select and load a suitable large model, which is used to perform preliminary annotation on the preprocessed data; 3) Automatic labeling: Use the loaded large model to automatically label the preprocessed data, and filter and optimize the preliminary labeling results; 4) Manual verification and correction: Verify the results of automatic annotation and correct incorrect annotations; 5) Integration and output of annotation results: Integrate the corrected annotation results and output the annotation dataset in an appropriate format; 6) Model optimization and iteration: If the annotation results still do not meet the requirements, optimize and iterate the large model based on the annotation results; then repeat steps 1)-5) until a satisfactory annotation result is obtained.

2. According to the data annotation method based on the large language model of claim 1, it is characterized in that: The data preparation specifically includes: Raw data collection, including text, images, audio, video, or point cloud data; Preprocessing: Perform necessary cleaning, format conversion, and enhancement on raw data to ensure data quality and consistency.

3. According to the data annotation method based on a large language model in claim 1, it is characterized in that: The model selection loading specifically includes: Select a suitable large model: According to the requirements of the data labeling task, select a suitable pre-trained large model, including ernie, GPT series, and image large model; Load pre-trained model: Load the selected pre-trained model into the annotation system.

4. The data annotation method based on a large language model according to claim 1, characterized in that: The automatic marking specifically includes: Based on the characteristics of the task, write prompt instructions to direct the large model and use the large model's strong understanding ability to perform preliminary batch labeling; Filtering and optimization of annotation results: Filter and optimize the preliminary annotation results to remove obviously erroneous annotations.

5. A data annotation method based on a large language model according to claim 4, characterized in that: The automatic annotation outputs the result in json format.

6. The data annotation method based on a large language model according to claim 1, characterized in that: The manual verification and correction specifically includes: Manual verification: Professional annotators verify the results of automatic annotation to ensure the accuracy and consistency of annotation; Correcting incorrect annotations: Errors or inaccurate annotations that occur during automatic annotation are manually corrected by the annotator.

7. The data annotation method based on a large language model according to claim 1, characterized in that: The integrated output of the annotation results specifically includes: Integrate annotation results: Integrate the manually verified and corrected annotation results to form the final annotation data set; Output labeled data: Output the labeled dataset in an appropriate format for subsequent model training or application.

8. A data annotation system based on a large language model, characterized in that: include: Data preparation module, used to collect raw data and perform preprocessing; A model selection and loading module is used to select and load a suitable large model, and the large model is used to perform preliminary annotation on the preprocessed data; The automatic labeling module uses the loaded large model to automatically label the preprocessed data and filter and optimize the preliminary labeling results; The manual verification and correction module is used to verify the results of automatic annotation and correct the incorrect annotations; The annotation result integration and output module is used to integrate the corrected annotation results and output the annotation data set in an appropriate format; The model optimization and iteration module is used to optimize and iterate the large model according to the annotation results when the annotation results do not meet the requirements until a satisfactory annotation result is obtained; The system specifically implements data labeling through the data labeling method based on a large language model as described in any one of claims 1 to 7.

9. A data annotation device based on a large language model, characterized in that: comprising at least one memory and a processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Large model labeling method based on data labeling rule

    CN120596535A

  • Data annotation overview method and device based on AI large model and medium

    CN121031819A

  • Multi-modal data labeling method and system based on large model pre-labeling

    CN121456489A

  • A Multimodal Data Labeling Method and System Based on Large Model Pre-labeling

    CN121456489B

  • Data labeling and processing method and system based on natural language model

    CN121481642A