Method and system for automatically labeling data based on model training

By building a closed-loop annotation process and an automatic annotation model, the problems of low efficiency and high cost in traditional data annotation methods are solved, efficient and automated data annotation are achieved, and the quality and consistency of the annotation are improved.

CN120296223AActive Publication Date: 2025-07-11HANGZHOU ZHENZHI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510772821.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Traditional data annotation methods rely on manual completion, which is expensive, inefficient and can easily lead to inconsistent annotation, affecting the model training effect.

Method used

Through the automatic data annotation method based on model training, including target content information collection, information screening and preliminary evaluation, screening data proofreading and feature feedback, differential feature extraction and labeling preparation, automatic labeling and error rate verification and update, a closed-loop annotation process is constructed, and dynamic feedback and error rate tracking is used to combine with automatic labeling model segmentation processing.

Benefits of technology

It improves the accuracy and automation of data processing, reduces the cost of manual intervention, improves the efficiency and consistency of labeling, and ensures the stability and robustness of labeling results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296223A_ABST
    Figure CN120296223A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of specific calculation models, and particularly relates to a method and a system for automatically labeling data based on model training. The method comprises the following steps: collecting target content information; performing screening according to the information timeliness and the relevancy to obtain first screening data; checking the first screening data, generating historical information and feeding back the historical information; extracting difference features as to-be-labeled data, and automatically labeling the to-be-labeled data; checking and updating the labeling result according to the historical error rate; the deviation degree is calculated, a monitoring evaluation interval is generated, and an output result is classified as frozen data or unfrozen data; the system comprises an acquisition module, a screening module, a proofreading module, a labeling module and a deviation degree analysis module. According to the scheme, the labeling efficiency and accuracy are improved, and the method is suitable for a multi-modal data processing scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of specific computing models, and particularly relates to a method and system for automatic data annotation based on model training. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, machine learning, and natural language processing, training high-precision intelligent models has put forward higher requirements for large-scale and high-quality annotated data. Especially in data-driven tasks such as image recognition, semantic understanding, and speech recognition, the quality of annotated data directly determines the performance of the model. However, traditional data annotation methods mainly rely on manual labor, which is not only costly and inefficient, but also prone to inconsistent annotations due to subjective judgment, thus affecting the model training effect.

[0003] Therefore, there is an urgent need to propose a method and system for automatic data annotation that can combine the model training process, automatically generate annotation information, and continuously optimize the annotation quality and model performance through a feedback mechanism, so as to reduce the dependence on manual labor, improve the annotation efficiency and consistency, and provide a high-quality and scalable data foundation for intelligent models. Summary of the Invention

[0004] In view of the above problems, the object of the present invention is to propose: A method for automatic data annotation based on model training, including: S1. Acquisition of target content information: Using an acquisition model to acquire target content information; wherein, the target content information includes at least one of text, images, and videos; The acquisition model is a RAG large language model, which performs semantic expansion according to the annotation tags input by the user to obtain an extended vocabulary similar in semantics to the annotation tags, and performs keyword retrieval on the text, images, and videos of the acquisition website based on the extended vocabulary. The retrieved text, images, and videos constitute the target content information; S2. Information screening and preliminary evaluation: Screening the target content information according to the first evaluation condition to obtain the first screened data; wherein, the first evaluation condition includes information timeliness and information relevance; S3. Proofreading of screened data and feature feedback: Performing proofreading processing on the first screened data to obtain historical information, and feeding it back to the subsequent target content information acquisition stage; wherein, the historical information includes proofread information and unproofread information; S4. Extraction of differential features and annotation preparation: Performing overflow screening on the historical information to obtain differential information, and calibrating the differential information as data to be annotated; S5. Automatic annotation and error rate verification and update: Performing segmentation processing on the data to be annotated; The segment processing specifically refers to: based on the annotation tags input by the user, with the hit point of the data to be annotated as the center, extending a necessary length forward and backward to obtain a semantically complete piece of text, an image, or a video segment, while truncating the text, image, or video segment containing non-annotation information; After segment processing, multiple segments to be annotated are obtained, and then the segments to be annotated are input into the automatic annotation model one by one to obtain the annotated data, and the annotated data is verified according to the historical information error rate, and the error rate is updated according to the verification result; S6. Deviation measurement and monitoring interval generation: Measure the deviation of the updated error rate, then generate a monitoring and evaluation interval based on the deviation, and update the error rate during the next data annotation, and classify the output result of the automatic annotation model into frozen data and unfrozen data according to the instantaneous value of the updated error rate and the monitoring and evaluation interval.

[0005] As a preferred solution, the S1 specifically includes: constructing a list of collection websites and issuing collection instructions to each collection website in the list of collection websites; calling the interfaces of the collection websites, accessing each collection website, and collecting output information of the collection websites to obtain target content information.

[0006] As a preferred solution, the S2 specifically includes: first evaluating the timeliness of each target content information according to the information timeliness in the first evaluation condition, and eliminating the data that does not meet the conditions; the information timeliness is the time interval of the target information set by the user; then calculating the information relevance of the remaining data, and screening out the content with a relevance not lower than the threshold as the first screened data; the information relevance is the semantic cosine value between the target information and the annotation tags input by the user.

[0007] As a preferred solution, the S3 includes: constructing a list of information to be matched and adding each of the first screened data to the list; using a matching model to match each piece of information one by one; if the match is successful, pre-marking is synchronously executed, if the match fails, continue to proofread, and optimize and update the historical information according to a preset strategy.

[0008] As a preferred solution, the historical information includes first historical information, which represents the data in the first screened data that remains consistent with the first screened data after proofreading; second historical information, which represents the data in the first screened data that is found to be inconsistent with the first screened data after proofreading; third historical information, which represents the newly added data that has not been matched with the historical annotation tags.

[0009] As a preferred solution, the error rate update process in S5 includes: constructing a list of tag information to be verified, extracting the update time and model input of each piece of tag information; determining unlabeled or incorrect tags based on the matching degree with historical labeled tags, setting the total number of labeling errors, calculating the deviation rate, and dynamically updating the error rate according to the unlabeling frequency, and incorporating the updated result back into the historical information.

[0010] As a preferred solution, the deviation measurement in S6 includes: dividing the updated error rate into multiple segment values in the time dimension, performing discretization preprocessing and sorting, and inputting them into the deviation measurement model in combination with the set smoothing coefficient to output the deviation result.

[0011] As a preferred solution, the method for generating the monitoring and evaluation interval includes: inputting the error rate into the verification model, obtaining verification information and performing wavelet transform and feature sorting, then performing first-order difference and inverse normalization processing, calculating the risk value based on the update bias and the set risk threshold, and generating the high interval value and the low interval value accordingly, and constructing the final monitoring and evaluation interval.

[0012] The present invention also provides a system for automatic data labeling based on model training for implementing the above method, including: Information collection module: used to construct a list of collection websites, issue collection instructions, and call interfaces to implement the capture of target content information such as text, images, and videos.

[0013] Information screening module: used to evaluate and process the target content information according to the timeliness and relevance of the information, remove the data that does not meet the conditions, and output the first screened data.

[0014] Data verification module: used to call the matching model to compare the first screened data with the historical information, generate verified and unverified information, and feedback the comparison result to the collection model.

[0015] Data overflow screening module: used to extract difference information from the historical information, identify the change trend, and label it as data to be labeled for use by the automatic labeling model.

[0016] Automatic labeling module: used to input the data to be labeled into the model in segments for labeling, and verify and correct the labeling result in combination with the historical error rate, and dynamically update the error rate parameter.

[0017] Deviation analysis module: used to perform deviation measurement on the updated error rate, generate a monitoring and evaluation interval, and classify the status of the model output result according to the deviation degree to realize the dynamic control of the labeling quality.

[0018] The present invention also provides an electronic device, including at least one processor and a memory communicatively connected to the processor. A computer program executable by the processor is stored in the memory, and the program is used to execute the method described above.

[0019] Advantageous effects: The present invention provides a method and system for automatic annotation of data based on model training and its application in an electronic device, having the following advantageous effects: From the perspective of the method: The present invention constructs a five-level closed-loop annotation process of "acquisition - screening - proofreading - annotation - evaluation", makes full use of historical information for dynamic feedback and error rate tracking, combines differential feature screening with segmented processing of the automatic annotation model, and improves the accuracy and automation of data processing. After annotation, an error rate verification and deviation monitoring mechanism is introduced, which can dynamically adjust the output quality of the model, effectively classify the annotation results, enhance the stability and robustness of the overall method, and significantly reduce the cost of manual intervention.

[0020] From the perspective of the system: The system structure of the present invention is clear, the functions of each module are clearly defined, covering the whole process of information acquisition, screening, proofreading, annotation and monitoring evaluation. The automatic annotation module and the deviation analysis module in the system work together, can implement closed-loop control and risk management for the annotated data, and ensure the continuous and stable operation of the data processing process. At the same time, the system has highly modular characteristics, is easy to deploy and expand, and adapts to data processing requirements of different scales and complexities.

[0021] From the perspective of the application: The present invention can be widely applied to multi-modal data scenarios such as text understanding, image recognition, video processing, etc., and has good generalization ability and practical feasibility in the big data environment. By embedding the method into the processor and memory modules in the electronic device, the real-time processing ability of collecting and annotating simultaneously can be achieved, meeting the integration requirements of various devices such as smart terminals, server clusters, industrial acquisition platforms, etc. This technology effectively solves the problems existing in traditional annotation such as low efficiency, high error rate, and untimely model update, and has significant promotion value in the fields of artificial intelligence, knowledge graph construction, search engine training, etc. Description of the Drawings

[0022] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Detailed Embodiments

[0023] To deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.

[0024] Embodiment 1 As Figure 1As shown in the figure, this embodiment provides a method for automatic annotation of data based on model training, which specifically includes: S1. Acquisition of target content information: Use an acquisition model to acquire target content information; where the target content information includes at least one of text, images, and videos; specifically includes: constructing a list of acquisition websites, and sending acquisition instructions to each acquisition website in the list of acquisition websites; calling the interfaces of the acquisition websites, accessing each acquisition website, and collecting output information from the acquisition websites to obtain target content information.

[0025] S2. Information screening and preliminary evaluation: Screen the target content information according to the first evaluation condition to obtain the first screened data; where the first evaluation condition includes information timeliness and information relevance; specifically includes: first, evaluate the timeliness of each target content information according to the information timeliness in the first evaluation condition, and eliminate the data that does not meet the conditions; the information timeliness is the time interval of the target information set by the user; then calculate the information relevance of the remaining data, and screen out the content with a relevance not lower than the threshold as the first screened data; the information relevance is the semantic cosine value between the target information and the annotation label input by the user.

[0026] S3. Proofreading of screened data and feature feedback: Proofread the first screened data to obtain historical information, and feedback it to the subsequent target content information acquisition stage; where the historical information includes proofread information and unproofread information; includes: constructing a list of information to be matched, and adding each of the first screened data to the list; using a matching model to match each piece of information one by one; if the match is successful, perform pre-labeling synchronously, if the match fails, continue to proofread, and optimize and update the historical information according to a preset strategy; the historical information includes: the first historical information, which represents the data that remains consistent with the first screened data after proofreading in the first screened data; the second historical information, which represents the data that is found to be inconsistent with the first screened data after proofreading in the first screened data; the third historical information, which represents the newly added data that has not been matched with the historical annotation label.

[0027] S4. Extraction of differential features and annotation preparation: Perform overflow screening on the historical information to obtain differential information, and label the differential information as data to be annotated. S5. Automatic annotation and error rate verification and update: Perform segmentation processing on the data to be annotated. The segmentation processing specifically refers to: based on the annotation label input by the user, with the hit point of the data to be annotated as the center, extending forward and backward by a necessary length to obtain a semantically complete piece of text, an image, or a video segment, and truncating the text, image, or video segment containing non-annotation information. After segmentation processing, multiple segments to be labeled are obtained. Then, the segments to be labeled are input into the automatic labeling model one by one to obtain the labeled data. The labeled data is verified according to the error rate of historical information, and the error rate is updated according to the verification result. The error rate update process includes: constructing a list of label information to be verified, extracting the update time and model input of each label information; determining unlabeled or incorrect labels according to the matching degree with historical labeled labels, setting the total number of labeling errors, calculating the deviation rate, and dynamically updating the error rate according to the unlabeled frequency. The updated result is re-incorporated into the historical information.

[0028] S6. Deviation measurement and monitoring interval generation: Measure the deviation of the updated error rate, then generate a monitoring and evaluation interval according to the deviation. And update the error rate during the next data labeling, and classify the output result of the automatic labeling model into frozen data and unfrozen data according to the instantaneous value of the updated error rate and the monitoring and evaluation interval. The deviation measurement includes: dividing the updated error rate into multiple segment values in the time dimension, performing discretization preprocessing and sorting, and inputting the smoothed coefficient into the deviation measurement model to output the deviation result. The method for generating the monitoring and evaluation interval includes: inputting the error rate into the verification model, obtaining verification information and performing wavelet transform and feature sorting, then performing first-order difference and inverse normalization processing, calculating the risk value according to the updated deviation degree and the set risk threshold, and generating the high interval value and the low interval value accordingly to construct the final monitoring and evaluation interval.

[0029] Embodiment 2 This embodiment provides a system for automatic data labeling based on model training, including: An information collection module, which is used to collect target content information by using a collection model. Among them, the target content information includes at least one of text, image and video; An information screening module, which is used to screen the target content information according to the first evaluation condition to obtain the first screened data. Among them, the first evaluation condition includes the timeliness and relevance of the information; A data proofreading module, which is used to proofread the first screened data to obtain historical information and feedback it to the subsequent target content information collection stage. Among them, the historical information includes the proofread information and the unproofread information; A data overflow screening module, which is used to screen the historical information for overflow to obtain difference information and label the difference information as data to be labeled; An automatic annotation module, which is used to segment the data to be annotated, obtain multiple segments to be annotated, and input the segments to be annotated into the automatic annotation model one by one to obtain the annotated data. The error rate of the annotated data is verified based on the historical information error rate, and the error rate is updated according to the verification result; A deviation analysis module, which is used to measure the deviation of the updated error rate, generate a monitoring and evaluation interval based on the deviation, update the error rate during the next data annotation, and classify the output result of the automatic annotation model into frozen data and unfrozen data according to the instantaneous value of the updated error rate and the monitoring and evaluation interval.

[0030] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatic data annotation based on model training, characterized in that, include: S1. Target content information collection: According to the annotation tags input by the user, the target content information is collected using the RAG large language model as the collection model; wherein the target content information includes at least one of text, image and video; S2, information screening and preliminary evaluation: screening the target content information according to the first evaluation condition to obtain first screening data; wherein the first evaluation condition includes information timeliness and information relevance; S3, screening data proofreading and feature feedback: proofreading the first screening data to obtain historical information, and providing feedback to the subsequent target content information collection stage; S4, difference feature extraction and labeling preparation: perform overflow screening on the historical information to obtain difference information, and label the difference information as data to be labeled; S5, automatic labeling and error rate verification and update: segment the data to be labeled with the data to be labeled to obtain labeled data, verify the labeled data according to the error rate of historical information and update the error rate; S6. Deviation calculation and monitoring interval generation: perform deviation calculation on the updated error rate, generate monitoring evaluation interval based on the deviation, update the error rate at the next data annotation, and classify the output results of the automatic annotation model into frozen data and unfrozen data based on the instantaneous value of the updated error rate and the monitoring evaluation interval.

2. The method for automatic data annotation based on model training according to claim 1, characterized in that: The S1 specifically includes: A collection website list is constructed, and a collection instruction is issued to each collection website in the collection website list; an interface of the collection website is called, each collection website is accessed, and output information of the collection website is collected to obtain target content information.

3. A method for automatic data annotation based on model training according to claim 1, characterized in that: The S2 specifically includes: First, the timeliness of each target content information is evaluated and processed according to the information timeliness in the first evaluation condition, and data that does not meet the condition is eliminated; the information timeliness is the timeliness interval of the target information set by the user; The information relevance of the remaining data is then calculated, and the content with information relevance not less than a threshold is screened out as the first screening data; the information relevance is the semantic cosine value between the target information and the annotation label input by the user.

4. A method for automatic data annotation based on model training according to claim 1, characterized in that: The S3 includes: Construct a list of information to be matched, and add each of the first screening data to the list; use the matching model to match the information one by one; if the match is successful, perform pre-marking synchronously, if it fails, continue to proofread, and optimize and update historical information according to the preset strategy.

5. The method for automatic data annotation based on model training according to claim 4, characterized in that: The historical information includes: The first historical information refers to the data in the first screening data that is still consistent with the first screening data after verification; The second historical information indicates data in the first screening data that is found inconsistent with the first screening data after proofreading; The third historical information represents newly added data that does not match the historical annotation tags.

6. A method for automatic data annotation based on model training according to claim 1, characterized in that: The error rate updating process in S5 includes: Construct a list of tag information to be verified, extract the update time and model input of each tag information; determine unlabeled or incorrect tags based on the matching degree with historical labeled tags, set the total number of labeling errors, calculate the deviation rate, and dynamically update the error rate according to the unlabeling frequency, and incorporate the updated results back into the historical information.

7. A method for automatic data annotation based on model training according to claim 1, characterized in that: The deviation measurement in S6 includes: Divide the updated error rate into multiple segment values in the time dimension, perform discretization preprocessing and sorting, input them into the deviation measurement model in combination with the set smoothing coefficient, and output the deviation result.

8. A method for automatic data annotation based on model training according to claim 1, characterized in that: The method for generating the monitoring and evaluation interval includes: Input the error rate into the verification model, obtain the verification information and perform wavelet transform and feature sorting, then perform first-order difference and inverse normalization processing, calculate the risk value based on the updated deviation degree and the set risk threshold, and generate the high interval value and the low interval value accordingly, and construct the final monitoring and evaluation interval.

9. A system for automatic data annotation based on model training, which is used to implement the method described in any one of claims 1 to 8, characterized in that, including: Information collection module: used to construct a list of collection websites, issue collection instructions, and call interfaces to implement the capture of text, image, and video target content information; Information screening module: used to evaluate and process the target content information according to the timeliness and relevance of the information, remove the data that does not meet the conditions, and output the first screened data; Data verification module: used to call the matching model to compare the first screened data with the historical information, generate verified and unverified information, and feedback the comparison result to the collection model; Data overflow screening module: used to extract differential information from historical information, identify the change trend, and label it as data to be labeled for use by the automatic labeling model; Automatic labeling module: used to input the data to be labeled into the model in segments for labeling, and verify and correct the labeling results in combination with the historical error rate, and dynamically update the error rate parameters; Deviation analysis module: used to perform deviation measurement on the updated error rate, generate a monitoring and evaluation interval, and classify the status of the model output results according to the degree of deviation to realize the dynamic control of the labeling quality.

10. An electronic device, characterized in that: It includes at least one processor and a memory communicatively connected to the processor. The memory stores a computer program executed by the processor, and the program is used to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Artificial intelligence data aggregation method based on big data

    CN119690974A

  • Non-transitory computer-readable recording medium storing machine learning program, machine learning method, and information processing apparatus

    US20240220776A1

Cited By

  • Image annotation method

    CN121121750A