A pre-marking system and automatic marking method thereof

By introducing data automatic labeling module, interface platform module, display module and workflow engine in the pre-labeling system, a closed-loop automated workflow is formed, which solves the problem of manual dependence and low degree of automation of the existing pre-labeling models, and achieves efficient data labeling and model generalization improvement.

CN118941238BActive Publication Date: 2025-05-09YITONG TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411003996.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-05-09
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

The data processing, training, release and deployment of existing pre-labeled models mainly rely on R&D personnel, resulting in high labor costs, low labeling efficiency, poor automation and generalization of the model, and poor model version management, which is prone to problems such as version confusion and performance degradation.

Method used

It provides a pre-labeling system, including a data automatic labeling module, an interface platform module, a display module and a workflow engine. It forms a closed-loop automated workflow by automatically identifying data types, calling corresponding pre-labeling models, automatically collecting and pre-processing labeling data, conducting model training and evaluation, and automatically publishing and deploying model versions.

Benefits of technology

It realizes automated annotation of the to-process data, improves the labeling efficiency and generalization of the model, reduces the workload of R&D personnel, and ensures the management and performance of model versions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118941238B_ABST
    Figure CN118941238B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and in particular to a pre-labeling system and an automated labeling method thereof, wherein the pre-labeling system includes a data automatic labeling module for identifying the data type of the data to be labeled; an interface platform module including a query interface and multiple reasoning interfaces, and different reasoning interfaces correspond to pre-labeling models of different data types; the query interface is connected to the data automatic labeling module for determining the reasoning interface corresponding to the data type of the data to be processed, so that the data automatic labeling module calls the pre-labeling model corresponding to the reasoning interface through the reasoning interface to automatically label the data to be processed and obtain the labeled data; a display module for displaying the labeled data; and a workflow engine including a data processing module, a model training and evaluation module, and a model publishing and deployment module. The present application realizes the automated labeling of the data to be processed, improves the labeling efficiency of the data to be processed, and at the same time optimizes the performance of the pre-labeling model and improves the generalization of the labeling model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a pre-labeling system and an automated labeling method thereof. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of pre-labeling models in data labeling is becoming more and more extensive. In order to realize the automatic labeling of target data by pre-labeling models, R&D personnel first need to develop a pre-labeling model, and manually label a batch of data, and then manually review the accuracy of the labeled data one by one, and make the approved labeled data into a data set to train the model. After the training is completed, the model reasoning service is manually deployed, and the deployed model reasoning service is used to realize automatic pre-labeling of data. However, in this solution, the data processing, training, publishing, and deployment of the pre-labeling model are mainly driven by R&D personnel, which increases labor costs, low data labeling processing efficiency, and low degree of automation of data labeling of the model; there is a lack of effective supervision mechanism for model publishing and reasoning, which is prone to problems such as model version confusion and performance degradation. In addition, since the data processing, training, publishing, and deployment of the pre-labeling model are mainly driven by R&D personnel, the capabilities of the pre-labeling model are fixed once it is deployed, which makes the flexibility, scalability, and generalization of the pre-labeling model poor; when the data tag type needs to be changed, it takes a lot of time and labor costs, and the operation and maintenance cost of the model is high. Summary of the invention

[0003] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application embodiment provides a pre-labeling system and an automated labeling method thereof, which realizes automated labeling of data to be processed, improves the labeling efficiency of data to be processed, and optimizes the performance of the pre-labeling model and improves the generalization of the labeling model.

[0004] In a first aspect, an embodiment of the present application provides a pre-marking system, including:

[0005] The data automatic labeling module is used to identify the data type of the data to be labeled;

[0006] The interface platform module includes a query interface and a plurality of reasoning interfaces, wherein different reasoning interfaces correspond to pre-annotated models of different data types; the query interface is connected to the data automatic annotation module, and is used to determine the reasoning interface corresponding to the data type of the data to be processed, so that the data automatic annotation module automatically annotates the data to be processed by calling the pre-annotated model corresponding to the reasoning interface through the reasoning interface to obtain annotated data;

[0007] A display module, used to display the marked data so that the reviewer can review and confirm the marked data;

[0008] Workflow engine, including data processing module, model training and evaluation module, and model publishing and deployment module.

[0009] The data processing module is used to collect the verified annotated data sent by the data push interface of the interface platform module, and pre-process the verified annotated data to obtain a training data set;

[0010] The model training and evaluation module is used to obtain a full training data set or an incremental training data set to iteratively train the pre-labeled model, and is also used to split the training data set into an evaluation data set, and evaluate the pre-labeled model according to the evaluation data set to determine an evaluation index, and perform an update or terminate an iteration operation according to the evaluation index;

[0011] The model publishing and deployment module is used to determine the publishing model version of the pre-labeled model according to the evaluation index, and is also used to update the query interface and the reasoning interface of the interface platform module according to the publishing model version.

[0012] According to some embodiments of the present application, the pre-labeling system also includes a timer, which is used to send a trigger signal at a preset time interval so that the data processing module receives the reviewed labeling data through the data push interface of the interface platform module after receiving the trigger signal.

[0013] In a second aspect, an embodiment of the present application provides an automatic labeling method of a pre-labeling system, which is applied to the pre-labeling system described in the technical solution of the first aspect, including:

[0014] Identify the data type of the data to be annotated;

[0015] Determine an inference interface corresponding to the data type of the data to be processed, so that the data automatic annotation module calls the pre-annotation model corresponding to the inference interface through the inference interface to automatically annotate the data to be processed, and obtain annotated data;

[0016] Displaying the marked data for reviewers to review and confirm the marked data;

[0017] Collecting the verified annotated data sent by the data push interface of the interface platform module, and preprocessing the verified annotated data to obtain a training data set;

[0018] Obtaining a full training data set or an incremental training data set to iteratively train the pre-labeled model;

[0019] Splitting the training data set into an evaluation data set, and evaluating the pre-labeled model according to the evaluation data set to determine an evaluation index, and performing an update or terminating an iteration operation according to the evaluation index;

[0020] The published model version of the pre-labeled model is determined according to the evaluation index, and the query interface and the reasoning interface of the interface platform module are updated according to the published model version.

[0021] According to some embodiments of the present application, the pre-marking system includes a timer, and the automatic marking method further includes:

[0022] A trigger signal is sent at a preset time interval, so that the data processing module receives the audited annotation data through the data push interface of the interface platform module after receiving the trigger signal.

[0023] According to some embodiments of the present application, after identifying the data type of the data to be annotated, the automatic annotation method further includes:

[0024] Displaying the marking method of the data to be processed on the display module, the marking method includes an automatic marking method and a manual marking method;

[0025] In response to the trigger information of the manual marking method, the data to be processed is displayed on the display module, so that the reviewer can mark and review the data to be processed, obtain the marked data, and store the marked data;

[0026] Determine a reasoning interface corresponding to the annotated data, and send the annotated data to a data processing module of the workflow engine through the reasoning interface;

[0027] In response to the trigger information of the automatic annotation mode, the data to be processed is automatically annotated through the reasoning interface to obtain the annotated data.

[0028] According to some embodiments of the present application, the query interface and the reasoning interface are encapsulated;

[0029] Registering the encapsulated query interface and the reasoning interface to an interface platform module, so that the interface platform module manages and calls the query interface and the reasoning interface;

[0030] When receiving the data to be processed, confirming the query interface that matches the data type of the data to be processed;

[0031] Calling a query interface that matches the data type of the data to be processed, wherein the query interface supports protocol conversion, and switches the interface type of the query interface to an http interface or a webservice interface according to the data type of the data to be processed;

[0032] The data in the query interface is parsed and configured to obtain a data parsing result, and the data parsing result in the query interface is saved to a database or a file system.

[0033] According to some embodiments of the present application, after obtaining the data analysis result, the automatic labeling method further includes:

[0034] Obtain data meta information according to the schema definition, and select annotated data that meets the quality and freshness standards according to the data meta information;

[0035] Loading data according to the path of the labeled data that meets the quality and freshness standards;

[0036] The method of loading data is determined according to the storage speed of the annotated data that meets the quality and freshness standards. When the storage speed of the annotated data that meets the quality and freshness standards is fast, the method of loading data is determined to be directly reading the data storage path or address through the pre-annotation model. When the storage speed of the annotated data that meets the quality and freshness standards is slow, the method of loading data is determined to be extracting and copying the annotated data to the environment where the pre-annotation model is trained by creating a local copy. When the annotated data that meets the quality and freshness standards has a customized storage method, the method of loading data is determined to be specifying the data storage address or path.

[0037] According to some embodiments of the present application, the loaded annotated data that meets the quality standard and the freshness standard is split into a training data set and an evaluation data set, and the pre-annotated model is evaluated in combination with the benchmark training set and the benchmark evaluation set to determine the evaluation index;

[0038] When the evaluation index of the pre-labeled model is less than a preset threshold, obtaining a full training data set to perform full training, wherein the full training data set includes the training data set and the benchmark training set;

[0039] When the evaluation index of the pre-labeled model is greater than a preset threshold, an incremental training data set is obtained to perform incremental training, where the incremental training data set includes the training data set loaded in the current round.

[0040] According to some embodiments of the present application, performing an update or terminating an iteration operation according to the evaluation indicator includes:

[0041] Recording the evaluation index of each iterative training of the pre-labeled model and storing it in an evaluation index database;

[0042] Compare the current evaluation index with any evaluation index in the evaluation index database, if the current evaluation index is the maximum evaluation index, take the current pre-labeled model version as the optimal model version, and perform an update deployment operation;

[0043] The current evaluation index is compared with any evaluation index in the evaluation index database, and if the current evaluation index is the minimum evaluation index, a termination iteration operation is performed.

[0044] In a third aspect, an embodiment of the present application provides a controller, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor runs the computer program, the automatic annotation method of the pre-annotation system as described in the technical solution of the second aspect above is executed.

[0045] The technical solution of the present application has at least one of the following advantages or beneficial effects: the data type of the data to be annotated is determined by the data automatic annotation module, the reasoning interface corresponding to the data type of the data to be processed is determined by the query interface of the interface platform module, different reasoning interfaces correspond to pre-annotation models of different data types, the data automatic annotation module automatically annotates the data to be processed by checking the determined reasoning interface corresponding to the data type to be annotated, and calls the pre-annotation model corresponding to the reasoning interface to obtain the annotated data, thereby realizing the automatic annotation of the data to be processed. By displaying the annotated data in the display module for the auditor to review and confirm, the accuracy of the annotation results of the data to be annotated is improved. The data processing module of the workflow engine realizes the automatic collection of annotated data and the automatic preprocessing of annotated data, the model training and evaluation module of the workflow engine realizes the iterative training of the pre-annotated model and the execution of update or termination of iteration operations, realizes the automatic training of the pre-annotated model, determines the release model version of the pre-annotated model by the model release and deployment module, and updates the query interface and reasoning interface of the interface platform module, realizes the automatic release and deployment of the model version of the pre-annotated model, the automatic update of the platform interface module, etc., greatly reduces the workload of the computing R&D personnel, and improves the marking efficiency of the data to be processed. The pre-labeling system of the embodiment of the present application forms a closed loop by connecting the automatic labeling module and the interface platform module, and connecting the interface platform module and the workflow engine, so as to realize the automatic labeling of the data to be processed, the automatic collection of the labeled data, the automatic training of the pre-labeling model, the automatic publishing of the model version of the pre-labeling model, the automatic updating of the query interface of the interface platform module, and the automatic updating of the reasoning interface, thereby realizing the automatic labeling of the data to be processed, improving the marking efficiency of the data to be processed, optimizing the performance of the pre-labeling model, and improving the generalization of the labeling model.

[0046] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a structural diagram of a pre-marking system provided in an embodiment of the present application;

[0048] Figure 2 is a flow chart of an automatic labeling method of a pre-labeling system provided in an embodiment of the present application;

[0049] Figure 3 is a flow chart of a method for obtaining labeled data provided in an embodiment of the present application;

[0050] Figure 4 is a flow chart of another automatic labeling method of a pre-labeling system provided in an embodiment of the present application;

[0051] Figure 5 is a flow chart of a method for determining loading data provided by an embodiment of the present application;

[0052] Figure 6 It is a flowchart of a method for determining a training data set of a pre-labeled model according to an evaluation index provided in an embodiment of the present application;

[0053] Figure 7 It is a flowchart of a method for performing an update or terminating an iterative operation according to an evaluation indicator provided in an embodiment of the present application;

[0054] Figure 8 It is a structural schematic diagram of a controller provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. In addition, the characteristics, operations or features described in the specification can be combined in any appropriate manner to form various implementation methods. At the same time, the steps or actions in the method description can also be replaced or adjusted in order in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the accompanying drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a necessary sequence, unless otherwise specified that a certain sequence must be followed.

[0056] In the description of this application, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0057] The serial numbers of the components in this document, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning. The "connection" and "coupling" mentioned in this application, unless otherwise specified, include direct and indirect connections (couplings).

[0058] With the rapid development of artificial intelligence technology, the application of pre-labeling models in data labeling is becoming more and more extensive. In order to realize the automatic labeling of target data by pre-labeling models, R&D personnel first need to develop a pre-labeling model, and manually label a batch of data, and then manually review the accuracy of the labeled data one by one, and make the approved labeled data into a data set to train the model. After the training is completed, the model reasoning service is manually deployed, and the deployed model reasoning service is used to realize automatic pre-labeling of data. However, in this solution, the data processing, training, publishing, and deployment of the pre-labeling model are mainly driven by R&D personnel, which increases labor costs, low data labeling processing efficiency, and low degree of automation of data labeling of the model; there is a lack of effective supervision mechanism for model publishing and reasoning, which is prone to problems such as model version confusion and performance degradation. In addition, since the data processing, training, publishing, and deployment of the pre-labeling model are mainly driven by R&D personnel, the capabilities of the pre-labeling model are fixed once it is deployed, which makes the flexibility, scalability, and generalization of the pre-labeling model poor; when the data tag type needs to be changed, it takes a lot of time and labor costs, and the operation and maintenance cost of the model is high.

[0059] Based on this, the embodiment of the present application provides a pre-labeling system and an automated labeling method thereof, which realizes automated labeling of data to be processed, improves the labeling efficiency of data to be processed, and at the same time optimizes the performance of the pre-labeling model and improves the generalization of the labeling model.

[0060] The embodiments of the present application are further described below with reference to the accompanying drawings.

[0061] Reference Figure 1 As shown, Figure 1 : is a schematic diagram of the structure of a pre-marking system provided in an embodiment of the present application, the pre-marking system includes:

[0062] An automatic data labeling module, an interface platform module, a display module and a workflow engine; the automatic data labeling module is used to identify the data type of the data to be labeled; the interface platform module includes a query interface and multiple reasoning interfaces, and different reasoning interfaces correspond to pre-labeling models of different data types; the query interface is connected to the automatic data labeling module, and the query interface is used to determine the reasoning interface corresponding to the data type of the data to be processed, so that the automatic data labeling module calls the pre-labeling model corresponding to the reasoning interface through the reasoning interface to automatically label the data to be processed and obtain the labeled data; the display module is used to display the labeled data for the auditor to review and confirm the labeled data. The workflow engine includes a data processing module, a model training and evaluation module, and a model publishing and deployment module, wherein the data processing module is used to collect the audited labeled data sent by the data push interface of the interface platform module, and pre-process the audited labeled data to obtain a training data set; the model training and evaluation module is used to obtain a full training data set or an incremental training data set to iteratively train the pre-labeled model, and is also used to split the training data set into an evaluation data set, and evaluate the pre-labeled model according to the evaluation data set to determine the evaluation index, and perform an update or terminate the iteration operation according to the evaluation index; the model publishing and deployment module is used to determine the published model version of the pre-labeled model according to the evaluation index, and is also used to update the query interface and reasoning interface of the interface platform module according to the published model version.

[0063] In the embodiment of the present application, the data type of the data to be annotated is determined by the data automatic annotation module, the reasoning interface corresponding to the data type of the data to be processed is determined by the query interface of the interface platform module, different reasoning interfaces correspond to pre-annotation models of different data types, and the data automatic annotation module automatically annotates the data to be processed by checking the determined reasoning interface corresponding to the data type to be annotated, and calling the pre-annotation model corresponding to the reasoning interface to obtain the annotated data, thereby realizing the automatic annotation of the data to be processed. By displaying the annotated data in the display module for the auditor to review and confirm, the accuracy of the annotation results of the data to be annotated is improved. The data processing module of the workflow engine realizes the automatic collection of annotated data and the automatic preprocessing of annotated data, the model training and evaluation module of the workflow engine realizes the iterative training of the pre-annotated model and the execution of the update or termination of the iteration operation, realizes the automatic training of the pre-annotated model, determines the release model version of the pre-annotated model by the model release and deployment module, and updates the query interface and reasoning interface of the interface platform module, realizes the automatic release and deployment of the model version of the pre-annotated model, the automatic update of the platform interface module, etc., greatly reduces the workload of the computing R&D personnel, and improves the marking efficiency of the data to be processed.

[0064] The pre-labeling system of the embodiment of the present application forms a closed loop by connecting the automatic labeling module and the interface platform module, and connecting the interface platform module and the workflow engine, so as to realize the automatic labeling of the data to be processed, the automatic collection of the labeled data, the automatic training of the pre-labeling model, the automatic publishing of the model version of the pre-labeling model, the automatic updating of the query interface of the interface platform module, and the automatic updating of the reasoning interface, thereby realizing the automatic labeling of the data to be processed, improving the marking efficiency of the data to be processed, optimizing the performance of the pre-labeling model, and improving the generalization of the labeling model.

[0065] In some embodiments of the present application, the data automatic labeling module includes multiple labeling tools, and different labeling tools are used to identify different data types, so that the data automatic labeling module can classify the data to be labeled of various different data types, thereby improving the data module and sending the classified data to be labeled to the interface platform module, so that the reasoning interface of the interface platform module calls the pre-labeling model corresponding to the reasoning interface to automatically label the data to be processed and obtain labeled data.

[0066] The data type of the data to be labeled is determined by the data automatic labeling module to confirm the pre-labeling model corresponding to the data type of the data to be labeled, thereby improving the accuracy of the labeling results of the data to be labeled.

[0067] In some embodiments of the present application, the pre-labeling system also includes a timer, which is used to send a trigger signal at a preset time interval so that the data processing module receives the reviewed labeling data through the data push interface of the interface platform module after receiving the trigger signal.

[0068] In this application, data collection is performed and processed by timer or interface triggering. Each time the annotated data is obtained, it can be screened according to metadata information and configuration information such as data time, data type, and data quality. It can be understood that data collection is performed by timer in the form of sending a trigger signal at a preset time interval. After the data processing module receives the trigger signal, the data processing module receives the audited annotated data through the data push interface of the interface platform module. In this application, the trigger signal is sent by the timer at a preset time interval to realize automatic data collection, thereby improving the automation of the pre-annotation system.

[0069] Reference Figure 2 As shown, Figure 2 is a flowchart of an automatic labeling method of a pre-labeling system provided in an embodiment of the present application, including but not limited to steps S100 to S160. Specifically,

[0070] Step S100: Identify the data type of the data to be labeled;

[0071] Step S110: determining an inference interface corresponding to the data type of the data to be processed, so that the data automatic annotation module calls the pre-annotated model corresponding to the inference interface through the inference interface to automatically annotate the data to be processed, and obtains annotated data;

[0072] Step S120: displaying the marked data for the auditor to review and confirm the marked data;

[0073] Step S130: collecting the verified annotated data sent by the data push interface of the interface platform module, and preprocessing the verified annotated data to obtain a training data set;

[0074] Step S140: Obtain a full training data set or an incremental training data set to perform iterative training on the pre-labeled model;

[0075] Step S150: split the training data set into evaluation data sets, and evaluate the pre-labeled model according to the evaluation data sets to determine the evaluation index, and perform an update or terminate the iteration operation according to the evaluation index;

[0076] Step S160: Determine the published model version of the pre-labeled model according to the evaluation index, and update the query interface and reasoning interface of the interface platform module according to the published model version.

[0077] In some embodiments of the present application, the automatic annotation method of the pre-annotation system includes: after receiving the data to be processed, identifying the data type of the data to be annotated, and determining the reasoning interface corresponding to the data type of the data to be processed, the data automatic annotation module calls the pre-annotation model corresponding to the reasoning interface through the reasoning interface corresponding to the data type of the data to be processed, and realizes automatic annotation of the data to be processed through the pre-annotation model to obtain the annotated data. After the annotated data to be processed is completed, the annotated data is displayed for the auditor to review and confirm the annotated data, thereby improving the accuracy of the annotation results of the data to be annotated. The audited annotated data sent by the data push interface is collected, and the audited annotated data is pre-processed to obtain a training data set, thereby realizing the automatic preprocessing of the annotated data. The full training data set or the incremental training data set is obtained to iteratively train the pre-annotated model and perform update or termination of the iteration operation, thereby realizing the automatic training of the pre-annotated model, determining the release model version of the pre-annotated model according to the evaluation index, and updating the query interface and reasoning interface of the interface platform module, thereby realizing the automatic release and deployment of the model version of the pre-annotated model, the automatic update of the platform interface module, etc., thereby greatly reducing the workload of the computing R&D personnel and improving the marking efficiency of the data to be processed.

[0078] The automatic labeling method of the pre-labeling system of the embodiment of the present application forms a closed loop with a series of workflows such as automatic labeling of data to be processed, automatic collection of labeled data, automatic training of the pre-labeling model, automatic publishing of the model version of the pre-labeling model, automatic updating of the query interface of the interface platform module, and automatic updating of the reasoning interface, thereby realizing automatic labeling of the data to be processed, improving the marking efficiency of the data to be processed, and optimizing the performance of the pre-labeling model and improving the generalization of the labeling model.

[0079] In some embodiments of the present application, the automatic labeling method of the pre-labeling system further includes step S200, specifically,

[0080] Step S200: sending a trigger signal at a preset time interval, so that the data processing module receives the verified annotation data through the data push interface of the interface platform module after receiving the trigger signal.

[0081] Data collection is collected and processed through timer or interface triggering. Each time the annotated data is obtained, it can be screened according to metadata information and configuration information such as data time, data type, and data quality. It can be understood that data collection is collected through a timer in a way that the timer sends a trigger signal at a preset time interval. After the data processing module receives the trigger signal, the data processing module receives the audited annotated data through the data push interface of the interface platform module. In this application, the trigger signal is sent by the timer at a preset time interval to realize automatic data collection, thereby improving the automation level of the pre-annotation system.

[0082] Reference Figure 3 As shown, Figure 3 is a flowchart of a method for obtaining annotated data provided in an embodiment of the present application, including but not limited to steps S101 to S104. Specifically,

[0083] Step S101: displaying a marking method of the data to be processed on a display module, wherein the marking method includes an automatic marking method and a manual marking method;

[0084] Step S102: in response to the trigger information of the manual marking method, displaying the data to be processed on the display module, so that the reviewer can mark and review the data to be processed, obtain the marked data, and store the marked data;

[0085] Step S103: determining the reasoning interface corresponding to the annotated data, and sending the annotated data to the data processing module of the workflow engine through the reasoning interface;

[0086] Step S104: In response to the trigger information of the automatic annotation mode, the data to be processed is automatically annotated through the reasoning interface to obtain annotated data.

[0087] The labeled data is obtained by labeling the data to be processed. The labeling method of the data to be processed includes an automatic labeling method and a manual labeling method. In some embodiments of the present application, the automatic labeling method of the pre-labeling system includes, after the data type of the labeled data is determined, confirming the labeling method of the data to be processed, and displaying the labeling method of the data to be processed on the display module, wherein the labeling method includes an automatic labeling method and a manual labeling method. The research and development personnel can select the required data labeling mode according to the actual situation to label the data to be processed. In response to the trigger information of the manual labeling method, the data to be processed is displayed on the display module for the auditor to label and review the data to be processed, and the labeled data is obtained. The auditor reviews the labeled data, which improves the accuracy of the labeling results of the labeled data, and stores the labeled data in the display module. Then, the reasoning interface corresponding to the labeled data is determined, and the labeled data is sent to the data processing module of the workflow engine through the reasoning interface. The data processing module realizes the automatic collection of the labeled data and the automatic preprocessing of the labeled data.

[0088] In response to the trigger information of the automatic labeling method, the reasoning interface corresponding to the data type of the data to be processed is determined, and the pre-labeling model corresponding to the reasoning interface is called through the reasoning interface to automatically label the data to be processed, so as to obtain the labeled data, thereby improving the labeling efficiency of the data to be processed.

[0089] The embodiments of the present application provide two methods for labeling the data to be processed to obtain labeled data. Research and development personnel can select the required labeling method according to actual conditions, thereby increasing the diversity of methods for labeling the data to be processed to obtain labeled data.

[0090] In one embodiment of the present application, the data to be processed is first uploaded and saved, and then the data to be processed is classified, because different types of data have different labeling requirements, such as target detection and image segmentation. If automatic labeling is required, the reasoning interface is obtained through the query interface of the interface platform module, and the matching reasoning interface can be queried through metadata such as data type, and then the pre-labeled model corresponding to the reasoning interface is called to automatically label the data to be processed to obtain the labeled data; and the labeled data returned by the pre-labeled model is received, and the labeled data is visualized in the display module, and the annotation of the labeled data is manually reviewed, calibrated or completed. Finally, after confirmation, the labeled data is submitted and the reasoning interface exposed by the interface platform module is called to send the metadata of the labeled data to the workflow engine. The workflow engine realizes automatic labeling of the data to be processed, automatic collection of labeled data, automatic training of the pre-labeled model, automatic release of the model version of the pre-labeled model, and automatic update of the query interface of the interface platform module, automatic update of the reasoning interface, and a series of workflows form a closed loop, thereby realizing automatic labeling of the data to be processed, improving the marking efficiency of the data to be processed, and optimizing the performance of the pre-labeled model and improving the generalization of the labeled model.

[0091] Reference Figure 4 As shown, Figure 4 is a flowchart of another automatic labeling method of a pre-labeling system provided in an embodiment of the present application, including but not limited to steps S300 to S340. Specifically,

[0092] Step S300: encapsulating the query interface and the reasoning interface;

[0093] Step S310: registering the encapsulated query interface and reasoning interface to the interface platform module, so that the interface platform module manages and calls the query interface and reasoning interface;

[0094] Step S320: when receiving the data to be processed, confirming the query interface that matches the data type of the data to be processed;

[0095] Step S330: calling a query interface that matches the data type of the data to be processed, the query interface supports protocol conversion, and the interface type of the query interface is switched to an http interface or a webservice interface according to the data type of the data to be processed;

[0096] Step S340: parse and configure the data in the query interface to obtain data parsing results, and save the data parsing results in the query interface to a database or a file system.

[0097] In some embodiments of the present application, the interface platform module is an optional module. The interface platform module is set by the pre-marking system to make the interface use more flexible and more standardized. Whether it is a query interface or a reasoning interface, it is uniformly accessed and managed by the interface platform module. The automatic marking method of the pre-marking system also includes encapsulating the query interface and the reasoning interface, and then registering the encapsulated query interface and reasoning interface to the interface platform module so that the interface platform module manages and calls the query interface and the reasoning interface. In this way, the interface platform module can list all registered query interfaces and reasoning interfaces, and can query the interface to be used and call it. When receiving the data to be processed, confirm the query interface that matches the data type of the data to be processed, and call the query interface that matches the data type of the data to be processed. If the data type of the data to be processed and the interface protocol of the query interface are inconsistent, the query interface supports protocol conversion, and the data in the query interface is parsed and configured. For example, the interface type of the query interface is switched to an http interface or a webservice interface. Then, the data in the query interface is parsed and configured to obtain the data parsing results, and the data parsing results in the query interface are saved to the database or file system to facilitate the automatic data collection and processing by the data processing module of the workflow engine; in addition, if the collection and processing of the data to be processed does not pass through the interface platform module, the data automatic annotation module and the workflow engine are required to manually align the data to be processed and the query interface, and agree on the interface protocol, interface parameters and access address.

[0098] Reference Figure 5 As shown, Figure 5 is a flowchart of a method for determining loading data provided by an embodiment of the present application, including steps S350 to S370. Specifically,

[0099] Step S350: Obtain data meta information according to the schema definition, and filter the labeled data that meets the quality and freshness standards according to the data meta information;

[0100] Step S360: Loading data according to the path of the annotated data that meets the quality and freshness standards;

[0101] Step S370: Determine a method for loading data based on the storage speed of the annotated data that meets the quality and freshness standards. When the storage speed of the annotated data that meets the quality and freshness standards is fast, the method for loading data is determined to be directly reading the data storage path or address through the pre-annotation model. When the storage speed of the annotated data that meets the quality and freshness standards is slow, the method for loading data is determined to be extracting and copying the annotated data to the pre-annotation model training environment by creating a local copy. When the annotated data that meets the quality and freshness standards has a customized storage method, the method for loading data is determined to be specifying the data storage address or path.

[0102] In some embodiments of the present application, first, a data push interface is provided and registered with an interface platform module. The data push interface can also be sent directly to a data automatic annotation module without going through the interface platform module. After the data automatic annotation module calls the data push interface to push the data, it obtains data metadata according to the schema definition, and then screens the annotated data that meets the quality and freshness standards according to the data metadata, and loads the data according to the path of the annotated data that meets the quality and freshness standards. The loaded data can be selected according to the storage speed of the annotated data that meets the quality and freshness standards. If the storage speed is fast, choose to let the pre-annotation model read the data address directly. If the storage speed of the annotated data that meets the quality and freshness standards is slow, choose to create a local copy to extract and copy the annotated data that meets the quality and freshness standards to the pre-annotation model training environment. If there is a custom storage method, directly specify the data storage address or path manually.

[0103] By loading data according to the path of the annotated data that meets the quality and freshness standards, and determining the method of loading data according to the storage speed of the annotated data that meets the quality and freshness standards, a storage method paired with the annotated data is provided, and the efficiency of loading data is improved.

[0104] It should be noted that in some embodiments of the present application, the workflow engine needs to pre-configure the storage access account and permissions so that when the storage speed of the annotated data that meets the quality and freshness standards is fast, the pre-annotated model can be directly selected to read the data address.

[0105] Reference Figure 6 As shown, Figure 6 4 is a flowchart of a method for determining a training data set of a pre-labeled model according to an evaluation index provided in an embodiment of the present application, including steps S400 to S420. Specifically,

[0106] Step S400: split the loaded annotated data that meets the quality and freshness standards into a training data set and an evaluation data set, and evaluate the pre-annotated model in combination with the benchmark training set and the benchmark evaluation set to determine the evaluation index;

[0107] Step S410: when the evaluation index of the pre-labeled model is less than a preset threshold, a full training data set is obtained to perform full training, where the full training data set includes a training data set and a benchmark training set;

[0108] Step S420: When the evaluation index of the pre-labeled model is greater than a preset threshold, an incremental training data set is obtained to perform incremental training, where the incremental training data set includes the training data set loaded in the current round.

[0109] In some embodiments of the present application, after loading the labeled data that meets the quality and freshness standards, the loaded labeled data that meets the quality and freshness standards are split into a training data set and an evaluation data set, and then the pre-labeled model is evaluated in combination with the benchmark training set and the benchmark evaluation set to determine the evaluation index, and the corresponding pre-labeled model training strategy is executed according to the evaluation index. It can be understood that the pre-labeled model training strategy can be configured as follows: when the evaluation index of the pre-labeled model drops below a preset threshold, the full training data set is obtained to perform full training, and the full training data set includes the training data set and the benchmark training set; when the evaluation index of the pre-labeled model is greater than the preset threshold, the incremental training data set is obtained to perform incremental training, and the incremental training data set includes the training data set loaded in the current round. Incremental training will only load the training data set produced in this round, so that the training time is short and the pre-labeled model capabilities can be quickly expanded, while the full training will train all rounds of training data sets and benchmark data sets together, so that the training generalization of the pre-labeled model can be guaranteed to be the best.

[0110] Reference Figure 7 As shown, Figure 7 is a flowchart of a method for performing an update or terminating an iteration operation according to an evaluation index provided by an embodiment of the present application, including steps S151 to S153. Specifically,

[0111] Step S151: Record the evaluation index of each iterative training of the pre-labeled model and store it in the evaluation index database;

[0112] Step S152: compare the current evaluation index with any evaluation index in the evaluation index database. If the current evaluation index is the maximum evaluation index, use the current pre-labeled model version as the optimal model version and perform an update deployment operation.

[0113] Step S153: compare the current evaluation index with any evaluation index in the evaluation index database. If the current evaluation index is the minimum evaluation index, terminate the iteration operation.

[0114] In some embodiments of the present application, the automatic labeling method of the pre-labeling system also includes: recording the evaluation indicators of each iterative training of the pre-labeling model and storing them in an evaluation indicator database. After the training of the pre-labeling model is completed, you can choose to evaluate according to the evaluation data set split out of this round, or you can evaluate the evaluation sets split out of all rounds and the benchmark evaluation sets together, and judge the pros and cons of the evaluation indicators, compare the current evaluation indicator with any evaluation indicator in the evaluation indicator database, if the current evaluation indicator is the maximum evaluation indicator, use the current pre-labeling model version as the optimal model version, and perform an update and deployment operation; if the current evaluation indicator is the minimum evaluation indicator, execute the termination iteration operation.

[0115] In the embodiment of the present application, if the evaluation index is better than that of the previous iteration, the model version is released and deployed, and the reasoning interface corresponding to the pre-labeled model generated by the deployment is registered or updated to the interface platform module. The reasoning interface can also be directly notified to the data automatic labeling module without going through the interface platform module. The data automatic labeling module calls the reasoning interface for automatic labeling. If the evaluation index of this round is worse than that of the previous version, the iteration is terminated, and the timer is waited again to obtain the latest data or new data is pushed through the reasoning interface to drive the pre-labeling model training; finally, the latest labeled data is automatically obtained to drive the continuous training of the pre-labeling model, and the performance of the pre-labeling model is continuously optimized to form a closed loop of continuous iteration.

[0116] Reference Figure 8 , Figure 8 It is a structural diagram of a controller 1000 provided in an embodiment of the present application, including a processor 1001, which can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the automatic labeling method of the pre-labeling system provided in the embodiment of the present application; a memory 1002, which can be implemented in the form of a read-only memory 1002 (Read Only Memory, ROM), a static storage device, a dynamic storage device or a random access memory 1002 (Random Access Memory, RAM) and the like. The memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the embodiments of this application; the input / output interface 1003 is used to implement information input and output; the communication interface 1004 is used to implement communication interaction between this device and other devices, and communication can be achieved through wired methods (such as USB, network cables, etc.) or wireless methods (such as mobile networks, WIFI, Bluetooth, etc.); a bus transmits information between various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003 and the communication interface 1004); wherein the processor 1001, the memory 1002, the input / output interface 1003 and the communication interface 1004 are connected to each other within the device through the bus.

[0117] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer-readable storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0118] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present application. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A pre-marking system, characterized in that: include: The data automatic labeling module is used to identify the data type of the data to be labeled; The interface platform module includes a query interface and a plurality of reasoning interfaces, wherein different reasoning interfaces correspond to pre-annotated models of different data types; the query interface is connected to the data automatic annotation module, and is used to determine the reasoning interface corresponding to the data type of the data to be processed, so that the data automatic annotation module automatically annotates the data to be processed by calling the pre-annotated model corresponding to the reasoning interface through the reasoning interface to obtain annotated data; A display module, used to display the marked data so that the reviewer can review and confirm the marked data; Workflow engine, including data processing module, model training and evaluation module, and model publishing and deployment module. The data processing module is used to collect the verified annotated data sent by the data push interface of the interface platform module, and pre-process the verified annotated data to obtain a training data set; The model training and evaluation module is used to obtain a full training data set or an incremental training data set to iteratively train the pre-labeled model, and is also used to split the training data set into an evaluation data set, and evaluate the pre-labeled model according to the evaluation data set to determine an evaluation index, and perform an update or terminate an iteration operation according to the evaluation index; The model publishing and deployment module is used to determine the publishing model version of the pre-labeled model according to the evaluation index, and is also used to update the query interface and the reasoning interface of the interface platform module according to the publishing model version.

2. The pre-marking system according to claim 1, characterized in that: It also includes a timer, which is used to send a trigger signal at a preset time interval, so that the data processing module receives the reviewed annotation data through the data push interface of the interface platform module after receiving the trigger signal.

3. An automatic labeling method for a pre-labeling system, applied to the pre-labeling system as claimed in any one of claims 1 to 2, characterized in that: include: Identify the data type of the data to be annotated; Determine an inference interface corresponding to the data type of the data to be processed, so that the data automatic annotation module calls the pre-annotation model corresponding to the inference interface through the inference interface to automatically annotate the data to be processed, and obtain annotated data; Displaying the marked data for reviewers to review and confirm the marked data; Collecting the verified annotated data sent by the data push interface of the interface platform module, and preprocessing the verified annotated data to obtain a training data set; Obtaining a full training data set or an incremental training data set to iteratively train the pre-labeled model; Splitting the training data set into an evaluation data set, and evaluating the pre-labeled model according to the evaluation data set to determine an evaluation index, and performing an update or terminating an iteration operation according to the evaluation index; The published model version of the pre-labeled model is determined according to the evaluation index, and the query interface and the reasoning interface of the interface platform module are updated according to the published model version.

4. The automatic labeling method of the pre-labeling system according to claim 3, characterized in that: The pre-marking system includes a timer, and the automatic marking method further includes: A trigger signal is sent at a preset time interval, so that the data processing module receives the audited annotation data through the data push interface of the interface platform module after receiving the trigger signal.

5. The automatic labeling method of the pre-labeling system according to claim 3, characterized in that: After identifying the data type of the data to be annotated, the automatic annotation method further includes: Displaying the marking method of the data to be processed on the display module, the marking method includes an automatic marking method and a manual marking method; In response to the trigger information of the manual marking method, the data to be processed is displayed on the display module, so that the reviewer can mark and review the data to be processed, obtain the marked data, and store the marked data; Determine a reasoning interface corresponding to the annotated data, and send the annotated data to a data processing module of the workflow engine through the reasoning interface; In response to the trigger information of the automatic annotation mode, the data to be processed is automatically annotated through the reasoning interface to obtain the annotated data.

6. The automatic labeling method of the pre-labeling system according to claim 3, characterized in that: Also includes: Encapsulating the query interface and the reasoning interface; Registering the encapsulated query interface and the reasoning interface to an interface platform module, so that the interface platform module manages and calls the query interface and the reasoning interface; When receiving the data to be processed, confirming the query interface that matches the data type of the data to be processed; Calling a query interface that matches the data type of the data to be processed, wherein the query interface supports protocol conversion, and switches the interface type of the query interface to an http interface or a webservice interface according to the data type of the data to be processed; The data in the query interface is parsed and configured to obtain a data parsing result, and the data parsing result in the query interface is saved to a database or a file system.

7. The automatic labeling method of the pre-labeling system according to claim 6, characterized in that: After obtaining the data analysis result, the automatic labeling method further includes: Obtain data meta information according to the schema definition, and select annotated data that meets the quality and freshness standards according to the data meta information; Loading data according to the path of the labeled data that meets the quality and freshness standards; The method of loading data is determined according to the storage speed of the annotated data that meets the quality and freshness standards. When the storage speed of the annotated data that meets the quality and freshness standards is fast, the method of loading data is determined to be directly reading the data storage path or address through the pre-annotation model. When the storage speed of the annotated data that meets the quality and freshness standards is slow, the method of loading data is determined to be extracting and copying the annotated data to the environment where the pre-annotation model is trained by creating a local copy. When the annotated data that meets the quality and freshness standards has a customized storage method, the method of loading data is determined to be specifying the data storage address or path.

8. The automatic labeling method of the pre-labeling system according to claim 7, characterized in that: Also includes: Splitting the loaded annotated data that meets the quality and freshness standards into a training data set and an evaluation data set, and evaluating the pre-annotated model in combination with a benchmark training set and a benchmark evaluation set to determine an evaluation index; When the evaluation index of the pre-labeled model is less than a preset threshold, obtaining a full training data set to perform full training, wherein the full training data set includes the training data set and the benchmark training set; When the evaluation index of the pre-labeled model is greater than a preset threshold, an incremental training data set is obtained to perform incremental training, where the incremental training data set includes the training data set loaded in the current round.

9. The automatic labeling method of the pre-labeling system according to claim 8, characterized in that: The updating or terminating the iteration operation according to the evaluation index comprises: Recording the evaluation index of each iterative training of the pre-labeled model and storing it in an evaluation index database; Compare the current evaluation index with any evaluation index in the evaluation index database, if the current evaluation index is the maximum evaluation index, take the current pre-labeled model version as the optimal model version, and perform an update deployment operation; The current evaluation index is compared with any evaluation index in the evaluation index database, and if the current evaluation index is the minimum evaluation index, a termination iteration operation is performed.

10. A controller, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the automatic annotation method of the pre-annotation system according to any one of claims 3 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Data annotation system

    CN113407980A

  • Data annotation, deep learning model training and service publishing system

    CN113706099A