Data annotation method, system and device based on multi-model cooperation and storage medium
Through the multi-model collaboration data annotation method, the problems of low efficiency and poor accuracy in the existing technology are solved, efficient and automated data annotation are realized, labor costs are reduced, and large-scale complex data processing is adapted to.
Patent Information
- Application Number
- CN202510413294.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The existing data annotation methods are inefficient, poorly accurate, and rely on manual annotation and are costly, making it difficult to meet the needs of large-scale complex data processing.
Through multi-model collaboration data annotation methods, including data cleaning and format conversion, pre-trained model screening and manual review, labeling model training, test data set screening and correction, new data set screening and low-similarity data reflow, a dynamic optimization closed loop of data is formed.
It improves the degree of automation and accuracy of data annotation, reduces the cost of manual intervention, and enhances the model's adaptability in large-scale and complex data scenarios.
Smart Images

Figure CN120337033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a data annotation method, system, device, and storage medium based on multi-model collaboration. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the collection, processing, and utilization of data have become the core of intelligent applications in all industries. The generation and wide application of large-scale data have promoted the innovation of data processing technologies. Especially in the fields of natural language processing, image recognition, speech recognition, etc., automated annotation and data cleaning have become key links to improve model performance and efficiency.
[0003] However, existing data processing methods often have problems of low efficiency and insufficient accuracy when facing large-scale complex data. Traditional data processing methods mainly rely on manual annotation and basic statistical analysis, which are not only time-consuming and laborious, but also difficult to cope with the diversity and complexity of data. Especially in data annotation technology, it still relies on manual annotation, which is inefficient and easily affected by human biases. In the scenario of large-scale data processing, the cost and time overhead of manual annotation are relatively high, making it difficult to meet the requirements of real-time and large-scale data processing. Summary of the Invention
[0004] The purpose of the present invention is to provide a data annotation method, system, device, and storage medium based on multi-model collaboration, which is used to solve the technical problems of low data annotation efficiency, poor accuracy, and high manual annotation cost in the prior art.
[0005] In a first aspect, the present invention provides a data annotation method based on multi-model collaboration, and the method includes:
[0006] Clean and convert the format of the original data to generate structured data;
[0007] Use a pre-trained model to screen the structured data, and conduct manual review on the screened target data to obtain a training data set;
[0008] Train an annotation model based on the training data set;
[0009] Run the trained annotation model on a test data set, screen and manually correct mislabeled data, and use the corrected data for re-training the model;
[0010] Conduct binary classification screening on the newly added data set, and screen out low-similarity data based on vectorization processing and similarity calculation, and return the low-similarity data to the training data set.
[0011] Optionally, the specific steps of cleaning and formatting the original data to generate structured data include:
[0012] Obtain the original data;
[0013] Perform noise removal, duplicate data removal, and missing value processing on the original data to obtain the cleaned data;
[0014] Perform format conversion on the cleaned data to generate the structured data;
[0015] Upload the structured data to the server cluster.
[0016] Optionally, the specific steps of using a pre-trained model to screen the structured data and manually review the screened target data to obtain a training dataset include:
[0017] Perform binary classification marking on the structured data through the pre-trained model to screen out the target data;
[0018] Perform secondary manual marking on the target data to obtain positive sample data and difficult negative sample data;
[0019] Combine the positive sample data and the difficult negative sample data to form the training dataset.
[0020] Optionally, the specific steps of running the trained annotation model on the test dataset, screening and manually correcting mislabeled data, and using the corrected data for retraining the model include:
[0021] Apply the trained annotation model to the test dataset, evaluate the annotation results, and identify mislabeled data;
[0022] Manually review and correct the mislabeled data;
[0023] Add the corrected data to the training dataset and retrain the annotation model based on the updated dataset.
[0024] Optionally, the specific steps of performing binary classification screening on the new dataset, screening out low-similarity data based on vectorization processing and similarity calculation, and returning the low-similarity data to the training dataset include:
[0025] Perform binary classification screening on the new dataset through the trained annotation model to obtain annotated data;
[0026] Perform vectorization processing on the annotated data and historical data through a vector model to convert the annotated data and the historical data into vector representations;
[0027] Calculate the vector similarity between the labeled data and the historical data, and filter out the low-similarity data with a similarity lower than a preset threshold;
[0028] Return the low-similarity data to the training data set.
[0029] Optionally, the specific steps of obtaining the original data include obtaining the original data from an online platform by means of API call, database query, or file download.
[0030] Optionally, the specific steps of missing value processing include identifying the missing values in the original data and processing them by means of interpolation, deletion, or filling.
[0031] In a second aspect, the present invention also provides a data annotation system adopting the above-mentioned data annotation method based on multi-model collaboration, and the system includes:
[0032] A data processing module, configured to clean and format-convert the original data to generate structured data;
[0033] A data screening module, configured to screen the structured data by using a pre-trained model and perform manual review on the screened target data to obtain a training data set;
[0034] A model training module, configured to train an annotation model based on the training data set;
[0035] A data correction module, configured to run the trained annotation model on a test data set, screen and manually correct the mislabeled data, and use the corrected data for re-training the model;
[0036] A data return module, configured to perform binary classification screening on new data, screen out low-similarity data by using vectorization processing and similarity calculation, and return the low-similarity data to the training data set.
[0037] In a third aspect, the present invention also provides a device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0038] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned data annotation method based on multi-model collaboration.
[0039] In a fourth aspect, the present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned data annotation method based on multi-model collaboration are implemented.
[0040] According to the solution of the present invention, through data cleaning and format conversion, the standardization and availability of data are improved, and a pre-trained model is combined with manual review to ensure the quality of the training data set. On this basis, the trained annotation model can effectively improve the annotation efficiency, and continuously optimize the model performance through the screening and correction mechanism of the test data set. In addition, by screening the newly added data set and returning the low-similarity data, a dynamic optimization closed loop of data is formed, thereby improving the automation degree and accuracy of data annotation, reducing the cost of manual intervention, and enhancing the adaptability of the model in large-scale and complex data scenarios.
[0041] Furthermore, the present invention improves the flexibility of data collection by supporting various data acquisition methods such as API calls, database queries, and file downloads, to adapt to the data collection requirements from different sources. For the noise, duplicate values, and missing values in the data, a systematic cleaning strategy is adopted, including noise removal, duplicate data removal, and methods for handling missing values such as interpolation, deletion, or filling, to ensure the integrity and consistency of the data. In addition, by manually reviewing the target data and optimizing the screening of positive samples and difficult negative samples, the training data set is made more accurate, and the effectiveness of model training is improved.
[0042] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following details the preferred embodiments of the present invention as follows. Brief Description of the Drawings
[0043] Figure 1 Shows a schematic flowchart of a data annotation method based on multi-model collaboration according to an embodiment of the present invention;
[0044] Figure 2 Shows Figure 1 A schematic flowchart of a method for cleaning and format-converting the original data in step S100 shown to generate structured data;
[0045] Figure 3 Shows Figure 1 A schematic flowchart of a method for screening the structured data using a pre-trained model in step S200 shown and manually reviewing the screened target data to obtain a training data set;
[0046] Figure 4 Shows Figure 1 A schematic flowchart of a method for running the trained annotation model on the test data set in step S400 shown, screening and manually correcting the mislabeled data, and using the corrected data for re-training the model;
[0047] Figure 5 Shows Figure 1Schematic flowchart of the method for performing binary classification screening on the newly added data set in step S500, screening low-similarity data based on vectorization processing and similarity calculation, and returning the low-similarity data to the training data set;
[0048] Figure 6 Shows a structural block diagram of a data annotation system based on multi-model collaboration according to an embodiment of the present invention;
[0049] Figure 7 Shows a structural block diagram of a data annotation device based on multi-model collaboration according to an embodiment of the present invention. Detailed implementation manners
[0050] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe the detailed implementation manners of the present application in conjunction with the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the present application are shown in the accompanying drawings rather than all the structures. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0051] The terms "including" and "having" and any variations thereof in the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0052] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0053] Figure 1 Shows a schematic flowchart of a data annotation method based on multi-model collaboration according to an embodiment of the present invention. As Figure 1 shown, the data annotation method based on multi-model collaboration includes:
[0054] Step S100, cleaning and format conversion of the original data to generate structured data.
[0055] In step S100, structured data is generated by cleaning and formatting the original data, making the original data meet the standardization requirements, improving the efficiency and accuracy of subsequent processing, and providing high-quality data input for data screening and model training.
[0056] Step S200: Use a pre-trained model to screen the structured data, and manually review the selected target data to obtain a training dataset.
[0057] In this step S200, a pre-trained model is used to screen the structured data, and a high-quality training dataset is constructed through manual review. By combining automatic screening and manual verification, the accuracy and reliability of data screening are improved, and the representativeness of the training dataset is ensured. This step S200 can reduce the interference of low-quality data and improve the effect of subsequent model training.
[0058] Step S300: Train the annotation model based on the training dataset.
[0059] Training the annotation model based on the training dataset enables the model to learn data features and optimize its annotation ability. Through iterative training, the generalization ability of the model is improved, enabling it to adapt to data of different categories and features. Through this step S300, the degree of automation of annotation is improved, the dependence on manual annotation is reduced, and thus the overall annotation efficiency is enhanced.
[0060] Step S400: Run the trained annotation model on the test dataset, screen and manually correct the mislabeled data, and use the corrected data for retraining the model.
[0061] Running the trained annotation model on the test dataset, screening and manually correcting the mislabeled data, and at the same time returning the corrected data to the training dataset. Through this step S400, the annotation accuracy of the model can be effectively improved, the impact of mislabeling can be reduced, and the stability and reliability of the model in practical applications can be ensured.
[0062] Step S500: Perform binary classification screening on the new dataset, and screen out low-similarity data based on vectorization processing and similarity calculation, and return the low-similarity data to the training dataset.
[0063] In this step S500, binary classification screening is performed on the new dataset, and low-similarity data is screened out through similarity calculation and returned to the training dataset. This process ensures the diversity of training data, reduces the interference of redundant data, enhances the adaptability of the model to different data distributions, and thus improves its annotation accuracy in various scenarios.
[0064] According to this embodiment, through data cleaning and format conversion, the standardization and usability of data are improved, and a pre-trained model is combined with manual review to ensure the quality of the training dataset. On this basis, the trained annotation model can effectively improve the annotation efficiency, and continuously optimize the model performance through the screening and correction mechanism of the test dataset. In addition, by screening the newly added dataset and returning low-similarity data, a dynamic optimization closed-loop of data is formed, thereby improving the automation degree and accuracy of data annotation, reducing the cost of manual intervention, and enhancing the adaptability of the model in large-scale and complex data scenarios.
[0065] Figure 2 shows Figure 1 a schematic flowchart of the method for cleaning and format-converting the original data in step S100 shown to generate structured data. As Figure 2 shown, this step S100 includes:
[0066] Step S110, obtaining the original data.
[0067] In some embodiments, the specific steps of obtaining the original data include obtaining the original data from an online platform through API calls, database queries, or file downloads. These methods can efficiently obtain data from different sources, ensure the comprehensiveness and timeliness of the data, and provide a basis for subsequent processing.
[0068] Step S120, performing noise removal, duplicate data removal, and missing value processing on the original data to obtain the cleaned data.
[0069] In this step, first, a pre-trained model is applied to perform binary classification on the data, thereby effectively identifying and removing noise information in the data, such as advertisements, spam, etc. This process ensures the quality of the dataset and prevents interference from irrelevant information in subsequent processing. Secondly, for the missing values in the data, methods such as interpolation, deletion, or filling are used to process them to ensure the integrity and consistency of the data, and to avoid affecting the analysis and model training effects due to missing data. Finally, duplicate records in the data are removed to ensure the uniqueness of each piece of data and prevent unnecessary biases in model training caused by duplicate data. Through these comprehensive data cleaning operations, irrelevant, redundant, or incomplete data can be effectively removed, ensuring the accuracy, reliability, and consistency of the data, and providing a high-quality and clear data basis for subsequent data analysis and model training.
[0070] Step S130, performing format conversion on the cleaned data to generate structured data.
[0071] In this step, the cleaned data is converted in format to generate unified and standardized structured data. For example, JSON data is converted into an XLSX file. Structured data is convenient for subsequent analysis and processing, can improve the data processing efficiency, and provide clear and standard data input for the training of the annotation model.
[0072] Step S140: Upload the structured data to the server cluster.
[0073] The structured data after format conversion is uploaded to the server cluster for centralized storage and efficient computing. Through the processing capabilities of the cloud server cluster, the security and scalability of data storage are ensured, and support is provided for subsequent data analysis and model training. In some embodiments, the data can be packaged by compressing it into a ZIP file or the like, and uploaded to the server through protocols such as FTP, SFTP, and HTTP.
[0074] In some embodiments, Figure 3 shows Figure 1 a schematic flowchart of the method for screening structured data using a pre-trained model and manually reviewing the screened target data to obtain a training dataset in step S200 shown. As Figure 3 shown, this step S200 includes:
[0075] Step S210: Perform binary classification labeling on the structured data through a pre-trained model to screen out target data.
[0076] First, use the pre-trained model to automatically process the input structured data, perform a binary classification task, and label the data as valuable data and worthless data. In this way, the efficiency of data screening can be effectively improved, and potentially useful target data can be automatically extracted from large-scale data, thus providing high-quality candidate data for subsequent manual review and further processing.
[0077] Step S220: Perform secondary manual labeling on the target data to obtain positive sample data and difficult negative sample data.
[0078] In this step S220, the manual annotator performs secondary verification and annotation on the target data according to the preliminary classification results of the pre-trained model. Specifically, the positive samples and the difficult-to-classify negative samples in the target data are respectively labeled to ensure the accuracy of the data labels. Through manual review, the details that the model may miss can be supplemented, and the accuracy and reliability of data annotation can be improved, especially in the case of complex data or fuzzy boundaries.
[0079] Step S230: Combine the positive sample data and the difficult negative sample data to form a training dataset.
[0080] In step S230, the positive sample data and hard negative sample data obtained in step S220 are effectively combined to form a final training data set. By combining these two types of data, it is ensured that the training data set contains representative and challenging data, which helps the annotation model improve its recognition performance in practical applications. Especially when dealing with complex and unclear boundaries, it can effectively improve the generalization ability and accuracy of the model.
[0081] In some embodiments, Figure 4 shows Figure 1 a schematic flowchart of the method for running the trained annotation model on the test data set in step S400, screening and manually correcting mislabeled data, and using the corrected data for retraining the model. As Figure 4 shown, this step S400 includes:
[0082] Step S410, applying the trained annotation model to the test data set, evaluating the annotation results and identifying mislabeled data.
[0083] In this step, by applying the trained annotation model to the test data set, the output of the model is comprehensively evaluated. The system automatically identifies the mislabeled data in the annotation results, including cases of incorrect annotation or inaccurate classification. In this way, potential problems in the annotation process can be quickly located, providing accurate mislabeled data for subsequent correction steps and ensuring the improvement of the model's accuracy and reliability.
[0084] Step S420, manually reviewing and correcting the mislabeled data.
[0085] The human annotator carefully reviews the mislabeled data identified in step S410 and makes corrections according to the actual situation. Manual review can ensure the high accuracy of data annotation, especially for data that is difficult for the model to correctly identify or classify. Through manual intervention, the biases and errors that may occur in the model can be eliminated, improving the data quality and providing more accurate annotated data for subsequent model training.
[0086] Step S430, adding the corrected data to the training data set and retraining the annotation model based on the updated data set.
[0087] This step S430 re-integrates the manually corrected annotated data into the training data set and retrains the annotation model using the updated data set. Through this iterative process, the model can be continuously optimized, improving its adaptability and accuracy to data. Especially when data changes or new types of data appear, it can quickly enhance the model's processing ability and generalization ability, thus ensuring the effectiveness and reliability of the model in practical applications.
[0088] In some embodiments, Figure 5 is shown Figure 1 a schematic flowchart of a method for performing binary classification screening on a newly added data set in step S500, screening low-similarity data based on vectorization processing and similarity calculation, and returning the low-similarity data to the training data set. As Figure 5 shown, step S500 includes:
[0089] Step S510, performing binary classification screening on the newly added data set through a trained annotation model to obtain annotated data.
[0090] This step applies the trained annotation model to the newly added data set, performs binary classification processing on the data, and automatically identifies and annotates the qualified data. Through this automatic screening, effective annotated data can be quickly and efficiently extracted from the newly added data, saving the time and cost of manual annotation and ensuring the efficiency of data processing.
[0091] Step S520, performing vectorization processing on the annotated data and historical data through a vector model, and converting the annotated data and historical data into vector representations.
[0092] In this step, a vector model is used to perform vectorization processing on the annotated data screened by the annotation model and the original historical data, and convert them into vector representations that can be processed by a computer. This process enables the data to be calculated and analyzed in a more concise and structured manner in subsequent steps. Vectorization processing enables the features of the dialogue data to be clearly expressed in a high-dimensional space, facilitating further operations such as similarity calculation and data clustering.
[0093] Step S530, calculating the vectorized similarity between the annotated data and the historical data, and screening out the low-similarity data with a similarity lower than a preset threshold.
[0094] By calculating the similarity between the vectorized annotated data and the historical data, data with a similarity lower than a preset threshold (e.g., 0.65) is screened out, and these data are used as training data. The purpose of similarity calculation is to identify and extract samples that are significantly different from or have anomalies compared to the existing data. In this way, data that is significantly different from the existing data set can be effectively identified, avoiding the impact of duplicate data on model training, ensuring the diversity and representativeness of the newly added data, and further improving the generalization ability of the model.
[0095] Step S540, returning the low-similarity data to the training data set.
[0096] By returning the filtered low-similarity data to the training dataset, the diversity of the training data can be increased, and the model can learn richer features during the training process. The addition of the returned data helps to improve the adaptability of the model, preventing the model from being trained only on common or high-similarity data, and thus enhancing the model's ability to recognize and process new types of data, enabling it to better handle different types of data in practical applications.
[0097] According to the above embodiments, by supporting various data acquisition methods such as API calls, database queries, and file downloads, the flexibility of data collection is improved to adapt to the data acquisition requirements from different sources. For the noise, duplicate values, and missing values in the data, a systematic cleaning strategy is adopted, including noise removal, duplicate data removal, and methods for handling missing values such as interpolation, deletion, or filling, to ensure the integrity and consistency of the data. In addition, by manually reviewing the target data and optimizing the screening of positive samples and difficult negative samples, the training dataset is made more accurate, enhancing the effectiveness of model training.
[0098] An embodiment of the present invention also provides a data annotation system based on multi-model collaboration. As Figure 6 shown, the data annotation system based on multi-model collaboration includes:
[0099] A data processing module 101, configured to clean and convert the format of the original data to generate structured data.
[0100] A data screening module 102, configured to screen the structured data using a pre-trained model and manually review the screened target data to obtain a training dataset.
[0101] A model training module 103, configured to train an annotation model based on the training dataset.
[0102] A data correction module 104, configured to run the trained annotation model on the test dataset, screen and manually correct the mislabeled data, and use the corrected data for retraining the model.
[0103] A data return module 105, configured to perform binary classification screening on the new data, screen the low-similarity data using vectorization processing and similarity calculation, and return the low-similarity data to the training dataset.
[0104] In the above data annotation system applicable to multi-model collaboration, for the specific implementation manners of data screening and return, refer to the relevant content of the embodiments in the above method, and details are not elaborated herein.
[0105] An embodiment of the present invention further provides a data annotation device based on multi-model collaboration, including: a processor 201, a memory 202, and a computer program stored in the memory 202 and configured to be executed by the processor 201. When the processor 201 executes the computer program, it implements the data annotation method based on multi-model collaboration as described in any of the above embodiments.
[0106] When the processor 201 executes the computer program, it implements the steps in the above embodiments of the data annotation method based on multi-model collaboration, such as Figure 1 all the steps of the data annotation method based on multi-model collaboration as shown. Alternatively, when the processor 201 executes the computer program, it implements the functions of each module / unit in the above data annotation system based on multi-model collaboration, such as Figure 6 the functions of each module of the data annotation system based on multi-model collaboration as shown.
[0107] Exemplarily, the computer program can be divided into one or more modules. One or more modules are stored in the memory 202 and executed by the processor 201 to complete the present invention. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the data annotation system based on multi-model collaboration.
[0108] The so-called processor 201 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 201 is the control center of the data annotation system based on multi-model collaboration, and connects various parts of the entire data annotation system based on multi-model collaboration through various interfaces and lines.
[0109] The memory 202 can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 202, and by invoking the data stored in the memory 202, the processor 201 realizes various functions of the multi-model collaboration-based data annotation system. The memory 202 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the multi-model collaboration-based data annotation system, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0110] Among them, if the module / unit of the multi-model collaboration-based data annotation system is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0111] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0112] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0113] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
Claims
1. A data annotation method based on multi-model collaboration, characterized in that, The method includes: Cleaning and formatting the original data to generate structured data; Using a pre-trained model to screen the structured data, and manually reviewing the screened target data to obtain a training data set; Training an annotation model based on the training data set; Running the trained annotation model on the test data set, screening and manually correcting mislabeled data, and using the corrected data for retraining the model; Performing binary classification screening on the new data set, and screening low-similarity data based on vectorization processing and similarity calculation, and returning the low-similarity data to the training data set.
2. The data annotation method based on multi-model collaboration according to claim 1, wherein The specific steps of cleaning and formatting the original data to generate structured data include: Obtaining the original data; Removing noise, duplicate data, and handling missing values from the original data to obtain cleaned data; Performing format conversion on the cleaned data to generate the structured data; Uploading the structured data to the server cluster.
3. The data annotation method based on multi-model collaboration according to claim 1, wherein The specific steps of using a pre-trained model to screen the structured data, and manually reviewing the screened target data to obtain a training data set include: Performing binary classification tagging on the structured data through the pre-trained model to screen out target data; Performing secondary manual tagging on the target data to obtain positive sample data and hard negative sample data; Combining the positive sample data and the hard negative sample data to form the training data set.
4. The data annotation method based on multi-model collaboration according to claim 1, wherein The specific steps of running the trained annotation model on the test data set, screening and manually correcting mislabeled data, and using the corrected data for retraining the model include: Applying the trained annotation model on the test data set, evaluating the annotation results and identifying mislabeled data; Manually reviewing and correcting the mislabeled data; Adding the corrected data to the training data set, and retraining the annotation model based on the updated data set.
5. The data annotation method based on multi-model collaboration according to claim 1, characterized in that The specific steps of performing binary classification screening on the new data set, and screening low-similarity data based on vectorization processing and similarity calculation, and returning the low-similarity data to the training data set include: Performing binary classification screening on the new data set through the trained annotation model to obtain annotated data; Performing vectorization processing on the annotated data and historical data through a vector model, and converting the annotated data and the historical data into vector representations; Calculating the vectorized similarity between the annotated data and the historical data, and screening out low-similarity data with a similarity lower than a preset threshold; Returning the low-similarity data to the training data set.
6. The data annotation method based on multi-model collaboration according to claim 2, characterized in that The specific steps of obtaining the original data include obtaining the original data from an online platform through API call, database query, or file download.
7. The data annotation method based on multi-model collaboration according to claim 2, wherein The specific steps of handling missing values include identifying missing values in the original data and using interpolation, deletion, or filling methods for processing.
8. A data annotation system adopting the data annotation method based on multi-model collaboration according to any one of claims 1-7, characterized in that, The system includes: A data processing module for cleaning and formatting the original data to generate structured data; A data screening module, which is used to screen structured data by using a pre-trained model, and perform manual review on the screened target data to obtain a training data set; A model training module, which is used to train an annotation model based on the training data set; A data correction module, which is used to run the trained annotation model on a test data set, screen and manually correct mislabeled data, and use the corrected data for retraining the model; A data feedback module, which is used to perform binary classification screening on new data, screen low-similarity data by using vectorization processing and similarity calculation, and feedback the low-similarity data to the training data set.
9. A device, characterized in that, Comprising: A processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the multi-model collaboration-based data annotation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the multi-model collaboration-based data annotation method according to any one of claims 1-7 are implemented.