Data processing method, device and equipment

By dividing the data set into multiple subsets and performing automatic cross-scoring of multi-model feedback, the problems of inefficiency and data leakage risks in large-scale data annotation are solved, and high-precision and efficient data pre-marking are achieved.

CN120030391APending Publication Date: 2025-05-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220196.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology has problems such as inefficiency, error prone and data leakage risks in the process of large-scale data labeling, and it is difficult to meet the needs of fast and high-quality labeling.

Method used

By obtaining the data set to be marked, dividing it into multiple data subsets, and label prediction processing and model fine-tuning are performed on each data subset based on the target model to realize an efficient data pre-marking mechanism of automatic cross-scoring of multi-model feedback.

Benefits of technology

It realizes high-precision and high-efficiency data pre-marking tasks, reduces labor costs, improves the quality and efficiency of data labeling, and reduces the risk of data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030391A_ABST
    Figure CN120030391A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method, device and equipment, and the method comprises the steps: obtaining a to-be-marked data set, and dividing the data set into a plurality of different data subsets; performing label prediction processing on the data in the plurality of different data subsets based on the target model to obtain a first prediction label subset corresponding to each data subset; performing fine tuning on the target model based on each data subset and the first prediction tag subset corresponding to each data subset to obtain a fine tuning target model corresponding to each data subset; performing label prediction processing on data in other data subsets based on the fine tuning target model corresponding to each data subset to obtain a second prediction label subset corresponding to each data subset; and based on the first prediction tag subset corresponding to each data subset and the second prediction tag subset corresponding to each data subset, determining tag information of data in the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of computer technology, and in particular to a data processing method, device and equipment. Background Art

[0002] At present, in the modeling tasks of machine learning and artificial intelligence, data labeling is indispensable for training and evaluating models. Especially at the current stage of rapid development of large model technology, large-scale data labeling tasks are even more important for improving model accuracy.

[0003] However, the usual labeling process requires a lot of professional training, manual participation and time cost. At the same time, subjective factors, human errors and fatigue of the labelers will affect the labeling results, resulting in errors in the labeling results, which directly affect the accuracy of model training. In addition, in large-scale data labeling tasks, the usual static manual labeling process lacks real-time feedback and correction, and it is difficult to cope with the label design that may change constantly for large-scale data, and thus cannot meet the needs of fast and high-quality labeling. It is inefficient and prone to errors. Moreover, as people pay more and more attention to their privacy data, large-scale manual labeling also increases the risk of data leakage. To this end, it is necessary to provide a better data labeling method, so as to achieve high-precision and high-efficiency data pre-labeling tasks, and complete high-efficiency and high-quality labeling tasks for large-scale data with minimal labor costs. Summary of the invention

[0004] The purpose of the embodiments of this specification is to provide a better data labeling method, so as to achieve high-precision and high-efficiency data pre-labeling tasks, and complete high-efficiency and high-quality labeling tasks for large-scale data with minimal labor costs.

[0005] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided in an embodiment of the present specification comprises: obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets. Based on the target model, label prediction processing is performed on the data in the plurality of different data subsets respectively, and a first prediction label subset corresponding to each data subset is obtained. Based on each data subset and the first prediction label subset corresponding to each data subset, the target model is fine-tuned respectively, and a fine-tuned target model corresponding to each data subset is obtained. Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets respectively, and a second prediction label subset corresponding to each data subset is obtained. Based on the first prediction label subset corresponding to each data subset and the second prediction label subset corresponding to each data subset, the label information of the data in the data set is determined.

[0006] A data processing device provided in an embodiment of the present specification comprises: a data acquisition module, which acquires a data set to be labeled and divides the data set into a plurality of different data subsets. A first label prediction module, which performs label prediction processing on the data in a plurality of different data subsets based on a target model, and obtains a first predicted label subset corresponding to each data subset. A model fine-tuning module, which fine-tunes the target model based on each data subset and the first predicted label subset corresponding to each data subset, and obtains a fine-tuned target model corresponding to each data subset. A second label prediction module, which performs label prediction processing on the data in other data subsets based on the fine-tuned target model corresponding to each data subset, and obtains a second predicted label subset corresponding to each data subset. A label determination module, which determines the label information of the data in the data set based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset.

[0007] A data processing device provided in an embodiment of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor: obtains a data set to be labeled, and divides the data set into a plurality of different data subsets. Based on the target model, label prediction processing is performed on the data in the plurality of different data subsets respectively, and a first predicted label subset corresponding to each data subset is obtained. Based on each data subset and the first predicted label subset corresponding to each data subset, the target model is fine-tuned respectively, and a fine-tuned target model corresponding to each data subset is obtained. Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets respectively, and a second predicted label subset corresponding to each data subset is obtained. Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, the label information of the data in the data set is determined.

[0008] The embodiments of this specification also provide a storage medium, which is used to store computer executable instructions. When the executable instructions are executed by the processor, they implement the following process: obtain a data set to be labeled, and divide the data set into multiple different data subsets. Based on the target model, label prediction processing is performed on the data in multiple different data subsets respectively to obtain a first predicted label subset corresponding to each data subset. Based on each data subset and the first predicted label subset corresponding to each data subset, the target model is fine-tuned respectively to obtain a fine-tuned target model corresponding to each data subset. Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets respectively to obtain a second predicted label subset corresponding to each data subset. Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, the label information of the data in the data set is determined.

[0009] The embodiments of this specification also provide a computer program product, including a computer program, which implements the following process when executed by a processor: obtain a data set to be labeled, and divide the data set into multiple different data subsets. Based on the target model, label prediction processing is performed on the data in multiple different data subsets respectively, and a first predicted label subset corresponding to each data subset is obtained. Based on each data subset and the first predicted label subset corresponding to each data subset, the target model is fine-tuned respectively to obtain a fine-tuned target model corresponding to each data subset. Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets respectively to obtain a second predicted label subset corresponding to each data subset. Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, the label information of the data in the data set is determined. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is an embodiment of a data processing method of this specification; Figure 2 Another data processing method embodiment of this specification; Figure 3 This is another data processing method embodiment of the present specification; Figure 4 This is another data processing method embodiment of the present specification; Figure 5 A schematic diagram of a data processing process of this specification; Figure 6 This is another data processing method embodiment of the present specification; Figure 7 This is an embodiment of a data processing device of the present specification; Figure 8 This is an embodiment of a data processing device in this specification. DETAILED DESCRIPTION

[0011] The embodiments of this specification provide a data processing method, device and equipment.

[0012] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0013] The embodiment of this specification provides an efficient data pre-labeling mechanism with multi-model feedback automatic cross-scoring. At present, in the modeling tasks of machine learning and artificial intelligence, the data labeling work performed for training and evaluating models is indispensable, especially at the current stage of rapid development of large model technology, large-scale data labeling tasks are more important for improving model accuracy. However, the usual labeling process requires a lot of professional training, manual participation and time cost. At the same time, the subjective factors, human errors and fatigue of the labelers will affect the labeling results, resulting in errors in the labeling results, which directly affect the accuracy of model training. In addition, in large-scale data labeling tasks, the usual static manual labeling process lacks real-time feedback and correction, and it is difficult to cope with the label design that may change constantly for large-scale data, and thus cannot meet the requirements of fast and high-quality labeling, which is inefficient and prone to errors. In the embodiment of this specification, an efficient data pre-labeling mechanism with automatic cross-scoring of multi-model feedback is designed, which realizes high-precision and high-efficiency data pre-labeling tasks by introducing multi-model feedback mechanisms and data set label update technologies, and completes high-efficiency and high-quality labeling tasks for large-scale data with minimal labor costs. For specific processing, please refer to the specific content in the following embodiments.

[0014] like Figure 1As shown, an embodiment of this specification provides a data processing method, and the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car device, etc.), etc., wherein the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server in the financial field or the online shopping field, or a background server of an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server situation processing, which will not be repeated here. The method can specifically include the following steps: In step S102, a data set to be labeled is obtained, and the data set is divided into a plurality of different data subsets.

[0015] Among them, the data set to be labeled can be a data set for which label information needs to be set. The data in the data set can include multiple types, for example, transaction data between different users, interaction data between different users, image data, audio data, video data, etc., which can be set specifically according to actual conditions.

[0016] In implementation, the data set to be labeled can be obtained in a variety of different ways. For example, the data to be labeled can be obtained from a specified database, and a data set can be constructed based on the obtained data to be labeled. The database can be a database for storing data to be labeled or a database for storing sample data for training a certain model, etc.; or, relevant data can be collected from different users in the process of using a specified service. When the above-mentioned relevant data is needed, all the data or a certain amount of data can be obtained from the server of the above-mentioned service, and the data set to be labeled can be constructed based on the obtained data, etc. The specific settings can be based on actual conditions.

[0017] In order to use different data to train the specified model separately in the future, the above data set can be divided into multiple different data subsets, for example, the data set can be divided into 2 data subsets, or the data set can be divided into 3 data subsets, or the data set can be divided into 5 data subsets, etc. In addition, the above data set can be randomly divided into multiple different data subsets, or the data set can be divided into multiple data subsets that meet the specified conditions based on the specified allocation rules. In addition, the above data set can be evenly divided into multiple different data subsets, at this time, the number of data contained in each data subset is the same, or the above data set can be divided into multiple data subsets whose numbers meet the specified conditions, for example, the data set is divided into 3 data subsets, wherein one data subset contains 1000 data, one data subset contains 1500 data, and another data subset contains 1200 data, or one data subset contains 1000 data, one data subset contains 1300 data, and another data subset contains 1000 data, etc., which can be set according to actual conditions.

[0018] In step S104, label prediction processing is performed on the data in a plurality of different data subsets based on the target model to obtain a first predicted label subset corresponding to each data subset.

[0019] Among them, the target model can be any model, for example, the target model can be a risk identification model, a facial recognition model, or a data classification model, etc. The target model can be constructed by a variety of different algorithms or networks. For example, the target model can be constructed by a clustering algorithm, or it can be constructed by a specified neural network, or it can be constructed by a specified large model (such as a large language model, etc.), or it can also be constructed by a tree-structured algorithm (specifically such as a decision tree algorithm or a random forest algorithm, etc.), etc., and can be set according to actual conditions.

[0020] In implementation, a target model can be obtained, which can be a pre-trained model that can classify and label data. For any data subset, each data in the data subset can be input into the target model respectively, and the label prediction processing is performed on each data through the target model to obtain the first predicted label information corresponding to each data. Finally, the first predicted label subset consisting of the first predicted label information corresponding to each data in the data subset can be obtained. In the above manner, label prediction processing can be performed on the data in multiple other different data subsets based on the target model to obtain the first predicted label subset corresponding to each other data subset, thereby obtaining the first predicted label subset corresponding to each data subset in multiple different data subsets.

[0021] In addition, for the processing of the above-mentioned steps S102 and S104, corresponding effects can also be obtained through other methods. Specifically, a data set to be labeled is obtained, and label prediction processing is performed on the data in the data set based on the target model to obtain predicted label information corresponding to the data set; the data set is divided into multiple different data subsets, and accordingly, the predicted label information corresponding to the data set is divided into corresponding predicted label subsets to obtain a first predicted label subset corresponding to each data subset in multiple different data subsets.

[0022] In step S106, the target model is fine-tuned based on each data subset and the first prediction label subset corresponding to each data subset, so as to obtain a fine-tuned target model corresponding to each data subset.

[0023] In implementation, after obtaining the first prediction label subset corresponding to each data subset in a plurality of different data subsets by the above method, the target model can be fine-tuned based on each data subset and the first prediction label subset corresponding to each data subset. Specifically, for any data subset, the target model can be further trained using the data in the data subset and the first prediction label information corresponding to the first prediction label subset corresponding to the data subset to fine-tune the model parameters of the target model and obtain the fine-tuned target model based on the data subset, i.e., the fine-tuned target model corresponding to the data subset. The fine-tuned target models corresponding to the other data subsets can be obtained by the above method. For example, the data set is divided into three data subsets, and the target model can be further trained using the data in the first data subset and the corresponding first prediction label information in the first prediction label subset corresponding to the first data subset to fine-tune the model parameters of the target model to obtain a target model fine-tuned based on the first data subset, that is, the fine-tuned target model corresponding to the first data subset. In the above manner, the fine-tuned target model corresponding to the second data subset and the fine-tuned target model corresponding to the third data subset can be obtained respectively.

[0024] In step S108, based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset.

[0025] Among them, other data subsets can be data subsets other than the data subsets used when fine-tuning the target model. For example, the data set is divided into three data subsets. The fine-tuning target model corresponding to the first data subset can perform label prediction processing on the data in the second data subset, and can perform label prediction processing on the data in the third data subset. The fine-tuning target model corresponding to the second data subset can perform label prediction processing on the data in the first data subset, and can perform label prediction processing on the data in the third data subset. The fine-tuning target model corresponding to the third data subset can perform label prediction processing on the data in the first data subset, and can perform label prediction processing on the data in the second data subset.

[0026] In implementation, for example, the data set is divided into three data subsets. Through the above processing, the fine-tuning target model corresponding to the first data subset, the fine-tuning target model corresponding to the second data subset, and the fine-tuning target model corresponding to the third data subset can be obtained. The fine-tuning target model corresponding to the first data subset can be used to perform label prediction processing on the data in the second data subset and the label prediction processing on the data in the third data subset, respectively, to obtain the second predicted label information corresponding to the data in the second data subset and the second predicted label information corresponding to the data in the third data subset, that is, to obtain the second predicted label subset corresponding to the second data subset and the second predicted label subset corresponding to the third data subset, and the fine-tuning target model corresponding to the second data subset can be used to perform label prediction processing on the data in the first data subset, respectively. The row label prediction processing is performed and the label prediction processing is performed on the data in the third data subset to obtain the second predicted label information corresponding to the data in the first data subset and the second predicted label information corresponding to the data in the third data subset, that is, the second predicted label subset corresponding to the first data subset and the second predicted label subset corresponding to the third data subset are obtained, and the fine-tuning target model corresponding to the third data subset can be used to perform label prediction processing on the data in the first data subset and the label prediction processing on the data in the second data subset, respectively, to obtain the second predicted label information corresponding to the data in the first data subset and the second predicted label information corresponding to the data in the second data subset, that is, the second predicted label subset corresponding to the first data subset and the second predicted label subset corresponding to the second data subset are obtained.

[0027] In step S110, label information of the data in the above data set is determined based on the first prediction label subset corresponding to each data subset and the second prediction label subset corresponding to each data subset.

[0028] In implementation, for any data subset, the first predicted label information and the second predicted label information corresponding to each data in the data subset can be compared. If the two are the same, it indicates that the predicted label information of the data is accurate. At this time, the first predicted label information or the second predicted label information can be used as the label information of the data. If the two are different, it indicates that the prediction of the label information of the data is inaccurate. At this time, the predicted label information of the data can be adjusted. For example, the predicted label information of the data can be adjusted through expert experience, or the data can be input into a new model or network for prediction to obtain new label information. The label information of the data can be determined based on the new label information, the first predicted label information and the second predicted label information. In the above manner, the label information of other data in the data subset can be determined, and then the label information of the data in the data subset can be obtained. In the above manner, the label information of the data in multiple different data subsets can be obtained, and finally, the label information of the data in the data set can be obtained.

[0029] The embodiments of the present specification provide a data processing method, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. The predicted label subset and the second predicted label subset corresponding to each data subset determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and complete the high-efficiency and high-quality labeling tasks of large-scale data with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and the high-efficiency and high-quality labeling tasks of large-scale data are completed with minimal labor costs.

[0030] In practical applications, a cold start phase can be set, and the target model can be trained in the cold start phase. Specifically, label prediction processing can be performed on data in multiple different data subsets based on the target model, and before obtaining the first predicted label subset corresponding to each data subset, the cold start phase processing can be performed, such as Figure 2 As shown, please refer to the processing of the following steps S112 to S116 for details.

[0031] In step S112, a first amount of data is sampled from the data set.

[0032] The first number can be set according to actual conditions, for example, 5000 or 10000.

[0033] In implementation, a small amount of data (i.e., a first amount of data) can be sampled from a large amount of data sets to be labeled. In the process of sampling the first amount of data, the first amount of data can be collected from a large amount of data sets to be labeled by random sampling, or by specified sampling rules. The sampling rules may include multiple types, for example, collecting data of a specified type, such as collecting data containing specified keywords or key words, or collecting data with a data volume greater than a preset threshold, or collecting data of users in a specified area, or collecting data generated within a specified time period, etc. The specific setting can be based on actual conditions, and the embodiments of this specification do not limit this.

[0034] In step S114, first label information of each data in the first amount of data is obtained.

[0035] In implementation, a first amount of data can be sent to a management terminal, and the first amount of data is used to trigger a user of the management terminal to mark each data in the first amount of data through preset marking rules. The marking rules include one or more of marking rules based on expert experience and marking rules constructed based on a marking network.

[0036] After the management terminal receives the first amount of data, the user of the management terminal can perform labeling processing on each data in the first amount of data based on the labeling rules of expert experience to obtain the first label information of each data in the first amount of data, or the user of the management terminal can input each data in the first amount of data into the labeling network for labeling processing to obtain the first label information of each data in the first amount of data, etc. After obtaining the first label information of each data in the first amount of data in the above manner, the management terminal can send the first label information of each data in the first amount of data to the server, and the server can obtain the first label information of each data in the first amount of data.

[0037] In step S116, model training is performed on the target model based on the first amount of data and the first label information of each data in the first amount of data to obtain a trained target model.

[0038] In implementation, the loss function can be pre-set according to the actual situation, such as the mean square error loss function, the cross entropy loss function, etc. Then, each data in the first amount of data can be input into the target model to obtain the corresponding output result. Based on the output result and the first label information of the corresponding data in the first amount of data, the corresponding loss information can be calculated by the above loss function. The model parameters of the target model can be adjusted based on the calculated loss information to train the target model until the above loss function converges to obtain the trained target model. The trained target model can be applied to the subsequent step S104 (i.e., label prediction processing is performed on the data in multiple different data subsets based on the trained target model to obtain the first predicted label subset corresponding to each data subset) and its subsequent steps.

[0039] In practical applications, the specific processing methods of the above step S110 can be varied. An optional processing method is provided below, such as Figure 3 As shown, the processing may specifically include the following steps S1102 and S1104.

[0040] In step S1102, for any data subset, the first prediction label subset is compared with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected to determine the label information of the data in any of the above data subsets.

[0041] In implementation, for any data subset, the first prediction label information and the second prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset can be compared. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, the different prediction label information can be corrected. For example, the data set is divided into three data subsets, namely the first data subset, the second data subset and the third data subset. The first data subset includes data A and data B. Data A corresponds to the first prediction label information and the second prediction label information, and data B corresponds to the first prediction label information and the second prediction label information. For the first data subset, the first prediction label information and the second prediction label information corresponding to data A can be compared, and the first prediction label information and the second prediction label information corresponding to data B can be compared. If the above comparison results of data A and data B are both different, the prediction label information of data A and data B can be corrected respectively to obtain the label information of the data in the first data subset.

[0042] In step S1104, based on the processing method of any of the data subsets, label information of the data in the data set is determined.

[0043] In implementation, by processing the first data subset, label information of the data in the second data subset and label information of the data in the third data subset can be obtained respectively, thereby obtaining label information of the data in the above data sets.

[0044] In addition, if the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold, it can be determined that the prediction label information corresponding to each data is more accurate. At this time, the first prediction label information or the second prediction label information can be directly used as the label information of the corresponding data, thereby obtaining the label information of the data in the above data set.

[0045] In practical applications, the specific processing methods of the above step S1102 can be varied. An optional processing method is provided below, such as Figure 4 As shown, please refer to the processing of steps S11022 to S11028 below for details.

[0046] In step S11022, for any data subset, the first prediction label subset is compared with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected, and the corrected prediction label information corresponding to the data in each data subset is determined to obtain the corrected prediction label subset corresponding to each data subset.

[0047] The specific processing process of the above step S11022 can be found in the above related content and will not be repeated here.

[0048] In step S11024, the revised prediction label subset corresponding to each data subset is used as the first prediction label subset corresponding to each data subset, and the target model is fine-tuned based on each data subset and the first prediction label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset.

[0049] In implementation, Figure 5 As shown, through the above processing, a revised prediction label subset corresponding to each data subset is obtained, and the processing of the above steps S106 to S110 can be executed in a loop, that is, the revised prediction label subset corresponding to each data subset is used as the first prediction label subset corresponding to each data subset, and the target model is fine-tuned based on each data subset and the first prediction label subset corresponding to each data subset, so as to obtain a fine-tuned target model corresponding to each data subset; based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets, so as to obtain a second prediction label subset corresponding to each data subset (that is, the following step S11026), and the specific processing method of step S110 can be processed by steps S1102 and S1104.

[0050] In step S11026, based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset.

[0051] The specific processing procedures of the above-mentioned steps S11024 and S11026 can be found in the aforementioned related content and will not be repeated here.

[0052] In step S11028, for any data subset, the first prediction label subset is compared with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold, the label information of the data in any of the above data subsets is determined.

[0053] In implementation, Figure 5 As shown (taking the example of dividing the data set into two data subsets for explanation), for any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds the preset number threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected, and the corrected prediction label information corresponding to the data in each data subset is determined to obtain the corrected prediction label subset corresponding to each data subset. The above-mentioned processing of S11022 to S11028 can be repeatedly performed until the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold. At this time, it indicates that the marking quality of the label information corresponding to the data in the data set has reached the preset requirement. Finally, the label information of the data in any data subset can be obtained, and then the label information of the data in each data subset can be obtained, thereby obtaining the label information of the data in the data set.

[0054] In actual applications, the specific processing methods of the above step S1102 can be varied. An optional processing method is provided below. For details, please refer to the following content: for any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset threshold, then select a second number of different prediction label information corresponding to the same data from the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset, correct the selected different prediction label information corresponding to the same data, and determine the label information of the data in any of the above data subsets.

[0055] The second quantity can be set according to actual conditions, such as 80% or 90% of the total quantity.

[0056] In implementation, in the process of correcting the prediction label information, the different prediction label information corresponding to the same data can be corrected, or some of them can be selected for correction. For example, if the prediction label information corresponding to 100 data in the data subset is different, the different prediction label information corresponding to 80 data can be selected for correction. In practical applications, selecting the different prediction label information of some data for correction can include multiple ways. For example, the difference value between the different prediction label information of the same data can be calculated, and the data can be sorted from large to small according to the calculated difference value, and the second number of data arranged in the front can be selected, and then the different prediction label information corresponding to the selected second number of data can be corrected, or the different prediction label information corresponding to the second number of data can be randomly selected for correction, etc., which can be set according to actual conditions. Then, the above steps S106 to S110 can be executed cyclically until the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold, and finally the label information of the data in any data subset is obtained, and then the label information of the data in each data subset is obtained, so as to obtain the label information of the data in the data set.

[0057] In practical applications, there may be various specific processing methods for correcting different predicted label information corresponding to the same data. An optional processing method is provided below, which may specifically include the following steps A2 and A4.

[0058] In step A2, different predicted label information corresponding to the same data is sent to the management terminal. The different predicted label information corresponding to the same data is used to trigger the user of the management terminal to correct the different predicted label information corresponding to the same data through preset label correction rules. The label correction rules include one or more of label correction rules based on expert experience and label correction rules constructed based on a label correction network.

[0059] Among them, the label correction network can include multiple types, such as convolutional neural network, recurrent neural network, etc., which can be set according to actual conditions.

[0060] In step A4, the corrected predicted label information returned by the management terminal is received.

[0061] In practical applications, the target model may include a model built based on a preset neural network, a model built based on a linear classifier, or a model built based on a Transformer module.

[0062] In practical applications, in order to simplify the processing process, the divided data subsets include two, which can be the first data subset and the second data subset. Based on this, the specific processing methods of the above step S108 can be various. The following is an optional processing method, such as Figure 6 As shown, the processing may specifically include the following steps S1082 and S1084.

[0063] In step S1082, label prediction processing is performed on the data in the second data subset based on the fine-tuned target model corresponding to the first data subset to obtain a second predicted label subset corresponding to the second data subset.

[0064] In step S1084, label prediction processing is performed on the data in the first data subset based on the fine-tuned target model corresponding to the second data subset to obtain a second predicted label subset corresponding to the first data subset.

[0065] Based on the above data subsets, there are two cases, such as Figure 5 As shown, the processing of the above step S104 may include: performing label prediction processing on the data in the first data subset based on the target model to obtain the first predicted label subset corresponding to the first data subset; performing label prediction processing on the data in the second data subset based on the target model to obtain the first predicted label subset corresponding to the second data subset. The processing of the above step S106 may include: fine-tuning the target model based on the first data subset and the first predicted label subset corresponding to the first data subset to obtain the fine-tuned target model corresponding to the first data subset; fine-tuning the target model based on the second data subset and the first predicted label subset corresponding to the second data subset to obtain the fine-tuned target model corresponding to the second data subset. The processing of the above step S110 may include: determining the label information of the data in the data set based on the first predicted label subset corresponding to the first data subset, the second predicted label subset corresponding to the first data subset, the first predicted label subset corresponding to the second data subset, and the second predicted label subset corresponding to the second data subset. Among them, the specific processing of determining the label information of the data in the data set based on the first predicted label subset corresponding to the first data subset, the second predicted label subset corresponding to the first data subset, the first predicted label subset corresponding to the second data subset and the second predicted label subset corresponding to the second data subset can be referred to the aforementioned related content and will not be repeated here.

[0066] The embodiments of the present specification provide a data processing method, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. The predicted label subset and the second predicted label subset corresponding to each data subset determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and complete the high-efficiency and high-quality labeling tasks of large-scale data with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and the high-efficiency and high-quality labeling tasks of large-scale data are completed with minimal labor costs.

[0067] In addition, during the cold start phase, a small amount of data is sampled for labeling to train the target model and provide initial label information for subsequent label prediction feedback of the target model and update of data label information. In addition, dual-model or multi-model automatic cross-prediction is provided: by dividing two or more data subsets, iterating two or more target models in real time, and dual-model or multi-model cross-prediction mechanisms, a small amount of risk data that is incorrectly labeled during the cold start phase or early stages can be quickly and efficiently located, and label correction rules can be accurately, timely and efficiently introduced for correction processing (such as manual correction of label information, etc.). In addition, the idea of ​​active learning is adopted, and the label information of the data set is updated with the principle of minimum manual input by cross-comparing the differences between the model prediction results and the historical label information, reducing the workload of manual labeling and improving the efficiency of data pre-labeling.

[0068] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 7 shown.

[0069] The data processing device includes: a data acquisition module 701, a first label prediction module 702, a model fine-tuning module 703, a second label prediction module 704 and a label determination module 705, wherein: The data acquisition module 701 acquires a data set to be labeled and divides the data set into a plurality of different data subsets; A first label prediction module 702 performs label prediction processing on the data in a plurality of different data subsets based on the target model to obtain a first predicted label subset corresponding to each data subset; A model fine-tuning module 703, which fine-tunes the target model based on each data subset and the first prediction label subset corresponding to each data subset, to obtain a fine-tuned target model corresponding to each data subset; The second label prediction module 704 performs label prediction processing on the data in other data subsets based on the fine-tuned target model corresponding to each data subset, so as to obtain a second predicted label subset corresponding to each data subset; The label determination module 705 determines label information of the data in the data set based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset.

[0070] In the embodiment of this specification, the device further includes: A data sampling module, sampling a first amount of data from the data set; A first label acquisition module, acquiring first label information of each data in the first amount of data; The model training module performs model training on the target model based on the first amount of data and the first label information of each data in the first amount of data to obtain a trained target model.

[0071] In the embodiment of this specification, the tag determination module 705 includes: A first label determination unit, for any data subset, compares the first prediction label subset with the second prediction label subset, and if the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, performs correction processing on the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset, and determines the label information of the data in any data subset; The second label determination unit determines label information of the data in the data set based on the processing method of any data subset.

[0072] In an embodiment of the present specification, the first label determination unit compares the first prediction label subset with the second prediction label subset for any data subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected, and the corrected prediction label information corresponding to the data in each data subset is determined to obtain the corrected prediction label subset corresponding to each data subset; the corrected prediction label subset corresponding to each data subset is used as the first prediction label corresponding to each data subset According to the invention, the target model is fine-tuned based on each data subset and the first prediction label subset corresponding to each data subset to obtain the fine-tuned target model corresponding to each data subset; the label prediction processing is performed on the data in other data subsets based on the fine-tuned target model corresponding to each data subset to obtain the second prediction label subset corresponding to each data subset; for any data subset, the first prediction label subset is compared with the second prediction label subset, and if the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold, the label information of the data in any data subset is determined.

[0073] In an embodiment of the present specification, the first label determination unit compares the first prediction label subset with the second prediction label subset for any data subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset threshold, a second number of different prediction label information corresponding to the same data is selected from the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset, and correction processing is performed on the selected different prediction label information corresponding to the same data to determine the label information of the data in any data subset.

[0074] In an embodiment of the present specification, the first label determination unit sends different predicted label information corresponding to the same data to a management terminal, and the different predicted label information corresponding to the same data is used to trigger a user of the management terminal to correct the different predicted label information corresponding to the same data through a preset label correction rule, wherein the label correction rule includes one or more of a label correction rule based on expert experience and a label correction rule constructed based on a label correction network; and receives the corrected predicted label information returned by the management terminal.

[0075] In the embodiments of this specification, the target model includes a model constructed based on a preset neural network, a model constructed based on a linear classifier, or a model constructed based on a Transformer module.

[0076] In the embodiment of this specification, the divided data subset includes a first data subset and a second data subset, and the second label prediction module 704 includes: A first label prediction unit, performing label prediction processing on the data in the second data subset based on the fine-tuned target model corresponding to the first data subset, to obtain a second predicted label subset corresponding to the second data subset; The second label prediction unit performs label prediction processing on the data in the first data subset based on the fine-tuned target model corresponding to the second data subset to obtain a second predicted label subset corresponding to the first data subset.

[0077] The embodiment of the present specification provides a data processing device, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets respectively to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets respectively to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. The predicted label subset and the second predicted label subset corresponding to each data subset determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and complete the high-efficiency and high-quality labeling tasks of large-scale data with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and the high-efficiency and high-quality labeling tasks of large-scale data are completed with minimal labor costs.

[0078] In addition, during the cold start phase, a small amount of data is sampled for labeling to train the target model and provide initial label information for subsequent label prediction feedback of the target model and update of data label information. In addition, dual-model or multi-model automatic cross-prediction is provided: by dividing two or more data subsets, iterating two or more target models in real time, and dual-model or multi-model cross-prediction mechanisms, a small amount of risk data that is incorrectly labeled during the cold start phase or early stages can be quickly and efficiently located, and label correction rules can be accurately, timely and efficiently introduced for correction processing (such as manual correction of label information, etc.). In addition, the idea of ​​active learning is adopted, and the label information of the data set is updated with the principle of minimum manual input by cross-comparing the differences between the model prediction results and the historical label information, reducing the workload of manual labeling and improving the efficiency of data pre-labeling.

[0079] The above is a data processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 8 shown.

[0080] The data processing device may provide a terminal device or a server, etc. for the above-mentioned embodiments.

[0081] The data processing device may have relatively large differences due to different configurations or performances, and may include one or more processors 801 and memory 802, and the memory 802 may store one or more storage applications or data. Among them, the memory 802 may be a short-term storage or a persistent storage. The application stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the data processing device. Furthermore, the processor 801 may be configured to communicate with the memory 802 and execute a series of computer executable instructions in the memory 802 on the data processing device. The data processing device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input and output interfaces 805, and one or more keyboards 806.

[0082] Specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs are configured to be executed by one or more processors, including computer executable instructions for performing the following: Obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets; Based on the target model, label prediction processing is performed on the data in multiple different data subsets to obtain a first predicted label subset corresponding to each data subset; Fine-tune the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, label information of the data in the data set is determined.

[0083] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0084] The embodiment of the present specification provides a data processing device, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets respectively to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets respectively to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, The predicted label subset and the second predicted label subset corresponding to each data subset determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and complete the high-efficiency and high-quality labeling tasks of large-scale data with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and the high-efficiency and high-quality labeling tasks of large-scale data are completed with minimal labor costs.

[0085] Furthermore, based on the above Figures 1 to 6In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets; Based on the target model, label prediction processing is performed on the data in multiple different data subsets to obtain a first predicted label subset corresponding to each data subset; Fine-tune the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, label information of the data in the data set is determined.

[0086] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0087] The embodiment of the present specification provides a storage medium, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets respectively to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets respectively to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. The measured label subset and the second predicted label subset corresponding to each data subset are used to determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and high-efficiency and high-quality labeling tasks for large-scale data are completed with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and high-efficiency and high-quality labeling tasks for large-scale data are completed with minimal labor costs.

[0088] Furthermore, based on the above Figures 1 to 6 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets; Based on the target model, label prediction processing is performed on the data in multiple different data subsets to obtain a first predicted label subset corresponding to each data subset; Fine-tune the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, label information of the data in the data set is determined.

[0089] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0090] The embodiments of the present specification provide a computer program product, which obtains a data set to be labeled and divides the data set into multiple different data subsets. Then, based on the target model, label prediction processing can be performed on the data in the multiple different data subsets to obtain a first predicted label subset corresponding to each data subset. After that, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. Then, based on the fine-tuned target model corresponding to each data subset, label prediction processing can be performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset. Finally, based on the first predicted label subset corresponding to each data subset, the target model can be fine-tuned based on each data subset and the first predicted label subset corresponding to each data subset to obtain a fine-tuned target model corresponding to each data subset. The predicted label subset and the second predicted label subset corresponding to each data subset determine the label information of the data in the data set. In this way, an efficient data pre-labeling mechanism for automatic cross-label prediction based on multi-model feedback is introduced to achieve high-precision and high-efficiency data pre-labeling tasks, and complete the high-efficiency and high-quality labeling tasks of large-scale data with minimal labor costs. In addition, by dividing the data set into multiple different data subsets and then performing iterative optimization of multiple models, and performing efficient and automatic cross-label prediction through the predicted label information fed back by each of the multiple models, a small amount of potentially mislabeled data with inconsistent predicted label information is mined for further correction, and the high-efficiency and high-quality labeling tasks of large-scale data are completed with minimal labor costs.

[0091] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0092] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0093] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0094] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0095] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0096] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0097] The embodiments of this specification are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0098] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0099] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0100] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0101] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0102] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0103] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0104] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0106] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0107] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.

Claims

1. A data processing method, the method comprising: Obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets; Based on the target model, label prediction processing is performed on the data in multiple different data subsets to obtain a first predicted label subset corresponding to each data subset; Fine-tune the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, label information of the data in the data set is determined.

2. According to the method of claim 1, before performing label prediction processing on the data in a plurality of different data subsets based on the target model to obtain a first predicted label subset corresponding to each data subset, the method further comprises: sampling a first amount of data from the data set; Obtaining first label information of each data in the first amount of data; Model training is performed on the target model based on the first amount of data and first label information of each data in the first amount of data to obtain a trained target model.

3. The method according to claim 2, wherein determining the label information of the data in the data set based on the first prediction label subset corresponding to each data subset and the second prediction label subset corresponding to each data subset comprises: For any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, perform correction processing on the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset to determine the label information of the data in any data subset. Based on the processing method of any one of the data subsets, label information of the data in the data set is determined.

4. The method according to claim 3, wherein for any data subset, the first prediction label subset is compared with the second prediction label subset, and if the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected to determine the label information of the data in the any data subset, including: For any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, perform correction processing on the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset, determine the corrected prediction label information corresponding to the data in each data subset, and obtain the corrected prediction label subset corresponding to each data subset; Taking the modified prediction label subset corresponding to each data subset as the first prediction label subset corresponding to each data subset, and fine-tuning the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; For any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset does not exceed the preset number threshold, determine the label information of the data in the any data subset.

5. The method according to claim 3, wherein for any data subset, the first prediction label subset is compared with the second prediction label subset, and if the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset number threshold, the different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset is corrected to determine the label information of the data in the any data subset, including: For any data subset, compare the first prediction label subset with the second prediction label subset. If the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset exceeds a preset threshold, select a second number of different prediction label information corresponding to the same data from the number of different prediction label information corresponding to the same data in the first prediction label subset and the second prediction label subset, perform correction processing on the selected different prediction label information corresponding to the same data, and determine the label information of the data in any of the data subsets.

6. The method according to claim 4 or 5, wherein the correction process is performed on different predicted label information corresponding to the same data, comprising: Sending different predicted label information corresponding to the same data to a management terminal, wherein the different predicted label information corresponding to the same data is used to trigger a user of the management terminal to perform correction processing on the different predicted label information corresponding to the same data using a preset label correction rule, wherein the label correction rule includes one or more of a label correction rule based on expert experience and a label correction rule constructed based on a label correction network; The corrected predicted label information returned by the management terminal is received.

7. According to the method described in any one of claims 1-5, the target model comprises a model constructed based on a preset neural network, a model constructed based on a linear classifier, or a model constructed based on a Transformer module.

8. According to the method of claim 7, the divided data subsets include a first data subset and a second data subset, and the fine-tuning target model corresponding to each data subset is used to perform label prediction processing on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset, including: Performing label prediction processing on the data in the second data subset based on the fine-tuned target model corresponding to the first data subset to obtain a second predicted label subset corresponding to the second data subset; Based on the fine-tuned target model corresponding to the second data subset, label prediction processing is performed on the data in the first data subset to obtain a second predicted label subset corresponding to the first data subset.

9. A data processing device, comprising: A data acquisition module, which acquires a data set to be labeled and divides the data set into a plurality of different data subsets; A first label prediction module performs label prediction processing on the data in a plurality of different data subsets based on the target model to obtain a first predicted label subset corresponding to each data subset; A model fine-tuning module, which fine-tunes the target model based on each data subset and the first prediction label subset corresponding to each data subset, to obtain a fine-tuned target model corresponding to each data subset; A second label prediction module performs label prediction processing on the data in other data subsets based on the fine-tuned target model corresponding to each data subset, to obtain a second predicted label subset corresponding to each data subset; The label determination module determines label information of the data in the data set based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset.

10. A data processing device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Obtaining a data set to be labeled, and dividing the data set into a plurality of different data subsets; Based on the target model, label prediction processing is performed on the data in multiple different data subsets to obtain a first predicted label subset corresponding to each data subset; Fine-tune the target model based on each data subset and the first prediction label subset corresponding to each data subset, respectively, to obtain a fine-tuned target model corresponding to each data subset; Based on the fine-tuned target model corresponding to each data subset, label prediction processing is performed on the data in other data subsets to obtain a second predicted label subset corresponding to each data subset; Based on the first predicted label subset corresponding to each data subset and the second predicted label subset corresponding to each data subset, label information of the data in the data set is determined.