Data processing method and device, electronic equipment and storage medium
By processing the integrity of the labeled data determined by the inference results in the online training mode, and using multiple data sources to improve the quality of the training data set, the algorithm model is trained online, the problem of low online fine-tuning model accuracy is solved, and higher model accuracy and training efficiency is achieved.
Patent Information
- Application Number
- CN202411932460.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, in the online training mode, model inference and training are deployed in the same system, resulting in a simple labeling device, a small size of labeling samples, a low quality, and a limited computing power, resulting in a low accuracy of the online fine-tuning model.
The first label data and the second label data of the first sample image are determined by the inference result of the first sample image, and the integrity of these label data is processed to obtain high-quality target label data. At the same time, the second sample images and third annotation data of different data sources are used to improve the sample diversity and quality of the training data set, so that the algorithm model is trained online through the training data set.
The accuracy of the online fine-tuning model is improved, and the model accuracy is obtained through the processing of finite samples, and the data redundancy is used to exchange for improved training efficiency in the inference process.
Smart Images

Figure CN120047764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In the actual application of current large CV models, there are two training modes. One is the traditional offline training mode, and the other is the online training mode, where the result data generated by inference can be directly used as training data for continuous iterative update. In the traditional offline training mode, model inference and training are independently deployed. However, the offline mode has high training costs, low efficiency, long cycles, high professionalism, and data security issues. In the online training mode, model inference and training are deployed in the same system, and the result data generated by model inference is directly used as training data for continuous iterative update. However, online mode training usually faces problems such as simple annotation devices, small amounts of annotated samples, low quality of annotated samples, and extremely limited computing power, resulting in low accuracy of online fine-tuning models. Summary of the Invention
[0003] Embodiments of the present invention provide a data processing method, aiming to provide a new data processing solution to improve the accuracy of online fine-tuning models. By using the inference result of the first sample image to determine the first annotation data and the second annotation data of the first sample image, and performing integrity processing on the first annotation data and the second annotation data to obtain the first target annotation data corresponding to the first sample image, the annotation quality of the first sample image is improved. At the same time, by using the second sample image from different data sources and the corresponding third annotation data, the second target annotation data corresponding to the second sample image is obtained. By using the second sample image and the second target annotation data, the sample diversity of the training dataset is improved, and the sample quality of the training dataset is improved. By using this training dataset to perform online training on the algorithm model, the accuracy of the online fine-tuning model can be improved, higher model accuracy can be obtained through the processing of limited samples, and at the same time, the training efficiency is improved by means of data redundancy in the inference link.
[0004] In a first aspect, embodiments of the present invention provide a data processing method, the method comprising the following steps:
[0005] Obtain first data to be processed and second data to be processed, where the first data to be processed includes a first sample image, and a first annotation data and a second annotation data corresponding to the first sample image, the first annotation data and the second annotation data being determined according to the inference result obtained after the algorithm model performs inference processing on the first sample image, the second data to be processed includes a second sample image, and a third annotation data corresponding to the second sample image, and the first sample image and the second sample image are images from different data sources;
[0006] Perform first data integrity processing on the first labeled data and the second labeled data to obtain the first target labeled data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third labeled data to obtain the second target labeled data corresponding to the second sample image;
[0007] Construct a training dataset based on the first sample image, the first target labeled data corresponding to the first sample image, the second sample image, and the second target labeled data corresponding to the second sample image, and the training dataset is used for online training of the target algorithm model.
[0008] Optionally, the obtaining of the first data to be processed includes:
[0009] Obtain the image to be inferred;
[0010] Perform inference processing on the image to be inferred through the algorithm model to obtain the inference result of the image to be inferred;
[0011] Determine the first sample image in the image to be inferred;
[0012] Based on the inference result corresponding to the first sample image, determine the first labeled data and the second labeled data corresponding to the first sample image.
[0013] Optionally, the determining of the first labeled data and the second labeled data corresponding to the first sample image based on the inference result corresponding to the first sample image includes:
[0014] Send the first sample image and the inference result corresponding to the first sample image to the first user, and obtain the first labeled data made by the first user for the first sample image;
[0015] Perform structured processing on the inference result corresponding to the first sample image to obtain the second labeled data corresponding to the first sample image.
[0016] Optionally, the performing of second data integrity processing on the second sample image and the third labeled data to obtain the second target labeled data corresponding to the second sample image includes:
[0017] Perform inference processing on the second sample image through the algorithm model to obtain the inference result of the second sample image;
[0018] Perform structured processing on the inference result corresponding to the second sample image to obtain the fourth labeled data corresponding to the second sample image;
[0019] Perform a second data integrity process on the third labeled data and the fourth labeled data to obtain the second target labeled data corresponding to the second sample image.
[0020] Optionally, constructing a training data set based on the first sample image, the first target labeled data corresponding to the first sample image, the second sample image, and the second target labeled data corresponding to the second sample image includes:
[0021] Perform preprocessing on the first sample image and the second sample image based on image quality to obtain the preprocessed first sample image and the preprocessed second sample image;
[0022] Construct a training data set based on the preprocessed first sample image, the first target labeled data corresponding to the preprocessed first sample image, the preprocessed second sample image, and the second target labeled data corresponding to the preprocessed second sample image.
[0023] Optionally, constructing a training data set based on the first sample image, the first target labeled data corresponding to the first sample image, the second sample image, and the second target labeled data corresponding to the second sample image includes:
[0024] Associate the first sample image with the first target labeled data corresponding to the first sample image, and associate the second sample image with the second target labeled data corresponding to the second sample image to obtain corresponding training data;
[0025] Divide the training data into multiple training data subsets according to scenarios;
[0026] Each time, extract one piece of the training data from the training data subsets and number it according to the extraction order to obtain the numbered training data;
[0027] Construct a training data set based on the numbered training data.
[0028] Optionally, the training data has a positive sample attribute or a negative sample attribute. After constructing the training data set based on the numbered training data, the method further includes:
[0029] Determine the ratio between the training data with the positive sample attribute and the training data with the negative sample attribute in the training data set;
[0030] If the ratio is less than a preset ratio threshold, the training data is replicated in the subset of the training data for the corresponding scenario to obtain replicated training data, and the number of the replicated training data is determined according to the number of the training data being replicated;
[0031] Based on the replicated training data, the training data set is supplemented to obtain a final training data set.
[0032] In a second aspect, an embodiment of the present invention further provides a data processing apparatus, where the data processing apparatus includes:
[0033] An acquisition module, configured to acquire first data to be processed and second data to be processed, where the first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image, the first annotation data and the second annotation data are determined according to an inference result obtained by performing an inference process on the first sample image by an algorithm model, the second data to be processed includes a second sample image, and third annotation data corresponding to the second sample image, and the first sample image and the second sample image are images from different data sources;
[0034] A first processing module, configured to perform first data integrity processing on the first annotation data and the second annotation data to obtain first target annotation data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third annotation data to obtain second target annotation data corresponding to the second sample image;
[0035] A second processing module, configured to construct a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, where the training data set is used for online training of the target algorithm model.
[0036] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the data processing method provided by the embodiment of the present invention are implemented.
[0037] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the data processing method provided by the embodiment of the invention are implemented.
[0038] In an embodiment of the present invention, first data to be processed and second data to be processed are obtained. The first data to be processed includes a first sample image, as well as first annotation data and second annotation data corresponding to the first sample image. The first annotation data and the second annotation data are determined according to the inference results obtained after the algorithm model performs inference processing on the first sample image. The second data to be processed includes a second sample image, as well as third annotation data corresponding to the second sample image. The first sample image and the second sample image are images from different data sources. Perform first data integrity processing on the first annotation data and the second annotation data to obtain first target annotation data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third annotation data to obtain second target annotation data corresponding to the second sample image. Based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, a training data set is constructed. The training data set is used for online training of the target algorithm model. By using the inference results of the first sample image to determine the first annotation data and the second annotation data of the first sample image, and performing integrity processing on the first annotation data and the second annotation data to obtain the first target annotation data corresponding to the first sample image, the annotation quality of the first sample image is improved. At the same time, by using the second sample image from a different data source and the corresponding third annotation data to obtain the second target annotation data corresponding to the second sample image, the sample diversity of the training data set is improved by the second sample image and the second target annotation data, and the sample quality of the training data set is improved. By using this training data set to perform online training on the algorithm model, the accuracy of the online fine-tuning model can be improved, a higher model accuracy can be obtained through the processing of limited samples, and at the same time, the training efficiency is improved by means of data redundancy in the inference link. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present invention;
[0041] Figure 2 is a flowchart of another data processing method provided by an embodiment of the present invention;
[0042] Figure 3 is a schematic structural diagram of a data processing device provided by an embodiment of the present invention;
[0043] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0045] As Figure 1 shown, Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present invention. The data processing method includes the steps of:
[0046] 101. Obtain first data to be processed and second data to be processed.
[0047] In the embodiments of the present invention, the above data processing method can be applied to a data processing platform. The above data processing platform can be built based on a server or a distributed server. The above data processing platform includes a data interface, a database, an algorithm model interface, and a data processing program. The above data interface can be used to obtain the first data to be processed and the second data to be processed. The database can be used to store data such as the first data to be processed, the second data to be processed, and a training data set. The above algorithm model interface can be used to call algorithm models in different scenarios for inference and receive the inference results of algorithm models in different scenarios. The above data processing program is used to execute the above data processing method. Different scenarios correspond to different event detection tasks, and different event detection tasks use different algorithm models for inference.
[0048] When the online training start condition of the algorithm model is triggered, the first data to be processed and the second data to be processed can be read from the database. The online training start condition can be started by a user, can be started regularly, or can be started when the accuracy rate of the algorithm model drops to a preset accuracy rate.
[0049] The first data to be processed includes a first sample image, as well as first annotation data and second annotation data corresponding to the first sample image. The first annotation data and the second annotation data are determined according to the inference results obtained after the algorithm model performs inference processing on the first sample image. The second data to be processed includes a second sample image, as well as third annotation data corresponding to the second sample image. The first sample image and the second sample image are images from different data sources.
[0050] The above first data to be processed may include multiple first sample images, and each first sample image corresponds to a first annotation data and a second annotation data. The above second data to be processed may include multiple second sample images, and each second sample image corresponds to a second annotation data. The above first data to be processed may be picture data stored in a database and corresponding event data after being inferred by an algorithm model for a corresponding scenario. The above first data to be processed may be referred to as return data, and the above second data to be processed may be referred to as external data.
[0051] It can be understood that the above first sample image is an image collected by an image device in a specific scenario. After being inferred and processed by an algorithm model for the corresponding scenario, the inference result of the first sample image can be obtained. The inference result can be sent to the user, and through the user's confirmation, modification, and supplementation of the inference result of the first sample image, the first annotation data can be formed. The inference result can be structurally processed, and the structurally processed inference result is used as the second annotation data.
[0052] 102. Perform first data integrity processing on the first annotation data and the second annotation data to obtain the first target annotation data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third annotation data to obtain the second target annotation data corresponding to the second sample image.
[0053] In the embodiment of the present invention, after obtaining the first data to be processed, the first data to be processed can be parsed to obtain the first sample image and the first annotation data and the second annotation data corresponding to the first sample image. Since the annotation sources of the above first annotation data and the second annotation data are different, the first annotation data and the second annotation data can be subjected to first integrity processing, and then the complete annotation data of the first sample image can be obtained. The complete annotation data of the first sample image is the first target annotation data corresponding to the first sample image.
[0054] After obtaining the second data to be processed, the second data to be processed can be parsed to obtain the second sample image and the third annotation data corresponding to the second sample image. Since the first sample image and the second sample image are images from different data sources, the third annotation data may be incomplete. Therefore, the second integrity processing can be performed on the third annotation data corresponding to the second sample image to obtain the complete annotation data of the second sample image. The complete annotation data of the second sample image is the second target annotation data corresponding to the second sample image.
[0055] The above integrity processing may be to complete the types of the labeled data, so that the target labeled data has integrity and improves the labeling quality of the labeled data. The above first target labeled data and the above second target labeled data may be labeled data with the same data structure.
[0056] 103. Based on the first sample image, the first target labeled data corresponding to the first sample image, the second sample image, and the second target labeled data corresponding to the second sample image, a training data set is constructed, and the training data set is used for online training of the target algorithm model.
[0057] In an embodiment of the present invention, after obtaining the first sample image and the first target labeled data corresponding to the first sample image, the first sample image and the first target labeled data may be associated to form a training data. After obtaining the second sample image and the second target labeled data corresponding to the second sample image, the second sample image and the first target labeled data may be associated to form a training data.
[0058] After associating all the first sample images with the corresponding first target labeled data and associating all the second sample images with the corresponding second target labeled data, all the training data is added to a set to obtain a training data set.
[0059] After obtaining the training data set, the training data set may be used for online training of the algorithm model, thereby making the performance of the algorithm model better.
[0060] In a possible embodiment, the main data structure of the above training data set is algorithm category (object detection, classification, segmentation); label index and label category mapping list; mapping list of picture file name and annotation information file; picture file directory (including multiple JPEG files, each JPEG file corresponds to a first sample image); annotation information file directory (including multiple annotation information files, each annotation information file corresponds to a target labeled data). Among them, the main data structure of the annotation information file is the width and height information of the picture (the width and height information of the first sample image); annotation list (target labeled data, the positive sample is the target box coordinate information, target label category; the negative sample target box is empty, the type is background). The main data structure of the picture file is the acquisition device, acquisition time, and inference result (including the algorithm name, and the target label, confidence level, target box of each target).
[0061] In an embodiment of the present invention, first data to be processed and second data to be processed are obtained. The first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image. The first annotation data and the second annotation data are determined according to inference results obtained after the first sample image is processed by an algorithm model. The second data to be processed includes a second sample image, and third annotation data corresponding to the second sample image. The first sample image and the second sample image are images from different data sources. The first annotation data and the second annotation data are subjected to first data integrity processing to obtain first target annotation data corresponding to the first sample image, and the second sample image and the third annotation data are subjected to second data integrity processing to obtain second target annotation data corresponding to the second sample image. Based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, a training data set is constructed. The training data set is used for online training of a target algorithm model. By using the inference results of the first sample image to determine the first annotation data and the second annotation data of the first sample image, and performing integrity processing on the first annotation data and the second annotation data to obtain the first target annotation data corresponding to the first sample image, the annotation quality of the first sample image is improved. At the same time, by using the second sample image from a different data source and the corresponding third annotation data to obtain the second target annotation data corresponding to the second sample image, and using the second sample image and the second target annotation data to improve the sample diversity of the training data set, the sample quality of the training data set is improved. By using this training data set to perform online training on the algorithm model, the accuracy of the online fine-tuning model can be improved, a higher model accuracy can be obtained through the processing of a limited number of samples, and at the same time, the training efficiency is improved by means of data redundancy in the inference process.
[0062] It can be understood that in the specific implementation of the present application, relevant data such as video stream data, image data, and annotation data are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data, as well as the training, deployment, and invocation of the algorithm model, all need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0063] Optionally, in the step of obtaining the first data to be processed, an image to be inferred can be obtained; the image to be inferred is processed by an algorithm model to obtain an inference result of the image to be inferred; the first sample image is determined from the image to be inferred; based on the inference result corresponding to the first sample image, the first annotation data and the second annotation data corresponding to the first sample image are determined.
[0064] In an embodiment of the present invention, a video stream can be obtained from an image device, the video stream is parsed to obtain an image to be inferred, and the image to be inferred is subjected to an inference process through an algorithm model corresponding to the scenario to obtain an inference result of the image to be inferred. The above first annotation data and second annotation data are annotation data from different annotation sources, and the above inference result can be processed through different annotation sources to obtain the first annotation data and second annotation data of different annotation sources.
[0065] Please refer to Figure 2 , Figure 2 which is a framework diagram of a data processing platform provided by an embodiment of the present invention. The above data processing platform includes an image device interface, an external data interface, an external database, an inference application, an event database, a picture database, a training application, and a fine-tuning model repository. Among them, the data processing platform is communicatively connected to a plurality of image devices through the image device interface to obtain a video stream transmitted by the image device, and the video stream is decoded and inferred by the inference application by calling a corresponding algorithm model to obtain a corresponding inference result. The above inference can be to perform event detection on the video stream to obtain a corresponding event detection result, and the event detection result can include picture data and event data, where the picture data can be stored in the picture database, and the event data can be stored in the event database. There is a corresponding mapping relationship between the event data and the picture data. One event data can correspond to a group of picture data, and a group of picture data includes at least one picture data. The event data can include a background image, associated targets and their information.
[0066] In a possible embodiment, the event detection can include object detection, object classification, object segmentation, etc., and corresponding algorithm categories can be used to perform the above object detection, object classification, object segmentation, etc. The above event data can include the algorithm name, the target label corresponding to each target, the confidence level, and the target box, etc.
[0067] After the inference process on the video stream, object tracking, pre-selection, and optimization can be performed to obtain picture data and event data related to the event. The picture data is stored in the picture database, and the event data is stored in the event database. When online training of the algorithm model is required, the first sample image is selected from the picture database, and the event data corresponding to the first sample image is matched from the event data stored in the event database, and the first annotation data and second annotation data corresponding to the first sample image are determined according to the event data.
[0068] In a possible embodiment, please continue to refer to Figure 2 , and in the process of inference of the event data and acquisition of the picture data, it can be implemented through the following steps:
[0069] The single-frame data decoded from the video is sent to a certain algorithm model for inference;
[0070] The original result of model inference is appended as metadata information of the decoded YUV video frame;
[0071] The inference application performs target tracking, optimization, preprocessing, etc.;
[0072] When the inference application identifies a target as an event, it is necessary to report the event data;
[0073] The inference application performs JPEG encoding on the background image associated with the target event data to be reported, encodes the background image into JPEG data through a hardware encoder, extracts the metadata information list of the current YUV video frame, performs json serialization, and places it as a custom APP14 MARKER at the head of the JPEG file;
[0074] Taking events as the granularity, persist the event database and persistently store the associated background images (image database).
[0075] Optionally, in the step of determining the first annotation data and the second annotation data corresponding to the first sample image based on the inference result corresponding to the first sample image, the first sample image and the inference result corresponding to the first sample image can be sent to the first user to obtain the first annotation data made by the first user for the first sample image; perform structured processing on the inference result corresponding to the first sample image to obtain the second annotation data corresponding to the first sample image.
[0076] In the embodiments of the present invention, the above inference result may be event data. When online training of the algorithm model is required, picture data can be selected from the picture database as the first sample image, and event data corresponding to the first sample image can be selected from the event database. The first sample image and the corresponding event data are sent to the first user. After receiving the first sample image and the corresponding event data, the first user can modify the event data to form the corresponding first annotation data. The above event data may include algorithm name, target label, confidence level, and target box corresponding to each target, etc. The above first annotation data includes positive and negative sample identifiers, target box coordinate information, and target label categories. Among them, the positive sample target box information is a coordinate box, the positive sample target category is the target label category, the negative sample target box is empty, and the target category is the background.
[0077] The above structured processing refers to representing the inference result in a structured manner, such as algorithm name, target label, confidence level, and target box corresponding to each target, etc. The above second annotation data includes positive and negative sample identifiers, target box coordinate information, and target label categories. Among them, the positive sample target box information is a coordinate box, the positive sample target category is the target label category, the negative sample target box is empty, and the target category is the background.
[0078] It should be noted that the above first annotation data is the data annotated by the first user based on the first sample image and the corresponding event data. The above second annotation data does not require the first user to participate in the processing. When performing integrity processing on the first annotation data and the second annotation data, the first annotation data is the main one. The data in the second annotation data that conflicts with the first annotation data is discarded, and the data in the second annotation data that does not exist in the first annotation data is supplemented to the first annotation data to obtain the first target annotation data.
[0079] Since in the inference stage, the complete inference information is retained when storing the event data (i.e., the return data) and the image data, to prevent the event data from filtering the non-subject detection results, this stage can be read and parsed through the custom data defined in the file header.
[0080] Specifically, the first annotation data T1 corresponding to the first sample image can be read; the annotation data T2 obtained by parsing the first sample image through the custom data format of json can be read; T2 and T1 are merged to obtain the final annotation data T (i.e., the first target annotation data) corresponding to the current first sample image. All the first sample images are processed item by item.
[0081] Optionally, in the step of performing second data integrity processing on the second sample image and the third annotation data to obtain the second target annotation data corresponding to the second sample image, the second sample image can be inferred through an algorithm model to obtain the inference result of the second sample image; the inference result corresponding to the second sample image is structured to obtain the fourth annotation data corresponding to the second sample image; the second data integrity processing is performed on the third annotation data and the fourth annotation data to obtain the second target annotation data corresponding to the second sample image.
[0082] In the embodiment of the present invention, the above second annotation data is annotated by the second user, and the corresponding second sample image is not the image collected by the above image device. It can be an image provided by the second user or selected from an open-source image annotation database. The role of the second sample image is to enrich the data source of the training data set and prevent the algorithm model from being trained only through a single data source, resulting in a decrease in the generalization performance.
[0083] After obtaining the second sample image, the second sample image can be inferred through the algorithm model corresponding to the scene to obtain the inference result of the second sample image. The inference result corresponding to the second sample image is structured to obtain the fourth annotation data corresponding to the second sample image.
[0084] The above-mentioned structured processing refers to the structured representation of the inference results, such as the algorithm name, the target label corresponding to each target, the confidence level, and the target box, etc. The above-mentioned fourth annotation data includes positive and negative sample identifiers, target box coordinate information, and target label categories. Among them, the positive sample target box information is a coordinate box, the positive sample target category is the target label category, the negative sample target box is empty, and the target category is the background.
[0085] Perform the second data integrity processing on the third annotation data and the fourth annotation data to obtain the second target annotation data corresponding to the second sample image. When performing the integrity processing on the third annotation data and the fourth annotation data, the fourth annotation data can be used as the main data, discard the data in the third annotation data that conflicts with the fourth annotation data, and supplement the data in the third annotation data that does not exist in the fourth annotation data to the fourth annotation data, so as to obtain the second target annotation data. Of course, in a possible embodiment, if the second annotation data has been verified by the first user, the third annotation data can be used as the main data, discard the data in the fourth annotation data that conflicts with the third annotation data, and supplement the data in the fourth annotation data that does not exist in the third annotation data to the third annotation data, so as to obtain the second target annotation data.
[0086] In a possible embodiment, it is possible to distinguish whether it is return flow data (the first sample image) or external data (the second sample image) by reading the picture file header, load the specified algorithm model through the inference algorithm interface; infer the inference result T1 of the current second sample image through the inference algorithm interface; merge the annotation data of T1 and the third annotation data T0 annotated by the second user, and the credibility of the third annotation data annotated by the second user is higher than the result of the model's secondary inference, to obtain the final annotation result T of the current sample; process all the second sample images item by item.
[0087] Optionally, in the step of constructing a training dataset based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, the first sample image and the second sample image can be preprocessed to obtain the preprocessed first sample image and the preprocessed second sample image; based on the preprocessed first sample image, the first target annotation data corresponding to the preprocessed first sample image, the preprocessed second sample image, and the second target annotation data corresponding to the preprocessed second sample image, construct a training dataset.
[0088] In an embodiment of the present invention, after obtaining the first sample image and the first target annotation data corresponding to the first sample image, the first sample image and the first target annotation data can be associated (associated through a mapping list of picture file names and annotation information files) to form a piece of training data. After obtaining the second sample image and the second target annotation data corresponding to the second sample image, the second sample image and the second target annotation data can be associated (associated through the mapping list of picture file names and annotation information files) to form a piece of training data. Multiple pieces of training data are used to construct a training dataset.
[0089] The main data structures of the above training dataset are algorithm categories (object detection, classification, segmentation); a mapping list of label indexes and label categories; a mapping list of picture file names and annotation information files; a picture file directory (containing multiple JPEG files, each JPEG file corresponding to a first sample image); an annotation information file directory (containing multiple annotation information files, each annotation information file corresponding to a target annotation data). Among them, the main data structure of the annotation information file is the width and height information of the picture (the width and height information of the first sample image); an annotation list (target annotation data, the positive sample is the target box coordinate information, the target label category; the negative sample target box is empty, and the type is background). The main data structure of the picture file is the acquisition device, acquisition time, and inference result (including the algorithm name, and the target label, confidence level, and target box of each target).
[0090] Furthermore, during the construction of the training dataset, the image quality of the training data can be evaluated to obtain the image quality of the sample images in the training data. The sample images in the training data can be the first sample image or the second sample image. Specifically, the picture size of each sample image can be determined, and the sample images with a picture size smaller than the set size threshold can be filtered out. The target area of the target object in each sample image can also be determined, and the sample images with a target area smaller than the set area threshold can be filtered out. In this way, the training data corresponding to the sample images with higher image quality can be retained, thereby improving the data quality of the training dataset.
[0091] In a possible embodiment, duplicate removal processing can also be performed on the sample images in the training data. By calculating the md5 value of the sample images, if the md5 values are the same, one copy is retained and the labeled data is merged. During the merging process, if there is conflicting labeled data, it is determined whether there is labeled data belonging to the first labeled data among the conflicting labeled data. If there is labeled data belonging to the first labeled data, the labeled data belonging to the first labeled data is retained as the standard. If there is no labeled data belonging to the first labeled data, the first user can be submitted for confirmation. By extracting the metadata of the picture file corresponding to the sample image, the picture acquisition device and the acquisition time are obtained. For samples collected by the same device and within a relatively small time window, it is determined whether the samples are duplicates by calculating the target overlap degree. When the overlap degree is greater than the set overlap degree threshold, one picture with more labeled targets is retained as the final sample. All sample images are processed item by item to obtain the deduplicated training data, and a training data set is constructed through the deduplicated training data. In a possible embodiment, for positive samples, the training data that is not retained in the duplicate sample images can be randomly modified corresponding labeled data, and the training data corresponding to the modified labeled data is added to the training data set as negative samples.
[0092] In a possible embodiment, the above preprocessing can also be implemented when constructing the picture database and the event database, and the above duplicate removal processing can also be implemented when constructing the picture database and the event database. The above preprocessing and duplicate removal processing can be used as the preprocessing before constructing the training data set to improve the accuracy of the training data set and reduce the repeatability of the training data.
[0093] In a possible embodiment, please continue to refer to Figure 2 , the construction and use of the training data set are as follows:
[0094] The first user previews the event data through the visualization page, and the background image associated with the event is synchronously displayed. The first user selectively confirms whether the event is correct or wrong.
[0095] The first user manually or the system automatically starts the training iteration task, and uses the data confirmed in the event database as the training data for the first round of screening to construct a basic training sample set (including pictures, positive and negative sample labels in the pictures, and target label types), which is submitted to the training application in the form of a compressed file through the API interface.
[0096] The training application decompresses the training sample set submitted by the user.
[0097] Perform sample duplicate removal processing on the training data in the training sample set.
[0098] Restore the target integrity of the feedback data (the first data to be processed).
[0099] Perform secondary reasoning on external data (second data to be processed) to supplement the target integrity.
[0100] Perform low-quality filtering on the training data.
[0101] Perform positive and negative sample statistics on the training data and calculate the positive and negative sample ratio.
[0102] Compare whether the positive and negative ratio of the current training sample set meets the preset minimum sample ratio. If it is less than the minimum sample ratio, start filling the training data from the system-predefined training data subset to reach the minimum control ratio.
[0103] Organize the above-cleaned training sample set into a training data set and send it to the training algorithm for training.
[0104] Optionally, in the step of constructing a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, the first sample image can be associated with the first target annotation data corresponding to it, and the second sample image can be associated with the second target annotation data corresponding to it to obtain the corresponding training data; divide the training data into multiple training data subsets according to the scenario; extract one training data from the training data subsets each time and number it in the extraction order to obtain the numbered training data; based on the numbered training data, construct a training data set.
[0105] In the embodiment of the present invention, after obtaining the first sample image and the first target annotation data corresponding to the first sample image, the first sample image can be associated with the first target annotation data (associated through the mapping list of the picture file name and the annotation information file) to form a piece of training data. After obtaining the second sample image and the second target annotation data corresponding to the second sample image, the second sample image can be associated with the second target annotation data (associated through the mapping list of the picture file name and the annotation information file) to form a piece of training data. Among multiple pieces of training data, divide them into multiple training data subsets according to the scenario (algorithm category). Each training data subset includes at least one piece of training data. Number all the training data, and the corresponding training data can be conveniently retrieved through the number.
[0106] Specifically, take one piece of training data from each training data subset and continuously number it in the order of the training data subsets until all the training data is numbered, so that the training data shows a certain scene discrete distribution in terms of sequential continuity. In this way, it is possible to avoid sampling too many training data of the same scene to construct the training data.
[0107] Optionally, when the training data has positive sample attributes or negative sample attributes, after the step of constructing a training data set based on the numbered training data, it is also possible to determine the ratio between the training data with positive sample attributes and the training data with negative sample attributes in the training data set; if the ratio is less than a preset ratio threshold, then copy the training data in the training data subset of the corresponding scenario to obtain replicated training data, and the number of the replicated training data is determined according to the number of the training data being replicated; supplement the training data set based on the replicated training data to obtain the final training data set.
[0108] In the embodiment of the present invention, the number of training data with positive sample attributes and the number of training data with negative sample attributes in the training data set can be determined. According to the number of training data with positive sample attributes and the number of training data with negative sample attributes, calculate the positive and negative sample ratio in the training data set (calculated by the number of sample images). If the total number of positive samples is M, the total number of negative samples is N, and the set minimum positive and negative sample ratio is Q, the current actual sample ratio is M:N. Both M and N are integers greater than 1. Generally, M is greater than N, and the specific values can be set according to experience. For example, M:N can be 5:1, 4:1, etc.
[0109] When M:N < Q, the number of positive samples to be supplemented, T = N*Q - M, then start replicating and supplementing positive samples from the training data set of the corresponding scenario, and the number of the replicated samples is 1 - T until M:N = Q is satisfied.
[0110] In the embodiment of the present invention, different strategies are adopted for the system return data and external data to ensure the integrity of annotation and reduce the intrusion and strong dependence on the application system data. Effectively utilize the existing inference calculation results, reduce the real-time inference calculation cost, and improve the training speed. The effective distribution of samples can be ensured through deduplication of sample data, low-quality cleaning and filtering, and dynamic balance of the positive and negative sample ratios. The fusion of manually annotated data and automatically annotated data through secondary inference improves the annotation quality and integrity. Higher model accuracy is obtained through the processing of limited samples, and at the same time, data redundancy is used in the inference link to improve the training efficiency.
[0111] As Figure 3 shown, the embodiment of the present invention provides a data processing device, and the data processing device includes:
[0112] An acquisition module 301, configured to acquire first data to be processed and second data to be processed, where the first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image, and the first annotation data and the second annotation data are determined according to inference results obtained by performing inference processing on the first sample image by an algorithm model, the second data to be processed includes a second sample image, and third annotation data corresponding to the second sample image, and the first sample image and the second sample image are images from different data sources;
[0113] A first processing module 302, configured to perform first data integrity processing on the first annotation data and the second annotation data to obtain first target annotation data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third annotation data to obtain second target annotation data corresponding to the second sample image;
[0114] A second processing module 303, configured to construct a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, where the training data set is used to perform online training on the target algorithm model.
[0115] Optionally, the acquisition module 301 is further configured to acquire an image to be inferred; perform inference processing on the image to be inferred by the algorithm model to obtain an inference result of the image to be inferred; determine a first sample image in the image to be inferred; and determine first annotation data and second annotation data corresponding to the first sample image based on the inference result corresponding to the first sample image.
[0116] Optionally, the acquisition module 301 is further configured to send the first sample image and the inference result corresponding to the first sample image to a first user, and acquire first annotation data made by the first user for the first sample image; and perform structured processing on the inference result corresponding to the first sample image to obtain second annotation data corresponding to the first sample image.
[0117] Optionally, the second processing module 303 is further configured to perform inference processing on the second sample image by the algorithm model to obtain an inference result of the second sample image; perform structured processing on the inference result corresponding to the second sample image to obtain fourth annotation data corresponding to the second sample image; and perform second data integrity processing on the third annotation data and the fourth annotation data to obtain second target annotation data corresponding to the second sample image.
[0118] Optionally, constructing a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image includes:
[0119] Preprocessing the first sample image and the second sample image based on image quality to obtain the preprocessed first sample image and the preprocessed second sample image;
[0120] Constructing a training data set based on the preprocessed first sample image, the first target annotation data corresponding to the preprocessed first sample image, the preprocessed second sample image, and the second target annotation data corresponding to the preprocessed second sample image.
[0121] Optionally, the second processing module 303 is further configured to associate the first sample image with the first target annotation data corresponding to the first sample image, and associate the second sample image with the second target annotation data corresponding to the second sample image to obtain corresponding training data; divide the training data into multiple training data subsets according to scenarios; extract one piece of the training data from the training data subsets each time and number it according to the extraction order to obtain the numbered training data; construct a training data set based on the numbered training data.
[0122] Optionally, the apparatus further includes:
[0123] A third processing module, configured to determine the ratio between the training data with the positive sample attribute and the training data with the negative sample attribute in the training data set;
[0124] A fourth processing module, configured to, if the ratio is less than a preset ratio threshold, copy the training data in the training data subset corresponding to the scenario to obtain copied training data, and the number of the copied training data is determined according to the number of the training data being copied;
[0125] A fifth processing module, configured to supplement the training data set based on the copied training data to obtain a final training data set.
[0126] As Figure 4 shown, an embodiment of the present invention further provides an electronic device, including a processor, and the above processor can execute any one of the above data processing methods.
[0127] Specifically, it includes a processor 401 and a memory 402, and a computer program for executing a data processing method stored on the memory 402 and capable of running on the processor 401, where:
[0128] The processor 401 runs a calculator program of a data processing method stored in the memory 402 and executes the following steps:
[0129] Obtain first data to be processed and second data to be processed. The first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image. The first annotation data and the second annotation data are determined according to inference results obtained by performing inference processing on the first sample image by an algorithm model. The second data to be processed includes a second sample image and third annotation data corresponding to the second sample image. The first sample image and the second sample image are images from different data sources;
[0130] Perform first data integrity processing on the first annotation data and the second annotation data to obtain first target annotation data corresponding to the first sample image, and perform second data integrity processing on the second sample image and the third annotation data to obtain second target annotation data corresponding to the second sample image;
[0131] Construct a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image. The training data set is used for online training of the target algorithm model.
[0132] Optionally, the obtaining of the first data to be processed executed by the processor 401 includes:
[0133] Obtain an image to be inferred;
[0134] Perform inference processing on the image to be inferred by the algorithm model to obtain an inference result of the image to be inferred;
[0135] Determine a first sample image in the image to be inferred;
[0136] Based on the inference result corresponding to the first sample image, determine first annotation data and second annotation data corresponding to the first sample image.
[0137] Optionally, the determining of the first annotation data and the second annotation data corresponding to the first sample image executed by the processor 401 includes:
[0138] Send the first sample image and the inference result corresponding to the first sample image to a first user, and obtain first annotation data made by the first user for the first sample image;
[0139] Perform structured processing on the inference result corresponding to the first sample image to obtain the second annotation data corresponding to the first sample image.
[0140] Optionally, the second data integrity processing performed by the processor 401 on the second sample image and the third annotation data to obtain the second target annotation data corresponding to the second sample image includes:
[0141] Perform inference processing on the second sample image through the algorithm model to obtain the inference result of the second sample image;
[0142] Perform structured processing on the inference result corresponding to the second sample image to obtain the fourth annotation data corresponding to the second sample image;
[0143] Perform second data integrity processing on the third annotation data and the fourth annotation data to obtain the second target annotation data corresponding to the second sample image.
[0144] Optionally, the construction of the training data set by the processor 401 based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image includes:
[0145] Perform preprocessing on the first sample image and the second sample image based on image quality to obtain the preprocessed first sample image and the preprocessed second sample image;
[0146] Construct a training data set based on the preprocessed first sample image, the first target annotation data corresponding to the preprocessed first sample image, the preprocessed second sample image, and the second target annotation data corresponding to the preprocessed second sample image.
[0147] Optionally, the construction of the training data set by the processor 401 based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image includes:
[0148] Associate the first sample image with the first target annotation data corresponding to the first sample image, and associate the second sample image with the second target annotation data corresponding to the second sample image to obtain corresponding training data;
[0149] Divide the training data into multiple training data subsets according to the scenario;
[0150] Each of the training data is extracted from the training data subset one by one and numbered according to the extraction order to obtain the numbered training data;
[0151] Based on the numbered training data, a training data set is constructed.
[0152] Optionally, the training data has a positive sample attribute or a negative sample attribute. After constructing the training data set based on the numbered training data, the method executed by the processor 401 further includes:
[0153] Determine the ratio between the training data with the positive sample attribute and the training data with the negative sample attribute in the training data set;
[0154] If the ratio is less than a preset ratio threshold, the training data is copied in the training data subset of the corresponding scenario to obtain copied training data, and the number of the copied training data is determined according to the number of the training data being copied;
[0155] Based on the copied training data, the training data set is supplemented to obtain a final training data set.
[0156] An embodiment of the present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements each process of the data processing method provided by the embodiment of the present invention and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0157] Those of ordinary skill in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0158] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A data processing method, characterized in that: The method comprises the following steps: Acquire first data to be processed and second data to be processed, wherein the first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image, the first annotation data and the second annotation data are determined according to an inference result obtained after an algorithm model performs inference processing on the first sample image, and the second data to be processed includes a second sample image, and third annotation data corresponding to the second sample image, and the first sample image and the second sample image are images from different data sources; Performing a first data integrity processing on the first labeled data and the second labeled data to obtain first target labeled data corresponding to the first sample image, and performing a second data integrity processing on the second sample image and the third labeled data to obtain second target labeled data corresponding to the second sample image; Based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, a training data set is constructed, and the training data set is used to perform online training on the target algorithm model.
2. The data processing method according to claim 1, characterized in that: The obtaining of first data to be processed comprises: Obtain the image to be inferred; Performing reasoning processing on the image to be inferred by using the algorithm model to obtain a reasoning result of the image to be inferred; Determining a first sample image in the image to be inferred; Based on the inference result corresponding to the first sample image, first annotation data and second annotation data corresponding to the first sample image are determined.
3. The data processing method according to claim 2, characterized in that: The determining, based on the inference result corresponding to the first sample image, first annotation data and second annotation data corresponding to the first sample image includes: Sending the first sample image and the inference result corresponding to the first sample image to a first user, and obtaining first annotation data made by the first user for the first sample image; The inference result corresponding to the first sample image is structured to obtain second annotation data corresponding to the first sample image.
4. The data processing method according to claim 1, characterized in that: The performing second data integrity processing on the second sample image and the third annotated data to obtain second target annotated data corresponding to the second sample image includes: Performing inference processing on the second sample image by using the algorithm model to obtain an inference result of the second sample image; Performing structural processing on the inference result corresponding to the second sample image to obtain fourth annotated data corresponding to the second sample image; A second data integrity processing is performed on the third labeled data and the fourth labeled data to obtain second target labeled data corresponding to the second sample image.
5. The data processing method according to claim 1, characterized in that: The constructing a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image includes: Preprocessing the first sample image and the second sample image based on image quality to obtain the preprocessed first sample image and the preprocessed second sample image; A training data set is constructed based on the preprocessed first sample image, the first target annotation data corresponding to the preprocessed first sample image, the preprocessed second sample image, and the second target annotation data corresponding to the preprocessed second sample image.
6. The data processing method according to any one of claims 1 to 5, characterized in that: The constructing a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image includes: Associating the first sample image with the first target annotation data corresponding to the first sample image, and associating the second sample image with the second target annotation data corresponding to the second sample image, to obtain corresponding training data; Dividing the training data into multiple training data subsets according to scenarios; Extracting one training data at a time from the training data subset and numbering them in the extraction order to obtain the training data with numbers; Based on the numbered training data, a training data set is constructed.
7. The data processing method according to claim 6, characterized in that: The training data has a positive sample attribute or a negative sample attribute. After constructing a training data set based on the numbered training data, the method further includes: Determining, in the training data set, a ratio between the training data having the positive sample attribute and the training data having the negative sample attribute; If the ratio is less than a preset ratio threshold, the training data is copied in the training data subset corresponding to the scene to obtain copied training data, and the number of the copied training data is determined according to the number of the copied training data; Based on the copied training data, the training data set is supplemented to obtain a final training data set.
8. A data processing device, characterized in that: The data processing device comprises: an acquisition module, configured to acquire first data to be processed and second data to be processed, wherein the first data to be processed includes a first sample image, and first annotation data and second annotation data corresponding to the first sample image, the first annotation data and the second annotation data are determined according to an inference result obtained after an algorithm model performs inference processing on the first sample image, and the second data to be processed includes a second sample image, and third annotation data corresponding to the second sample image, and the first sample image and the second sample image are images from different data sources; a first processing module, configured to perform a first data integrity processing on the first labeled data and the second labeled data to obtain first target labeled data corresponding to the first sample image, and to perform a second data integrity processing on the second sample image and the third labeled data to obtain second target labeled data corresponding to the second sample image; The second processing module is used to construct a training data set based on the first sample image, the first target annotation data corresponding to the first sample image, the second sample image, and the second target annotation data corresponding to the second sample image, and the training data set is used to perform online training on the target algorithm model.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the data processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the data processing method according to any one of claims 1 to 7 are implemented.