Data processing method and device, server and medium
By using two models with different accuracy to process image data and filtering data based on the differences in processing results, the problem of poor processing effect of the autonomous driving algorithm model in extreme cases is solved, and efficient data screening and model accuracy are achieved.
Patent Information
- Application Number
- CN202311597277.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
In high-level autonomous driving, it is difficult for algorithmic models to process data in extreme cases efficiently and accurately, resulting in poor processing of models in extreme cases.
By using two models with different accuracy to process image data separately, the required data is filtered out from the original data based on the differences between the processing results of the two models. The specific steps include acquiring image data, processing data using the first model and the second model respectively, and determining the required data based on the differences in the processing results.
It realizes efficiently screening out the data required for model training, improving the processing accuracy of the model in extreme cases.
Smart Images

Figure CN120047700A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicles, and particularly to a data processing method, apparatus, server, and medium. Background Art
[0002] In high-level autonomous driving, the performance of the algorithm model can affect the effect of autonomous driving. With the gradual maturity of perception technology and computing platforms, the algorithm model has a good processing effect for common cases that occur during driving. However, for corner cases that occur during driving, since corner cases are not common, it is difficult for the algorithm model to process corner cases efficiently and accurately.
[0003] In order to optimize the algorithm model, relevant data in extreme cases is required to train the algorithm model. The industry urgently needs a method that can mine relevant data in extreme cases from a large amount of data. Summary of the Invention
[0004] This application provides a data processing method, which can efficiently screen out the data required for model training from a large amount of image data. This application also provides an apparatus, a server, and a medium corresponding to the above method.
[0005] In a first aspect, this application provides a data processing method. The method includes:
[0006] Obtain first image data;
[0007] Process the first image data by using a first model and a second model respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model, where the model accuracy of the second model is higher than that of the first model;
[0008] Determine second image data from the first image data according to the difference between the first processing result and the second processing result.
[0009] In a second aspect, this application provides a data processing apparatus. The apparatus includes:
[0010] An obtaining module, configured to obtain first image data;
[0011] A processing module, configured to process the first image data by using a first model and a second model respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model;
[0012] A determining module, configured to determine second image data from the first image data according to the difference between the first processing result and the second processing result.
[0013] In a third aspect, the present application provides a server. The server includes a processor and a memory. Instructions are stored in the memory, and the processor executes the instructions to cause the server to execute the method described in the first aspect or any implementation manner of the first aspect of the present application.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a server, it causes the server to execute the method described in the first aspect or any implementation manner of the first aspect above.
[0015] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
[0016] Based on the above description, it can be seen that the technical solution of the present application has the following beneficial effects:
[0017] Specifically, the method first obtains first image data, processes the first image data using a first model and a second model respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model. Among them, the model accuracy of the second model is higher than that of the first model. Then, according to the difference between the first processing result and the second processing result, second image data is determined from the first image data.
[0018] In this method, the first image data is processed using two models with different accuracies. According to the difference between the processing results of the two models, second image data with better processing effect of the model with higher accuracy and poor processing effect of the model with lower accuracy can be obtained. In this way, efficient data screening is achieved, the image data required for model training is obtained, and thus the model accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present application will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0020] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0021] Figure 2 It is a schematic flowchart of another data processing method provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic structural diagram of a data processing platform provided by an embodiment of the present application;
[0023] Figure 4 A structural schematic diagram of a data processing device provided by an embodiment of the present application;
[0024] Figure 5 A structural schematic diagram of a server for implementing data processing provided by an embodiment of the present application. Detailed implementation manners
[0025] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0026] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0027] It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.
[0028] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".
[0029] To facilitate understanding of the technical solution of the present application, the specific application scenarios in the present application will be described below.
[0030] In high-level autonomous driving, algorithm models (such as visual perception models) are usually used to process image data. During the development process of visual perception models, a large amount of training data is relied on, and the quality and quantity of the training data have a crucial impact on the development efficiency and performance of visual perception models.
[0031] As the amount of data continues to increase, the processing effect of the visual perception model for common cases may reach a stable state. At this time, the key to improving the model performance lies in the data of corner cases. For example, during autonomous driving, the recognition of ordinary cars belongs to common cases, and the recognition of special-shaped vehicles (such as trucks and buses) belongs to corner cases.
[0032] Since the collection of training data is usually non-directional, the image data belonging to common cases still occupies the main part of the training data. However, the image data belonging to common cases is redundant for the visual perception model, which will not only increase the cost of data storage and processing, but also reduce the training speed of the model. When the image data belonging to corner cases is used as training data, it can improve the accuracy of the visual perception model, but this type of image data is often very scarce and requires a lot of effort and cost to be screened out from the database.
[0033] Based on this, an embodiment of the present application provides a data processing method. Specifically, the method first obtains first image data, and uses a first model and a second model to process the first image data respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model, where the model accuracy of the second model is higher than that of the first model. Then, according to the difference between the first processing result and the second processing result, second image data is determined from the first image data.
[0034] In this method, by using two models with different accuracies to process the first image data respectively, according to the difference between the processing results of the two models, it is possible to obtain second image data with better processing effect for the model with higher accuracy and poor processing effect for the model with lower accuracy. In this way, efficient data screening is realized, the image data required for model training is obtained, and thus the model accuracy can be improved.
[0035] Next, the data processing method provided by the embodiment of the present application will be described in detail with reference to the accompanying drawings.
[0036] See Figure 1 The flowchart of a data processing method shown, which specifically includes the following steps:
[0037] S101: Obtain first image data.
[0038] The first image data refers to the data used for model training. In other words, the first image data is unfiltered and includes image data belonging to common cases and corner cases.
[0039] In some possible implementation manners, the first image data may be multiple preprocessed images, and the first image data can be output through a data iterator.
[0040] S102: Process the first image data using the first model and the second model respectively to obtain the first processing result corresponding to the first model and the second processing result corresponding to the second model.
[0041] Among them, the first model and the second model can be models for performing autonomous driving tasks, such as neural network models, and the model accuracy of the second model is higher than that of the first model. In some embodiments, the first model can be a lightweight model actually deployed in a vehicle for engineering applications during the iterative development of autonomous driving, and the second model can be a high-precision model that does not consider model size and computing power.
[0042] In the embodiments of the present application, the first image data is input into the first model and the second model respectively, so that the first model and the second model perform inferences on the first image data respectively to obtain the first processing result and the second processing result.
[0043] Further, after the inferences of the first model and the second model are completed, the first processing result and the second processing result can be saved in the same format for subsequent determination of the differences between the processing results obtained by different models.
[0044] S103: Determine the second image data from the first image data according to the difference between the first processing result and the second processing result.
[0045] Among them, the difference between the first processing result and the second processing result refers to the image data where there are differences between the first processing result and the second processing result. Specifically, when implementing, the difference processing result can be determined according to the first processing result and the second processing result, and the image data corresponding to the difference processing result is determined as the second image data.
[0046] It can be understood that since the model accuracy of the second model is relatively high, by comparing the first processing result and the second processing result, the image data corresponding to the differences in the processing results of the two models can be screened out. This image data is the image data with poor processing effect of the first model, that is, the training data that can improve the model performance.
[0047] According to the different autonomous driving tasks for which the first model and the second model are used, the process of determining the difference between the first processing result and the second processing result can be different. The following will be combined with Figure 2 be described in detail.
[0048] In some embodiments, the first model and the second model are object detection models. The object detection model is used to detect at least one detection target in the image. In the scenario of autonomous driving, the detection targets can be vehicles, pedestrians, obstacles, etc.
[0049] Such asFigure 2 As shown, after the first model and the second model complete the processing of the first image data, the first processing result and the second processing result can be image-matched, that is, the first processing result and the second processing result under the same first image data are matched. Then, for each detection target, the detection result of the first processing result for this detection target is compared with the detection result of the second processing result for this detection target to determine the missed detection image data of the first model and the mis-detection image data of the first model, and the missed detection image data and the mis-detection image data are determined as the differential processing result.
[0050] It can be understood that since the model accuracy of the second model is relatively high, the reliability of the second processing result is higher than that of the first processing result. Therefore, in the embodiments of the present application, the second processing result can be regarded as the ground truth (GT), so as to obtain the missed detection image data of the first model and the mis-detection image data of the first model.
[0051] In some possible implementation manners, the detection result for the detection target can be presented in the form of a bounding box. In other words, the first processing result and the second processing result can be image data including multiple bounding boxes, and the area of the bounding box corresponds to the area of the detection target. Information such as the name and score of the detection target can be presented in the bounding box. Among them, the score is used to characterize the reliability of the detection result. The higher the score, the more reliable the detection result.
[0052] For the missed detection image data, in specific implementation, first, from the detection results of the first processing result for this detection target, the first detection results for this detection target are screened out, where the score of the first detection result is greater than the first score threshold. Then, from the detection results of the second processing result for this detection target, the second detection results for this detection target are screened out, where the score of the second detection result is greater than the second score threshold. Then, for each first image data, the matching degree between the first detection result and the second detection result is determined, and the image data corresponding to the detection result with the matching degree between the first detection result and the second detection result less than the first matching threshold is determined as the missed detection image data of the first model.
[0053] In some embodiments, since the model accuracy of the second model is relatively high, when screening the second processing result to obtain the second detection result, a relatively high score threshold (i.e., the second score threshold) can be set to obtain a reliable benchmark. And when screening the first processing result to obtain the first detection result, a relatively low score threshold (i.e., the first score threshold) can be set, so that more detection results are included in the first detection result, and the missed detection image data is avoided from being missed.
[0054] Since the detection result for the detection target can be presented in the form of a rectangular box, the matching degree between the first detection result and the second detection result can be characterized by the intersection over union (IoU) of the rectangular boxes. Among them, the intersection over union refers to the area of the intersection of two rectangular boxes divided by the area of the union. Thus, for the same detection target in the same image data, if the matching degree is greater than or equal to the first matching threshold, it indicates that the first model has detected the detection target; if the matching degree is less than the first matching threshold, it indicates that the first model has not detected the detection target, and this image data belongs to the undetected image data of the first model.
[0055] For the misdetected image data, in specific implementation, first, from the detection results for the detection target in the first processing result, filter out the third detection result for the detection target, where the confidence of the third detection result is greater than the third confidence threshold. Then, from the detection results for the detection target in the second processing result, filter out the fourth detection result for the detection target, where the confidence of the fourth detection result is greater than the fourth confidence threshold. Then, for each first image data, determine the matching degree between the third detection result and the fourth detection result, and determine the image data corresponding to the detection result with a matching degree less than the second matching threshold between the third detection result and the fourth detection result as the misdetected image data of the first model.
[0056] In some embodiments, since the model accuracy of the second model is relatively high, when filtering the second processing result to obtain the fourth detection result, a relatively low confidence threshold (i.e., the fourth confidence threshold) can be set, so that the fourth detection result as a benchmark contains more detection results. And when filtering the first processing result to obtain the third detection result, a relatively high confidence threshold (i.e., the third confidence threshold) can be set, so as to filter out the third detection results with high confidence, in order to determine the misdetected image data that is detected by the first model but does not exist in the detection results of the second model.
[0057] Similarly, the matching degree between the third detection result and the fourth detection result can also be characterized by the intersection over union of the rectangular boxes. Thus, for the same detection target in the same image data, if the matching degree is greater than or equal to the second matching threshold, it indicates that both the first model and the second model have detected the detection target; if the matching degree is less than the second matching threshold, it indicates that the second model has not detected the detection target, but the first model has detected the detection target, and this image data belongs to the misdetected image data of the first model.
[0058] In some other embodiments, the first model and the second model are semantic segmentation models. The semantic segmentation model is used to identify at least one recognition category in an image and can obtain a pixel-level recognition result. In the scenario of autonomous driving, the recognition category can be a lane line or the like.
[0059] As Figure 2 shown, after the first model and the second model complete the processing of the first image data, the first processing result and the second processing result can be image-matched, that is, the first processing result and the second processing result under the same first image data are matched. Then, for each first image data, the label map in the first processing result and the label map in the second processing result are compared to determine the region type of the label map, where the label map includes the recognition results of at least one recognition category, and the differential processing result is determined according to the region type of the label map.
[0060] In the embodiments of the present application, by comparing the label maps under different models, the regions where the recognition results for the same recognition type are different under different models can be obtained, so as to determine the differential processing result.
[0061] In specific implementation, the label map in the first processing result and the label map in the second processing result can be compared. When the recognition result in the label map in the first processing result is the same as the recognition result in the label map in the second processing result, the region with the same recognition result is determined as the overlapping region. When the recognition result in the label map in the first processing result is different from the recognition result in the label map in the second processing result, the region with the different recognition result is determined as the non-overlapping region. In this way, the first image data corresponding to the processing result including the non-overlapping region in the label map is determined as the differential processing result.
[0062] It can be understood that for the same image data, when the recognition result in the label map in the first processing result is the same as the recognition result in the label map in the second processing result, it indicates that the processing results of the first model and the second model are consistent, and this region belongs to the overlapping region. When the recognition result in the label map in the first processing result is different from the recognition result in the label map in the second processing result, for example, the second model obtains a recognition result for a certain recognition category in this region, while the first model does not have a recognition result in this region, or the recognition result of the second model in this region is A, while the recognition result of the first model in this region is B, it indicates that the processing results of the first model and the second model are inconsistent, and this region belongs to the non-overlapping region. The image data with the non-overlapping region is the data valuable for model optimization.
[0063] In addition, to perform model optimization more specifically, before comparing the label mask graphs of the processing results, the processing results can also be filtered first. During specific implementation, determine the matching degree between the label mask graph in the first processing result and the label mask graph in the second processing result, and filter out the first image data corresponding to the processing results with a matching degree less than the third matching threshold.
[0064] It can be understood that the matching degree between the label mask graph in the first processing result and the label mask graph in the second processing result can represent the difference degree of the processing results of the first model and the second model for a certain first image data. If the matching degree between the label mask graph in the first processing result and the label mask graph in the second processing result is less than the third matching threshold, it indicates that there is a large difference between the processing results of the first model and the second model. And the image data with a large difference between the processing results of the first model and the second model has high value for model optimization. Therefore, the corresponding first image data can be filtered out.
[0065] After determining the difference between the first processing result and the second processing result and determining the second image data, considering that the number of the second image data may still be large, and there may be similar or duplicate images. Since a large number of duplicate image data is redundant for model training, in order to avoid wasting resources, the second image data can be de-duplicated.
[0066] Specifically, the hash values corresponding to each second image data can be determined. According to the hash values corresponding to each second image data, determine the distance between any two second image data in the second image data, and filter out the second image data with a distance less than the distance threshold.
[0067] In the embodiments of the present application, similar second image data can be filtered through the difference Hashing (dHash) algorithm. For example, the second image data can be first reduced to the same size, such as a size of 9*8, and the reduced second image data is grayscale processed. In this way, the hash value of the second image data can be determined through the following steps: for each pixel point, compare the pixel value of the pixel point on its right. If the pixel value of the current pixel point is greater than the pixel value of the pixel point on the right, assign 1 to the current pixel point, otherwise assign 0. Thus, the values corresponding to 64 pixel points can be obtained, and the 0-1 sequence with a length of 64 is the hash value of the second image data. Then, according to the hash values corresponding to each second image data, calculate the Hamming distance between any two second image data. If the Hamming distance is less than the distance threshold, it indicates that the similarity of the two second image data is high, and one of the second image data can be filtered out to achieve de-duplication and reduce the resource consumption of similar image data in data storage and model training.
[0068] The above data processing method supports local operation. However, considering that a local single server only supports serial operation, when the data volume is large, the time consumption is huge, seriously affecting the development efficiency. In addition, when multiple local servers run in parallel, due to different time-consuming of different nodes and the mutual dependence of data in each link, developers need to regularly check the progress and manually transfer data, which will increase the workload of developers and reduce efficiency. Moreover, for different tasks and different models, the dependent configuration environments are different, and local operation requires frequent manual switching of the configuration environment, with a cumbersome process. The embodiments of the present application also support automatically implementing the data processing link by using a data processing platform.
[0069] See Figure 3 As shown in the structural schematic diagram of a data processing platform, developers can, in the data processing platform, configure the environments required for the modules on multiple servers according to task requirements, such as the environments required for the first model processing module, the second model processing module, the difference determination module, and the image deduplication module. Then, register the multiple servers in a continuous integration and continuous delivery (CICD) tool, connect to form a local area network, configure the dependency relationships between different modules, generate a configuration file, and submit the task. In this way, the data processing platform can automatically generate a computational graph, automatically complete the data processing process and save the results, improving the development efficiency and simplifying the process at the same time.
[0070] Based on the above description, the embodiments of the present application provide a data processing method. The method first obtains first image data, and uses a first model and a second model to process the first image data respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model. Among them, the model accuracy of the second model is higher than that of the first model. Then, according to the difference between the first processing result and the second processing result, second image data is determined from the first image data.
[0071] In this method, the first image data is processed by using two models with different accuracies. According to the difference between the processing results of the two models, it is possible to obtain second image data with better processing effect for the model with higher accuracy and poor processing effect for the model with lower accuracy. In this way, efficient data screening is realized, the image data required for model training is obtained, and the model accuracy can be improved.
[0072] Based on the above method provided by the embodiments of the present application, the embodiments of the present application also provide a data processing device corresponding to the above method. The units / modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the unit / module does not constitute a limitation to the unit / module itself in some cases.
[0073] See Figure 4 the structural schematic diagram of the data processing device shown. The device 400 includes:
[0074] An acquisition module 401, configured to acquire first image data;
[0075] A processing module 402, configured to process the first image data by using a first model and a second model respectively, to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model;
[0076] A determination module 403, configured to determine second image data from the first image data according to the difference between the first processing result and the second processing result.
[0077] In some possible implementation manners, the determination module 403 is specifically configured to:
[0078] Determine a difference processing result according to the first processing result and the second processing result;
[0079] Determine the image data corresponding to the difference processing result as the second image data.
[0080] In some possible implementation manners, the first model and the second model are object detection models, and the determination module 403 is specifically configured to:
[0081] For each detection target, compare the detection result for the detection target in the first processing result with the detection result for the detection target in the second processing result, to determine the undetected image data of the first model and the misdetected image data of the first model;
[0082] Determine the undetected image data and the misdetected image data as the difference processing result.
[0083] In some possible implementation manners, the determination module 403 is specifically configured to:
[0084] From the detection results for the detection target in the first processing result, screen out a first detection result for the detection target, where the confidence level of the first detection result is greater than a first confidence level threshold;
[0085] From the detection results of the detection target in the second processing result, filter out the second detection results of the detection target, where the confidence of the second detection results is greater than the second confidence threshold;
[0086] For each of the first image data, determine the degree of matching between the first detection result and the second detection result;
[0087] Determine the image data corresponding to the detection results with a matching degree less than the first matching threshold between the first detection result and the second detection result as the missed detection image data of the first model.
[0088] In some possible implementation manners, the determining module 403 is specifically configured to:
[0089] From the detection results of the detection target in the first processing result, filter out the third detection results of the detection target, where the confidence of the third detection results is greater than the third confidence threshold;
[0090] From the detection results of the detection target in the second processing result, filter out the fourth detection results of the detection target, where the confidence of the fourth detection results is greater than the fourth confidence threshold;
[0091] For each of the first image data, determine the degree of matching between the third detection result and the fourth detection result;
[0092] Determine the image data corresponding to the detection results with a matching degree less than the second matching threshold between the third detection result and the fourth detection result as the mis-detection image data of the first model.
[0093] In some possible implementation manners, the first model and the second model are semantic segmentation models, and the determining module 403 is specifically configured to:
[0094] For each of the first image data, compare the label mask map in the first processing result and the label mask map in the second processing result, and determine the region type of the label mask map, where the label mask map includes at least one recognition result of a recognition category;
[0095] Determine the difference processing result according to the region type of the label mask map.
[0096] In some possible implementation manners, the determining module 403 is specifically configured to:
[0097] Compare the label mask graphs in the first processing result and the label mask graphs in the second processing result. When the recognition results in the label mask graphs in the first processing result are the same as the recognition results in the label mask graphs in the second processing result, determine the regions with the same recognition results as overlapping regions;
[0098] When the recognition results in the label mask graphs in the first processing result are different from the recognition results in the label mask graphs in the second processing result, determine the regions with different recognition results as non-overlapping regions;
[0099] Determine the first image data corresponding to the processing result including the non-overlapping region in the label mask graph as the differential processing result.
[0100] In some possible implementation manners, the determining module 403 is further configured to:
[0101] Determine the matching degree between the label mask graph in the first processing result and the label mask graph in the second processing result;
[0102] Filter out the first image data corresponding to the processing result with a matching degree less than the third matching threshold.
[0103] In some possible implementation manners, the determining module 403 is further configured to:
[0104] Determine the hash values corresponding to each of the second image data;
[0105] Determine the distance between any two of the second image data according to the hash values corresponding to each of the second image data;
[0106] Filter the second image data with a distance less than the distance threshold.
[0107] According to the data processing apparatus 400 in the embodiments of the present application, it can correspondingly execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the data processing apparatus 400 are respectively for implementing Figure 1 or Figure 2 the corresponding processes of each method in the embodiments shown, and for the sake of brevity, they will not be described in detail here.
[0108] The functions described above in this article can be at least partially executed by one or more hardware logic components. Refer to Figure 5 the structural schematic diagram of the server 500 for implementing data processing shown. It should be noted that Figure 5 the server shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0109] AsFigure 5 As shown, the server 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the server 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0110] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the server 500 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 5 the server 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0111] This application also provides a computer-readable storage medium, also referred to as a machine-readable medium. In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0113] The above computer-readable medium carries one or more programs, which, when executed by the server, cause the server to: obtain first image data; process the first image data using a first model and a second model respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model; and determine second image data from the first image data according to the difference between the first processing result and the second processing result.
[0114] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device. When the computer program is executed by a processing device, it executes the above functions defined in the method of the embodiment of the present application.
[0115] Although the subject matter has been described in language specific to structural features and / or method logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0116] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0117] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but also covers other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. A data processing method, It is characterized in that The method comprises: Acquiring first image data; Processing the first image data using a first model and a second model respectively to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model, wherein the model accuracy of the second model is higher than the model accuracy of the first model; Second image data is determined from the first image data according to a difference between the first processing result and the second processing result.
2. The method according to claim 1, It is characterized in that The determining the second image data from the first image data according to the difference between the first processing result and the second processing result includes: Determine a difference processing result according to the first processing result and the second processing result; The image data corresponding to the difference processing result is determined as the second image data.
3. The method according to claim 2, It is characterized in that The first model and the second model are target detection models, and determining a difference processing result according to the first processing result and the second processing result includes: For each detection target, compare the detection result for the detection target in the first processing result with the detection result for the detection target in the second processing result to determine missed detection image data of the first model and false detection image data of the first model; The missed detection image data and the false detection image data are determined as difference processing results.
4. The method according to claim 3, It is characterized in that The comparing the detection result for the detection target in the first processing result with the detection result for the detection target in the second processing result to determine the missed detection image data of the first model includes: Filtering out a first detection result for the detection target from the detection results for the detection target in the first processing result, wherein the confidence of the first detection result is greater than a first confidence threshold; Filtering out a second detection result for the detection target from the detection results for the detection target in the second processing result, wherein the confidence of the second detection result is greater than a second confidence threshold; For each of the first image data, determining a matching degree between the first detection result and the second detection result; Image data corresponding to detection results for which the matching degree between the first detection result and the second detection result is less than a first matching threshold is determined as missed-detection image data of the first model.
5. The method according to claim 3, It is characterized in that The comparing the detection result for the detection target in the first processing result with the detection result for the detection target in the second processing result to determine the false detection image data of the first model includes: Filtering out a third detection result for the detection target from the detection results for the detection target in the first processing result, wherein the confidence of the third detection result is greater than a third confidence threshold; Filtering out a fourth detection result for the detection target from the detection results for the detection target in the second processing result, wherein the confidence of the fourth detection result is greater than a fourth confidence threshold; For each of the first image data, determining a matching degree between the third detection result and the fourth detection result; The image data corresponding to the detection result whose matching degree between the third detection result and the fourth detection result is less than the second matching threshold is determined as the false detection image data of the first model.
6. The method according to claim 2, It is characterized in that The first model and the second model are semantic segmentation models, and determining a difference processing result according to the first processing result and the second processing result includes: For each of the first image data, compare the label mask map in the first processing result with the label mask map in the second processing result to determine the region type of the label mask map, wherein the label mask map includes a recognition result of at least one recognition category; According to the region type of the label mask image, it is determined to be a difference processing result.
7. The method according to claim 6, It is characterized in that The comparing the label mask map in the first processing result and the label mask map in the second processing result to determine the region type of the label mask map includes: Comparing the label mask map in the first processing result with the label mask map in the second processing result, when the recognition result in the label mask map in the first processing result is the same as the recognition result in the label mask map in the second processing result, determining the area with the same recognition result as the overlapping area; When the recognition result in the label mask map in the first processing result is different from the recognition result in the label mask map in the second processing result, determining the area with different recognition results as a non-overlapping area; The determining, according to the region type of the label mask image, as a difference processing result includes: The first image data corresponding to the processing result including the non-overlapping area in the label mask image is determined as the difference processing result.
8. The method according to claim 6, It is characterized in that Before comparing the label mask map in the first processing result with the label mask map in the second processing result for each of the first image data to determine the region type of the label mask map, the method further includes: Determining a degree of match between the label mask map in the first processing result and the label mask map in the second processing result; The first image data corresponding to the processing result whose matching degree is less than the third matching threshold is screened out.
9. The method according to any one of claims 1 to 8, It is characterized in that The method further comprises: Determine a hash value corresponding to each of the second image data; determining, according to the hash values corresponding to the respective second image data, a distance between any two second image data in the second image data; The second image data whose distance is less than the distance threshold is filtered.
10. A data processing device, It is characterized in that The device comprises: An acquisition module, used for acquiring first image data; A processing module, configured to process the first image data using a first model and a second model respectively, to obtain a first processing result corresponding to the first model and a second processing result corresponding to the second model; A determination module is used to determine second image data from the first image data according to a difference between the first processing result and the second processing result.
11. A server, It is characterized in that The server includes a processor and a memory, wherein instructions are stored in the memory, and the processor executes the instructions so that the server executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, It is characterized in that The method comprises computer-readable instructions, which, when executed on a server, cause the server to execute the method according to any one of claims 1 to 9.