Object Detection Model Processing Method, Device, Equipment and Storage Medium
By performing difficult data mining and server model training on edge devices, the problem of low detection accuracy of edge devices is solved, improving detection accuracy and saving network bandwidth.
Patent Information
- Application Number
- CN202111645457.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-29
AI Technical Summary
When edge devices detect objects, their accuracy is low, mainly due to the large difference between the training data and the actual detection data.
By obtaining difficult-to-example data on edge devices, performing difficult-to-example mining, uploading it to the server for model training, and updating the trained model to the edge device.
Improves the accuracy of edge devices in object detection tasks and reduces network bandwidth usage.
Smart Images

Figure CN114299030B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to an object detection model processing method, apparatus, device, and storage medium. Background Art
[0002] Currently, edge computing has become an emerging task processing method. Since the computing resources of edge computing are closer to the location where the computing tasks are generated, it can respond to computing tasks more quickly. At the same time, since edge computing can be processed by edge devices, it can also save network bandwidth. Object detection, as a common computing task, can also be implemented using the method of edge computing.
[0003] Currently, an object detection model can be deployed on an edge device so that the edge device has the ability to detect objects.
[0004] However, after the object detection model is deployed on the edge device, since the data obtained by the edge device when processing the object detection task is quite different from the training data during the training of the object detection model, the edge device will have a problem of low accuracy when using the method of edge computing to process the object detection task. Summary of the Invention
[0005] This application provides an object detection model processing method, apparatus, device, and storage medium to solve the problem of low accuracy when an edge device processes an object detection task.
[0006] In a first aspect, this application provides an object detection model processing method, including:
[0007] Obtain data to be detected, and input the data to be detected into an initial object detection model to obtain a model detection result, where the data to be detected is a video; perform hard example mining on the model detection result based on the video stream to obtain hard example data; upload the obtained hard example data to a server, so that when the number of received hard example data reaches a preset standard, the server creates a data set according to the hard example data and basic data, where the basic data is data pre-obtained by the server, and use the data set for model training to obtain a trained object detection model; receive the trained object detection model sent by the server.
[0008] In a possible implementation manner, performing hard example mining on the model detection result based on the video stream to obtain hard example data includes: extracting detection frames of each frame in the video of the model detection result based on the video stream; performing target tracking on each frame in the model detection result to add identifiers to the detection frames of each frame to obtain a tracking result; performing hard example mining on the tracking result to obtain hard example data.
[0009] In a possible implementation, hard example mining is performed on the tracking results to obtain hard example data, including: if it is detected that the first preset number of frames before and after any frame in the tracking results contain detection boxes corresponding to the same identifier, and any frame does not contain a detection box corresponding to the same identifier, then any frame is determined as a to-be-determined missed detection hard example frame; the union of all detection boxes in the second preset number of frames before and after the to-be-determined missed detection hard example frame is obtained to get a to-be-mapped detection box; the to-be-mapped detection box is mapped to the to-be-determined missed detection hard example frame to obtain a missed detection hard example candidate area for the to-be-determined missed detection hard example frame; the first image feature points of the image in the missed detection hard example candidate area are extracted; the image feature points of the images in the detection boxes of each frame in the third preset number of frames before and after the to-be-determined missed detection hard example frame are extracted; the first image feature points are matched with the image feature points of the images in the detection boxes of all frames in the third preset number of frames before and after the to-be-determined missed detection hard example frame to determine a plurality of first matching items; if the number of the plurality of first matching items is greater than the fourth preset number or greater than the first preset ratio, then a first circumscribed rectangle is determined according to the first image feature points, where the first circumscribed rectangle is the circumscribed rectangle of the first image feature points; the image content corresponding to the first circumscribed rectangle and the attribute of the first circumscribed rectangle are determined as a missed detection hard example, where the attribute of the circumscribed rectangle is the category of the object inside the circumscribed rectangle.
[0010] In a possible implementation, hard example mining is performed on the tracking results to obtain hard example data, including: if it is detected that any frame in the tracking results contains a detection box corresponding to any identifier, and the first preset number of frames before and after any frame do not contain a detection box corresponding to any identifier, then any frame is determined as a to-be-determined false detection hard example frame; the detection box in the to-be-determined false detection hard example frame is enlarged by a preset multiple to obtain a false detection hard example candidate area; the false detection hard example candidate area is mapped to the first preset number of frames before and after the to-be-determined false detection hard example frame to obtain a plurality of comparison areas; the second image feature points of the image in the false detection hard example candidate area are extracted; the image feature points of the images in the plurality of comparison areas of each frame in the first preset number of frames before and after the to-be-determined false detection hard example frame are extracted; the second image feature points are matched with the image feature points of each frame in the first preset number of frames of the to-be-determined false detection hard example frame to obtain a plurality of second matching items; if the number of the plurality of second matching items is greater than the seventh preset number or greater than the second preset ratio, then a second circumscribed rectangle is determined according to the second image feature points, where the second circumscribed rectangle is the circumscribed rectangle of the second image feature points; the image content corresponding to the second circumscribed rectangle and the attribute of the second circumscribed rectangle are determined as a false detection hard example, where the attribute of the circumscribed rectangle is the category of the object inside the circumscribed rectangle.
[0011] In a second aspect, the present application provides an object detection model processing method, including:
[0012] Receive the hard example data sent by the edge device; when the number of received hard example data reaches a preset standard, create a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server; use the data set for model training to obtain a trained object detection model; send the trained object detection model to the edge device.
[0013] In a possible implementation manner, when the number of received hard example data reaches a preset standard, creating a data set according to the hard example data and the basic data includes: when the data volume of the received hard example data reaches a preset value, or the ratio of the number of hard example data to the number of basic data reaches a third preset ratio, mix the hard example data and the basic data and divide them into a training set and a test set, where the training set and the test set form the data set.
[0014] In a possible implementation manner, using the data set for model training to obtain a trained object detection model includes: based on the training set, use transfer learning and / or early stopping strategy to train the initial object detection model to obtain a to-be-verified object detection model; based on the test set, calculate the performance index of the to-be-verified object detection model, and if the performance index reaches the preset standard, determine the to-be-verified object detection model as the trained model.
[0015] In a possible implementation manner, after using the data set for model training to obtain a trained object detection model, it further includes: using network pruning and / or quantization technology to simplify the trained object detection model to obtain a simplified object detection model; correspondingly, sending the trained object detection model to the edge device further includes: sending the simplified object detection model to the edge device.
[0016] In a third aspect, the present application provides an object detection model processing device, including:
[0017] A detection result obtaining module, configured to obtain the data to be detected, and input the data to be detected into the initial object detection model to obtain a model detection result, where the data to be detected is a video; a hard example data obtaining module, configured to perform hard example mining on the model detection result based on the video stream to obtain hard example data; a hard example data uploading module, configured to upload the obtained hard example data to the server, so that when the number of received hard example data reaches a preset standard, the server creates a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server, and uses the data set for model training to obtain a trained object detection model; a detection model receiving module, configured to receive the trained object detection model sent by the server.
[0018] In a fourth aspect, the present application provides an object detection model processing device, including:
[0019] The difficult example data receiving module is used to receive difficult example data sent by the edge device; the data set creation module is used to create a data set according to the difficult example data and the basic data when the number of received difficult example data reaches a preset standard, where the basic data is the data pre-obtained by the server; the detection model acquisition module is used to perform model training using the data set to obtain a trained object detection model; the detection model sending module is used to send the trained object detection model to the edge device.
[0020] In a fifth aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the object detection model processing method described in the first aspect.
[0021] In a sixth aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the object detection model processing method described in the second aspect.
[0022] In a seventh aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the object detection model processing method described in the first aspect.
[0023] In an eighth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the object detection model processing method described in the second aspect.
[0024] In a ninth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the object detection model processing method described in the first aspect.
[0025] In a tenth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the object detection model processing method described in the first aspect.
[0026] The object detection model processing method, device, equipment, and storage medium provided by the present application complete the mining of difficult example data by monitoring data on the edge device and obtaining difficult example data in the detection results, and realize the update of the model by sending the difficult example data to the server and receiving the trained object detection model sent by the server, achieving the effect of improving the accuracy of the edge device when detecting objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0028] Figure 1 Schematic diagram of the application scenario of the object detection model processing method provided by the embodiment of the present application;
[0029] Figure 2 Flow diagram of the object detection model processing method provided by the embodiment of the present application Figure 1 ;
[0030] Figure 3 Flow diagram of the object detection model processing method provided by the embodiment of the present application Figure 2 ;
[0031] Figure 4 Schematic diagram of the interaction process of the object detection model processing method provided by the embodiment of the present application;
[0032] Figure 5 Schematic diagram of an object detection model processing device provided by the embodiment of the present application Figure 1 ;
[0033] Figure 6 Schematic diagram of an object detection model processing device provided by the embodiment of the present application Figure 2 ;
[0034] Figure 7 Schematic diagram of an object detection model processing device provided by the embodiment of the present application Figure 3 ;
[0035] Figure 8 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application.
[0036] Through the above accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0038] Object detection by current edge devices has become a daily routine, from detecting flower and plant varieties in daily life, detecting commodities during shopping to detecting road signs in intelligent driving, etc. The traditional detection method is to upload images or videos to the server. After the server detects objects in the pictures or videos, it then sends the detection results to the intelligent devices at the edge.
[0039] Using the traditional detection method will reduce the computing power at the edge, but it requires greater network support and occupies more bandwidth. Currently, with the improvement of the computing power of edge devices, the detection task can be directly carried out on the edge devices, that is, deploying an object detection model on the edge devices to enable the edge devices to have the object detection ability. However, the actual usage scenarios faced by edge devices are more complex. The objects detected by edge devices are very likely to be those that the models contained in the edge devices cannot detect, because similar situations were not encountered during the model training process, which will result in relatively low object detection accuracy.
[0040] To address the above technical problems, the present application proposes the following technical concept. Data is acquired through edge devices, and difficult example data in the data, that is, data that is more likely to be misdetected or missed detected, is directly found on the edge devices. Then the difficult example data is uploaded to the server, and the server uses the difficult example data to optimize the model and then sends it back to the edge devices to complete the update of the edge device model, improving the accuracy of the edge device when performing object detection tasks.
[0041] The present application is applied to the scenario of processing object detection models. In the technical solution of the present application, the acquisition, storage, and application of the involved data, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0042] Figure 1 It is a schematic diagram of the application scenario of the object detection model processing method provided by the embodiment of the present application. As Figure 1 shown, in this scenario, it includes: an edge device 101 and a server 102.
[0043] The edge device 101 can include mobile phones, tablets, personal digital assistants (PDAs), computers, intelligent vehicles, and notebooks, etc., which can perform input of image data and deployment of models.
[0044] The server 102 can be implemented by a single server or a cluster of multiple servers with more powerful processing capabilities and higher security. Where possible, computers or laptops with relatively strong computing capabilities can also be used for replacement.
[0045] The connection between the edge device 101 and the server 102 can be a wired connection or a wireless network connection. The network used for the wireless network connection can include various types of wired and wireless networks, such as but not limited to: the Internet, local area network, Wireless Fidelity (WIFI), Wireless Local Area Networks (WLAN), cellular communication networks (General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), 2G / 3G / 4G / 5G cellular networks), satellite communication networks, and so on.
[0046] In the specific implementation process, the edge device 101 is used to obtain the data to be detected, detect the data to be detected, perform hard example mining on the detected results to obtain hard example data, send the hard example data to the server 102, and finally receive the trained object detection model sent by the server 102.
[0047] The server 102 is used to receive the hard example data sent by the edge device 101, establish a data set based on the hard example data and the pre-obtained basic data, train the initial object detection model using the data set to obtain a trained object detection model, and finally send the trained object detection model to the edge device 101.
[0048] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the processing of the object detection model. In some other feasible embodiments of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or different component arrangements, which can be specifically determined according to the actual application scenario and will not be limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0049] The following uses specific embodiments to detail the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the drawings.
[0050] Figure 2 Flow schematic of the object detection model processing method provided by the embodiments of the present application Figure 1 The execution subject of the embodiments of the present application can be Figure 1 the edge device 101 among them. As Figure 2 shown, the method includes:
[0051] S201: Acquire data to be detected, and input the data to be detected into an initial object detection model to obtain a model detection result, wherein the data to be detected is a video.
[0052] In this step, the data to be detected can be obtained by using the camera provided by or connected to the edge device 101, or it can be obtained from the network, and this application does not limit this. The above video can be regarded as a splicing of multiple images. The model detection result includes the detection frame of the object in each frame image, the corresponding confidence of the detection frame, and the classification probability corresponding to the detection frame.
[0053] For example, if there are three different objects in the current video, namely, a dog, a cat, and a person, the detection results are the corresponding detection frames of the three objects. The dog is in the dog detection frame, the dog detection frame has the corresponding confidence, the probability of the detection frame being classified as a dog, and the probability of being classified as other objects. The cat is in the cat detection frame, the cat detection frame also has the corresponding confidence, and the cat detection frame has the corresponding probability of being classified as a cat and the probability of being classified as other objects.
[0054] S202: Perform hard example mining on the model detection results based on the video stream to obtain hard example data.
[0055] In this step, hard example mining may include mining missed hard examples and mining falsely detected hard examples.
[0056] Among them, missed detection hard examples are objects that are difficult to detect in the data to be detected, and false detection hard examples are objects that are easy to detect incorrectly in the data to be detected.
[0057] S203: Upload the obtained difficult example data to the server, so that when the amount of received difficult example data reaches a preset standard, the server creates a data set based on the difficult example data and basic data, wherein the basic data is data pre-acquired by the server, and uses the data set to perform model training to obtain a trained object detection model.
[0058] In this step, the obtained difficult example data is uploaded to the server, which may include the image content corresponding to the difficult example data, the label of whether it is a missed difficult example or a falsely detected difficult example, the identifier corresponding to the difficult example data, and the object corresponding to the difficult example.
[0059] For example, if the content of the difficult example data is a cat, and it is a missed difficult example, then the picture content corresponding to the cat, the label of the missed difficult example, the "Cat 1" logo, and the difficult example being a cat are uploaded to the server.
[0060] S204: Receive the trained object detection model sent by the server.
[0061] In this step, the received trained object detection model can be the trained object detection model sent by the network download server, or it can be to first receive the download address of the trained object detection model sent by the server, and then pull the trained object detection model according to the download address. The trained object detection model in this step can be regarded as the initial object detection model when the above step S201 is executed again.
[0062] From the description of the above embodiments, it can be seen that the embodiment of the present application completes the mining of hard example data by monitoring data on the edge device and obtaining hard example data in the detection results, sends the hard example data to the server and receives the trained object detection model sent by the server, realizes the update of the model, and improves the accuracy of the edge device when detecting objects.
[0063] In a possible implementation manner, in the above step S202, hard example mining is performed on the model detection results based on the video stream to obtain hard example data, including:
[0064] S2021: Extract the detection frames of each frame in the video of the model detection results based on the video stream.
[0065] In this step, extracting the detection frames of each frame in the video of the model detection results means finding all the detection frames in the model detection.
[0066] Specifically, obtaining the detection frame can be to obtain the diagonal vertices of the detection frame to obtain the position of each detection frame.
[0067] For example, in the first frame of the video, the cat is on the left, the dog is in the middle, and the person is on the right, and each of them has a detection frame. In the second frame of the video, the positions of the above cat, dog, and person may change, and the positions of the corresponding detection frames will also change or increase or decrease. This step will find all the detection frames in all frames.
[0068] S2022: Perform object tracking on each frame in the model detection results to add identifiers to the detection frames of each frame to obtain the tracking results.
[0069] In this step, the method of object tracking can be to perform object tracking on each frame using a multi-object tracking algorithm (MOT, multiple objective track). The content to be tracked can be the detection frames of each frame, or the objects in the detection frames of each frame.
[0070] Specifically, for example, in step S2021, if the cat moves from the left to the right from the second frame to the fifth frame, and the positions of the dog and the person do not change, then the tracking can determine that the cat in the second frame to the fifth frame is still the cat in the first frame, and add the identifier "cat1" to the detection frames corresponding to the cat in the first to fifth frames. The tracking results include the detection frames and the corresponding identifiers.
[0071] In a possible implementation, the tracking result further includes the confidence corresponding to the above detection box and the classification probability corresponding to the detection box.
[0072] S2023: Mine hard examples from the tracking results to obtain hard example data.
[0073] In this step, hard example mining can be obtained by analyzing the tracking results of any frame or any consecutive several frames before and after in the video.
[0074] As can be seen from the description of the above embodiments, by first tracking the model detection results, the embodiments of the present application can more accurately find hard example data, provide a more accurate method for edge devices to find hard example data, and thus further improve the accuracy of the finally obtained trained object detection model.
[0075] In a possible implementation, in the above step S2023, mining hard examples from the tracking results to obtain hard example data includes:
[0076] S20231A: If it is detected that the first preset number of frames before and after any frame in the tracking results all contain detection boxes corresponding to the same identifier, and any frame does not contain a detection box corresponding to the same identifier, then determine any frame as a to-be-determined missed detection hard example frame.
[0077] In this step, the object may not be detected or the tracking is lost due to being blocked or moving too fast, so that there is no corresponding detection box in a certain frame. When the object appears again, it means that the object has not disappeared from the actual scenario. This frame may contain an object that should have been detected but was not detected. Therefore, this frame is determined as a to-be-determined missed detection hard example frame.
[0078] Specifically, the first preset number is usually 1, but it can also be other values such as 2. It should not be too large, otherwise there will be a situation where the object moves back into the picture after a period of time, and although the actual detection is correct, the frame where the object is outside the picture is determined as a to-be-determined missed detection hard example frame during processing.
[0079] S20232A: Take the union of all detection boxes in the second preset number of frames before and after the to-be-determined missed detection hard example frame to obtain a to-be-mapped detection box.
[0080] In this step, taking the union can be to stack all detection boxes in the second preset number of frames before and after the to-be-determined missed detection hard example frame according to their positions in the image, or it can be to find the minimum bounding rectangle of the stacked figure after stacking.
[0081] Specifically, if the stacking method is to find the minimum bounding rectangle of the stacked figure after stacking, the position of the obtained minimum bounding rectangle can also be represented by the coordinates of the diagonal points.
[0082] S20233A: Map the detection box to be mapped to the frame with potential missed detection and difficult examples to obtain the candidate area of potential missed detection and difficult examples in the frame with potential missed detection and difficult examples.
[0083] In this step, mapping the detection box to be mapped to the frame with potential missed detection and difficult examples means setting the obtained detection box to be mapped at the same position in the frame with potential missed detection and difficult examples.
[0084] S20234A: Extract the first image feature points of the image within the candidate area of potential missed detection and difficult examples.
[0085] In this step, the method for extracting the first image feature points can be one or more of the SIFT (Scale - invariant feature transform) algorithm, Harris corner extraction algorithm, SURF (Speeded Up Robust Features) algorithm, FAST (Features from accelerated segment test) algorithm, BRIEF algorithm, and ORB (Oriented FAST and Rotated BRIEF) algorithm. Among them, the Harris algorithm can be used to detect corners, the SIFT algorithm can be used to detect blobs, the SURF algorithm can be used to detect corners, the FAST algorithm can be used to detect corners, the BRIEF algorithm can be used to detect blobs, and the ORB algorithm represents the FAST algorithm with direction and the BRIEF algorithm with rotational invariance.
[0086] S20235A: Extract the image feature points of the images within the detection boxes of each frame in the third preset number of frames before and after the frame with potential missed detection and difficult examples.
[0087] In this step, the method for extracting image feature points used in this step is similar to that in S20234 and will not be elaborated here. The third preset number can be 1, or larger values such as 2, 3, etc.
[0088] S20236A: Match the first image feature points with the image feature points of the images within the detection boxes of all frames in the third preset number of frames before and after the frame with potential missed detection and difficult examples to determine multiple first matching items.
[0089] In this step, the image feature points are matched. The matching can be performed using the brute force matching method or the fast library for approximate nearest neighbors (FLANN) method. If there are successfully matched feature points, the successfully matched feature points are taken as a first matching item, and multiple first matching items are finally obtained.
[0090] S20237A: If the number of multiple first matching items is greater than a fourth preset number or greater than a first preset ratio, a first circumscribed rectangle is determined according to the first image feature points, where the first circumscribed rectangle is the circumscribed rectangle of the first image feature points.
[0091] In this step, the number of first matching items is the total number of first matching items. The present application does not limit the specific values of the fourth number and the first preset ratio. The first circumscribed rectangle can be determined by the positions of all the first image feature points, and the first circumscribed rectangle needs to contain all the first image feature points within itself.
[0092] In a possible implementation, the first circumscribed rectangle is the minimum circumscribed rectangle of the first image feature points.
[0093] S20238A: The image content corresponding to the first circumscribed rectangle and the attributes of the first circumscribed rectangle are determined as missed detection hard examples, where the attributes of the circumscribed rectangle are the categories of the objects within the circumscribed rectangle.
[0094] In this step, the image content corresponding to the first circumscribed rectangle is the image within the first circumscribed rectangle. The categories of the objects within the circumscribed rectangle can include the confidence level of the object category and the classification probability.
[0095] As can be seen from the description of the above embodiments, the embodiments of the present application further provide a method for mining missed detection hard examples. By taking the union of the detection frames of a preset number of frames before and after the frame to be determined as a missed detection hard example and mapping, the missed detection hard example candidate region of any of the above frames can be obtained. Then, the matching items obtained by using image feature point matching are used to further narrow down the missed detection hard example candidate region to obtain the first circumscribed rectangle, so that the finally obtained missed detection hard examples are more accurate.
[0096] In a possible implementation, any frame in the above step S2023A can be any number of frames. That is, if the detection frames corresponding to the same identifier in the first preset number of frames before and after more than one frame all contain the detection frames corresponding to the same identifier, and these frames do not contain the detection frames corresponding to the same identifier, then these frames are determined as the frames to be determined as missed detection hard examples.
[0097] Accordingly, in step S20232A, the union of all detection boxes in the second preset number of frames before the first frame among these frames and the second preset number of frames after the last frame among these frames can be obtained to get the detection boxes to be mapped for these frames.
[0098] Accordingly, the above step S20233A can be: mapping the detection boxes to be mapped to these frames to be determined as missed detection hard example frames to obtain multiple missed detection hard example candidate regions for these frames to be determined as missed detection hard example frames.
[0099] Accordingly, the above step S20234A can be: extracting multiple first image feature points of the images in multiple missed detection hard example candidate regions.
[0100] Accordingly, the above step S20235A can be: extracting the image feature points of the images within the detection boxes of each frame in the third preset number of frames before and after the multiple frames to be determined as missed detection hard example frames.
[0101] Accordingly, the above step S20236A can be: matching each first image feature point among the multiple first image feature points with the image feature points of the images within the detection boxes of all frames in the third preset number of frames before and after the frames to be determined as missed detection hard example frames to determine multiple first matching items corresponding to each first image feature point.
[0102] Accordingly, the above step S20237A can be: if the number of multiple first matching items corresponding to the first image feature point is greater than the fourth preset number or greater than the first preset ratio, then determining a first circumscribed rectangle according to the first image feature point, where the first circumscribed rectangle is the circumscribed rectangle of the first image feature point.
[0103] Specifically, this step can obtain the circumscribed rectangles of the first image feature points of each frame among these frames to be determined as missed detection hard example frames.
[0104] Accordingly, the above step S20238A can be: determining the image content corresponding to all first circumscribed rectangles and the attributes of the first circumscribed rectangles as missed detection hard examples, where the attribute of the circumscribed rectangle is the category of the object within the circumscribed rectangle.
[0105] In a possible implementation manner, the above step S2023 for hard example mining of the tracking result to obtain hard example data includes:
[0106] S20231B: If it is detected that any frame in the tracking result contains a detection box corresponding to any identifier, and none of the fifth preset number of frames before and after any frame contains a detection box corresponding to any identifier, then determining any frame as a frame to be determined as a misdetection hard example frame.
[0107] In this step, if a detection box appears in a single frame of the model detection result but does not appear in the previous and subsequent frames, this independent detection box in this single frame may be a false detection, that is, an object that did not appear is regarded as an object that appeared.
[0108] S20232B: Enlarge the detection box in the frame to be determined as a difficult false detection case by a preset multiple to obtain a candidate area for the difficult false detection case.
[0109] In this step, the specific value of the preset multiple in this application embodiment is not limited, but generally does not exceed the range of this frame of image after the detection box is enlarged, that is, the candidate area for the difficult case is within the corresponding image range of this frame.
[0110] S20233B: Map the candidate area for the difficult false detection case to the first six preset number of frames before and after the frame to be determined as a difficult false detection case to obtain multiple comparison areas.
[0111] In this step, the mapping method is similar to that in step S20233A above and will not be elaborated here.
[0112] S20234B: Extract the second image feature points of the image within the candidate area for the difficult false detection case.
[0113] In this step, the extraction method of the second image feature points is similar to the extraction method of the first image feature points in step S20234A above and will not be elaborated here.
[0114] S20235B: Extract the image feature points of the images within the multiple comparison areas of each of the first six preset number of frames before and after the frame to be determined as a difficult false detection case.
[0115] In this step, the extraction method of the image feature points can be the extraction method in step S20234A above.
[0116] S20236B: Match the second image feature points with the image feature points of each of the first six preset number of frames of the frame to be determined as a difficult false detection case to obtain multiple second matching items.
[0117] In this step, the matching method of the image feature points can adopt the matching method in step S20236A above.
[0118] S20237B: If the number of multiple second matching items is greater than the seventh preset number or greater than the second preset ratio, determine a second circumscribed rectangle according to the second image feature points, where the second circumscribed rectangle is the circumscribed rectangle of the second image feature points.
[0119] In this step, the second circumscribed rectangle can be the minimum circumscribed rectangle of the second image feature points.
[0120] S20238B: Determine the image content corresponding to the second circumscribed rectangle and the attributes of the second circumscribed rectangle as misdetection hard examples, where the attributes of the circumscribed rectangle are the categories of the objects within the circumscribed rectangle.
[0121] This step is similar to the above step S20238A and will not be elaborated here.
[0122] As can be seen from the description of the above embodiments, the embodiments of the present application provide a method for mining misdetection hard examples. By mining misdetection hard examples, training on misdetection hard examples can be added in the subsequent training process to improve the accuracy of the object detection model trained subsequently.
[0123] Figure 3 Schematic flow of the object detection model processing method provided by the embodiments of the present application Figure 2 . The execution subject of the embodiments of the present application can be Figure 1 the server 102 in. As Figure 3 shown, the method includes:
[0124] S301: Receive hard example data sent by the edge device.
[0125] In this step, the description of the hard example data can be found in the relevant descriptions of the above steps S202 and S20238A and will not be elaborated here.
[0126] S302: When the number of received hard example data reaches a preset standard, create a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server.
[0127] In this step, the preset standard can be a preset quantity or a preset ratio. Creating a data set according to the hard example data and the basic data can be mixing the hard example data with the basic data and dividing them into a training set and a test set in proportion after mixing. The basic data can be the data pre-obtained for training an initial object detection model.
[0128] S303: Use the data set for model training to obtain a trained object detection model.
[0129] In this step, the present application does not limit the training method used for model training.
[0130] S304: Send the trained object detection model to the edge device.
[0131] In this step, the sending process can be sending the download address of the trained object detection model to the edge device so that the edge device can pull the trained object detection model according to the download address.
[0132] As can be seen from the description of the above embodiments, in this embodiment of the present application, the training, updating, and sending of the object detection model are completed by receiving hard example data and combining it with the basic data. Bandwidth is saved by only receiving hard example data, and the object detection model of the edge device is updated to make the edge device more accurate in detecting objects.
[0133] In a possible implementation manner, in the above step S302, when the number of received hard example data reaches a preset standard, a data set is created according to the hard example data and the basic data, including:
[0134] When the data volume of the received hard example data reaches a preset value, or the ratio of the number of hard example data to the number of basic data reaches a third preset ratio, the hard example data and the basic data are mixed and then divided into a training set and a test set, where the training set and the test set form the data set.
[0135] Among them, mixing the hard example data and the basic data can be mixing the hard example data and the basic data in proportion. At the same time, the proportion of undetected hard examples and misdetected hard examples in the hard example data can also be set during mixing. Dividing into a training set and a test set can also be done in a preset ratio. Usually, the training set accounts for a larger proportion.
[0136] Specifically, for example, mixing the hard example data and the basic data can be mixing in a ratio of 1:1, 1:2, 1:4, etc. The ratio of undetected hard examples and misdetected hard examples can be 1:1:2:3, 2:2, etc. The ratio of the training set to the test set can be 7:3, 8:2, 6:4, etc. The present application does not impose special restrictions on the various mixing ratios during mixing and the division ratios during division.
[0137] As can be seen from the description of the above embodiments, the embodiment of the present application provides a method for creating a data set. A reasonable data set can make the trained object detection model more accurate, so that the edge device is more accurate in object detection.
[0138] In a possible implementation manner, in the above step S303, the data set is used for model training to obtain a trained object detection model, including:
[0139] S3031: Based on the training set, the initial object detection model is trained using transfer learning and / or early stopping strategy to obtain an object detection model to be verified.
[0140] In this step, the early stopping strategy is to use the model with better performance obtained during the training process as the object detection model to be verified.
[0141] In a possible implementation manner, it can be to use the backbone replacement training method for model training to reduce the memory occupation of the obtained model.
[0142] S3032: Based on the above test set, calculate the performance metrics of the object detection model to be verified. If the performance metrics meet the preset criteria, determine the object detection model to be verified as a trained model.
[0143] In this step, based on the above test set, calculating the performance metrics of the object detection model to be verified means inputting the test set into the object detection model to be verified to obtain the corresponding results, and then using the corresponding results to calculate the performance metrics of the model.
[0144] Among them, the performance metrics include the mAP (Mean Average Precision) value, precision, recall, F1-score, and inference latency, etc. The preset criteria can be a preset value or the current performance value of the initial object detection model. If the performance metrics do not meet the preset value or the current performance value of the initial object detection model, the above steps S302 and 303 can be performed again.
[0145] Among them, the calculation method of precision is:
[0146] precision k = TP / (TP + FP)
[0147] Among them, precision k is the precision under a single category, TP is the number of correctly classified under a single category, and FP is the number of misclassifying other categories as this category under a single category.
[0148] The calculation method of recall is:
[0149] recall k = TP / (TP + FN)
[0150] Among them, recall k is the recall under a single category, TP is the number of correctly classified under a single category, and FN is the number of classifying this category as other categories under a single category.
[0151] The calculation method of F1-score under a single category is:
[0152]
[0153] Among them, fl k is the F1-score under a single category, precision k is the precision under a single category, and recall k is the recall under a single category.
[0154] The calculation method of the final F1-score is:
[0155]
[0156] Among them, score is the final F1-score, n is the number of categories, and f1 k is the F1-score for a single category.
[0157] The above inference latency is the time required for the calculation of each frame.
[0158] As can be seen from the description of the above embodiments, the embodiments of the present application can reduce the model training time by adopting transfer learning and early stopping strategies during the model training process, and by calculating the performance metrics of the object detection model to be verified, it can be determined whether the obtained object detection model to be verified can achieve better detection effects, realizing the quality inspection of the trained model.
[0159] In a possible implementation manner, after the above step S303 uses a data set to train a model to obtain a trained object detection model, it further includes:
[0160] S303A: Simplify the trained object detection model by using network pruning and / or quantization techniques to obtain a simplified object detection model.
[0161] In this step, network pruning can be to remove connections and links in some networks of the model or reduce the number of convolution kernels. Quantization is to reduce the floating-point precision of storage and calculation.
[0162] Correspondingly, after the above step S304 sends the trained object detection model to the edge device, it further includes:
[0163] S304A: Send the simplified object detection model to the edge device.
[0164] This step is similar to the above step S304 and will not be elaborated here.
[0165] As can be seen from the description of the above embodiments, by simplifying the model, the present application embodiments can reduce the memory occupied by the model and make the model adaptable to the smaller memory capacity of the edge device.
[0166] Figure 4 It is a schematic interaction flow diagram of the object detection model processing method provided by the embodiments of the present application. This embodiment describes the interaction process between the edge device and the server, as Figure 4 shown, the method includes:
[0167] S401: The edge device obtains the data to be detected, and inputs the data to be detected into the initial object detection model to obtain a model detection result, where the data to be detected is a video.
[0168] S402: The edge device performs hard example mining on the model detection results based on the video stream to obtain hard example data.
[0169] S403: The edge device uploads the obtained hard example data to the server.
[0170] S404: When the number of received hard example data reaches a preset standard, the server creates a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server.
[0171] S405: The server uses the data set for model training to obtain a trained object detection model.
[0172] S406: The server sends the trained object detection model to the edge device.
[0173] Figure 5 Schematic diagram of an object detection model processing device provided by an embodiment of the present application Figure 1 As Figure 5 shown, the object detection model processing device 500 includes: a detection result obtaining module 501, a hard example data obtaining module 502, a hard example data uploading module 503, and a detection model receiving module 504.
[0174] The detection result obtaining module 501 is configured to obtain data to be detected, and input the data to be detected into an initial object detection model to obtain a model detection result, where the data to be detected is a video.
[0175] The hard example data obtaining module 502 is configured to perform hard example mining on the model detection results based on the video stream to obtain hard example data.
[0176] The hard example data uploading module 503 is configured to upload the obtained hard example data to the server, so that when the number of received hard example data reaches a preset standard, the server creates a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server, and uses the data set for model training to obtain a trained object detection model.
[0177] The detection model receiving module 504 is configured to receive the trained object detection model sent by the server.
[0178] In a possible implementation manner, the above hard example data obtaining module 502 is specifically configured to: extract the detection frames of each frame in the video of the model detection result based on the video stream. Perform object tracking on each frame in the model detection result to add identifiers to the detection frames of each frame to obtain a tracking result. Perform hard example mining on the tracking result to obtain hard example data.
[0179] In a possible implementation, the above-mentioned hard example data acquisition module 502 is specifically configured to: if it is detected that the first preset number of frames before and after any frame in the tracking result all contain detection frames corresponding to the same identifier, and any frame does not contain a detection frame corresponding to the same identifier, then determine any frame as a to-be-determined missed detection hard example frame. Take the union of all detection frames in the second preset number of frames before and after the to-be-determined missed detection hard example frame to obtain a to-be-mapped detection frame. Map the to-be-mapped detection frame to the to-be-determined missed detection hard example frame to obtain a missed detection hard example candidate area of the to-be-determined missed detection hard example frame. Extract the first image feature points of the image in the missed detection hard example candidate area. Extract the image feature points of the images in the detection frames of each of the third preset number of frames before and after the to-be-determined missed detection hard example frame. Match the first image feature points with the image feature points of the images in the detection frames of all frames among the third preset number of frames before and after the to-be-determined missed detection hard example frame to determine a plurality of first matching items. If the number of the plurality of first matching items is greater than the fourth preset number or greater than the first preset ratio, then determine a first circumscribed rectangle according to the first image feature points, where the first circumscribed rectangle is the circumscribed rectangle of the first image feature points. Determine the image content corresponding to the first circumscribed rectangle and the attribute of the first circumscribed rectangle as a missed detection hard example, where the attribute of the circumscribed rectangle is the category of the object in the circumscribed rectangle.
[0180] In a possible implementation, the above-mentioned hard example data acquisition module 502 is specifically configured to: if it is detected that any frame in the tracking result contains a detection frame corresponding to any identifier, and the first preset number of frames before and after any frame do not contain a detection frame corresponding to any identifier, then determine any frame as a to-be-determined false detection hard example frame. Enlarge the detection frame in the to-be-determined false detection hard example frame by a preset multiple to obtain a false detection hard example candidate area. Map the false detection hard example candidate area to the first preset number of frames before and after the to-be-determined false detection hard example frame to obtain a plurality of comparison areas. Extract the second image feature points of the image in the false detection hard example candidate area. Extract the image feature points of the images in the plurality of comparison areas of each of the first preset number of frames before and after the to-be-determined false detection hard example frame. Match the second image feature points with the image feature points of each frame among the first preset number of frames of the to-be-determined false detection hard example frame to obtain a plurality of second matching items. If the number of the plurality of second matching items is greater than the seventh preset number or greater than the second preset ratio, then determine a second circumscribed rectangle according to the second image feature points, where the second circumscribed rectangle is the circumscribed rectangle of the second image feature points. Determine the image content corresponding to the second circumscribed rectangle and the attribute of the second circumscribed rectangle as a false detection hard example, where the attribute of the circumscribed rectangle is the category of the object in the circumscribed rectangle.
[0181] Figure 6 Schematic diagram of an object detection model processing device provided by an embodiment of the present application Figure 2 As Figure 6As shown, the object detection model processing device 600 includes: a hard example data receiving module 601, a data set creation module 602, a detection model acquisition module 603, and a detection model sending module 604.
[0182] The hard example data receiving module 601 is configured to receive hard example data sent by an edge device.
[0183] The data set creation module 602 is configured to create a data set according to the hard example data and basic data when the number of received hard example data reaches a preset standard, where the basic data is data pre-acquired by a server.
[0184] The detection model acquisition module 603 is configured to perform model training using the data set to obtain a trained object detection model.
[0185] The detection model sending module 604 is configured to send the trained object detection model to the edge device.
[0186] In a possible implementation manner, the data set creation module 602 is specifically configured to: when the data volume of the received hard example data reaches a preset value, or the ratio of the number of hard example data to the number of basic data reaches a third preset ratio, mix the hard example data and the basic data and divide them into a training set and a test set, where the training set and the test set form the data set.
[0187] In a possible implementation manner, the data set creation module 602 is specifically configured to: based on the training set, train an initial object detection model using transfer learning and / or early stopping strategy to obtain a to-be-verified object detection model. Based on the test set, calculate the performance index of the to-be-verified object detection model, and if the performance index reaches the preset standard, determine the to-be-verified object detection model as the trained model.
[0188] Figure 7 Schematic diagram of an object detection model processing device provided by an embodiment of this application Figure 3 As Figure 7 shown, the object detection model processing device 600 further includes:
[0189] The simplified model obtaining module 605 is configured to simplify the trained object detection model using network pruning and / or quantization technology to obtain a simplified object detection model;
[0190] The simplified model sending module 606 is configured to send the simplified object detection model to the edge device.
[0191] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of this application. For example, please refer to Figure 8As shown, the electronic device 800 may include a processor 801 and a memory 802 communicatively connected to the processor 801.
[0192] The memory 802 stores computer-executable instructions.
[0193] The processor 801 executes the computer-executable instructions stored in the memory 802 to implement the object detection model processing method provided in any of the above embodiments.
[0194] Optionally, the memory 802 may be either independent or integrated with the processor 801. When the memory 802 is a device independent of the processor 801, the electronic device may further include: a bus for connecting the memory 802 and the processor 801.
[0195] This application also provides a computer-readable storage medium storing computer-executable instructions. When the processor executes the computer-executable instructions, the technical solution of the object detection model processing method in any of the above embodiments is implemented. The implementation principle and beneficial effects are similar to those of the object detection model processing method. For details, refer to the implementation principle and beneficial effects of the object detection model processing method, which will not be elaborated here.
[0196] This application also provides a computer program product including a computer program. When the computer program is executed by the processor, the technical solution of the object detection model processing method in any of the above embodiments is implemented. The implementation principle and beneficial effects are similar to those of the object detection model processing method. For details, refer to the implementation principle and beneficial effects of the object detection model processing method, which will not be elaborated here.
[0197] In all embodiments of this application, the nth preset value only represents a numerical magnitude. This application does not impose specific limitations on the numerical magnitude. The "nth" is only used to distinguish numerical values and does not indicate a sequence.
[0198] In all embodiments of this application, the circumscribed rectangle and the minimum circumscribed rectangle can both be circumscribed rectangles with the same direction as the image, that is, the sides of the circumscribed rectangle can be parallel to the sides of the frame.
[0199] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.
[0200] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0201] In addition, each functional module in various embodiments of the present application can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.
[0202] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in various embodiments of the present application.
[0203] It should be understood that the above processor can be a Central Processing Unit (CPU for short), and can also be other general-purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.
[0204] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.
[0205] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.
[0206] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a Static Random Access Memory (SRAM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), an Erasable Programmable Read-Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0207] An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.
[0208] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical disks that can store program codes.
[0209] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0210] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for processing an object detection model, characterized in that, Applied to edge devices, including: Obtain the data to be detected, and input the data to be detected into the initial object detection model to obtain the model detection result, where the data to be detected is a video; Perform object tracking on each frame in the video of the model detection result based on the video stream to obtain the tracking result; If it is detected that a preset number of frames before and after any frame in the tracking result contain detection frames corresponding to the same identifier and the any frame does not contain the detection frame corresponding to the same identifier, or if it is detected that any frame in the tracking result contains a detection frame corresponding to any identifier and a preset number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, then determine the any frame as a candidate hard example frame to be determined; Obtain the hard example candidate region in the candidate hard example frame to be determined; Extract the image feature points of the image in the hard example candidate region; Extract the image feature points of the image in the target region; where, for the case where a preset number of frames before and after any frame contain detection frames corresponding to the same identifier and the any frame does not contain the detection frame corresponding to the same identifier, the target region is the detection frames of each frame in a preset number of frames before and after the candidate hard example frame to be determined; for the case where any frame contains a detection frame corresponding to any identifier and a preset number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, the target region is the comparison region of each frame in a preset number of frames before and after the candidate hard example frame to be determined, and the comparison region is obtained according to the mapping of the hard example candidate region in a preset number of frames before and after the candidate hard example frame to be determined; Match the image feature points of the image in the hard example candidate region and the image feature points of the image in the target region to obtain a plurality of matching items; If the number of the plurality of matching items is greater than a preset number or a preset ratio, then determine the minimum bounding rectangle according to the image feature points of the image in the hard example candidate region; Determine the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as hard example data; Upload the obtained hard example data to the server, so that when the number of the received hard example data reaches a preset standard, the server creates a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server, and uses the data set for model training to obtain a trained object detection model; Receive the trained object detection model sent by the server.
2. The method according to claim 1, characterized in that, The performing object tracking on each frame in the video of the model detection result based on the video stream to obtain the tracking result includes: Extract the detection frames of each frame in the video of the model detection result based on the video stream; Perform object tracking on each frame in the model detection result to add identifiers to the detection frames of each frame to obtain the tracking result.
3. The method according to claim 1 or 2, characterized in that The if it is detected that a preset number of frames before and after any frame in the tracking result contain detection frames corresponding to the same identifier and the any frame does not contain the detection frame corresponding to the same identifier, then determine the any frame as a candidate hard example frame to be determined includes: If it is detected that the first preset number of frames before and after any frame in the tracking result both contain detection frames corresponding to the same identifier, and the any frame does not contain the detection frame corresponding to the same identifier, then determine the any frame as a to-be-determined missed detection hard example frame; The obtaining of the hard example candidate region in the to-be-determined hard example frame includes: Take the union of all detection frames in the second preset number of frames before and after the to-be-determined missed detection hard example frame to obtain a to-be-mapped detection frame; Map the to-be-mapped detection frame to the to-be-determined missed detection hard example frame to obtain the missed detection hard example candidate region of the to-be-determined missed detection hard example frame; The extracting of the image feature points in the hard example candidate region includes: Extract the first image feature points of the image in the missed detection hard example candidate region; The extracting of the image feature points in the target region includes: Extract the image feature points of the image in the detection frame of each frame in the third preset number of frames before and after the to-be-determined missed detection hard example frame; The matching of the image feature points in the hard example candidate region and the image feature points in the target region to obtain a plurality of matching items includes: Match the first image feature points with the image feature points of the image in the detection frames of all frames in the third preset number of frames before and after the to-be-determined missed detection hard example frame to obtain a plurality of first matching items; If the number of the plurality of matching items is greater than the preset number or greater than the preset ratio, then determine the minimum bounding rectangle according to the image feature points in the hard example candidate region, including: If the number of the plurality of first matching items is greater than the fourth preset number or greater than the first preset ratio, then determine the first bounding rectangle according to the first image feature points, where the first bounding rectangle is the minimum bounding rectangle of the first image feature points; The determining of the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as hard example data includes: Determine the image content corresponding to the first bounding rectangle and the attributes of the first bounding rectangle as missed detection hard example data, where the attribute of the bounding rectangle is the category of the object inside the bounding rectangle.
4. The method according to claim 1 or 2, characterized in that The if it is detected that any frame in the tracking result contains a detection frame corresponding to any identifier, and the first preset number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, then determine the any frame as a to-be-determined hard example frame, includes: If it is detected that any frame in the tracking result contains a detection frame corresponding to any identifier, and the fifth preset number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, then determine the any frame as a to-be-determined false detection hard example frame; The obtaining of the hard example candidate region in the to-be-determined hard example frame includes: Enlarge the detection frame in the to-be-determined false detection hard example frame by a preset multiple to obtain a false detection hard example candidate region; The extracting of the image feature points in the hard example candidate region includes: Extract the second image feature points of the image in the false detection hard example candidate region; The extracting of the image feature points in the target region includes: Map the false detection hard example candidate region to the sixth preset number of frames before and after the to-be-determined false detection hard example frame to obtain a plurality of comparison regions; Extract the image feature points of the images in each of the first and sixth preset numbers of frames before and after the to-be-determined misdetection hard example frame; Match the image feature points of the image in the hard example candidate area with the image feature points of the image in the target area to obtain a plurality of matching items, including: Match the second image feature points with the image feature points of each frame in the first and sixth preset numbers of frames before and after the to-be-determined misdetection hard example frame to obtain a plurality of second matching items; If the number of the plurality of matching items is greater than a preset number or a preset ratio, determine the minimum bounding rectangle according to the image feature points of the image in the hard example candidate area, including: If the number of the plurality of second matching items is greater than a seventh preset number or a second preset ratio, determine a second bounding rectangle according to the second image feature points, where the second bounding rectangle is the minimum bounding rectangle of the second image feature points; Determine the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as hard example data, including: Determine the image content corresponding to the second bounding rectangle and the attributes of the second bounding rectangle as misdetection hard example data, where the attribute of the bounding rectangle is the category of the object inside the bounding rectangle.
5. A method for processing an object detection model, characterized in that Applied to a server, including: Receive the hard example data sent by the edge device; When the number of the received hard example data reaches a preset standard, create a data set according to the hard example data and the basic data, where the basic data is the data pre-obtained by the server; Use the data set for model training to obtain a trained object detection model; Send the trained object detection model to the edge device; Among them, the hard example data is obtained by the edge device in the following manner: The edge device obtains a model detection result according to the data to be detected, and performs target tracking on each frame in the video of the model detection result based on the video stream. After obtaining the tracking result; If the edge device detects that the frames with a preset number before and after any frame in the tracking result all contain detection frames corresponding to the same identifier and the any frame does not contain the detection frame corresponding to the same identifier, or if the edge device detects that any frame in the tracking result contains a detection frame corresponding to any identifier and the frames with a preset number before and after the any frame do not contain the detection frame corresponding to the any identifier, then determine the any frame as a to-be-determined hard example frame; The edge device obtains the hard example candidate regions in the to-be-determined hard example frames, extracts the image feature points of the images in the hard example candidate regions, and extracts the image feature points of the images in the target regions; wherein, for the cases where the front and back preset number of frames of any one frame contain detection frames corresponding to the same identifier and the any one frame does not contain the detection frame corresponding to the same identifier, the target region is the detection frames of each frame in the front and back preset number of frames of the to-be-determined hard example frame; for the cases where any one frame contains a detection frame corresponding to any identifier and the front and back preset number of frames of the any one frame do not contain the detection frame corresponding to the any identifier, the target region is the comparison regions of each frame in the front and back preset number of frames of the to-be-determined hard example frame, and the comparison region is obtained by mapping the hard example candidate region in the front and back preset number of frames of the to-be-determined hard example frame; The edge device matches the image feature points of the images in the hard example candidate region and the image feature points of the images in the target region to obtain a plurality of matching items; If the number of the plurality of matching items is greater than a preset number or greater than a preset ratio, the edge device determines the minimum bounding rectangle according to the image feature points of the images in the hard example candidate region, and determines the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as hard example data; Wherein, the data to be detected is a video.
6. The method according to claim 5, wherein When the quantity of the received hard example data reaches a preset standard, creating a data set according to the hard example data and the basic data, including: When the data volume of the received hard example data reaches a preset value, or the ratio of the quantity of the hard example data to the quantity of the basic data reaches a third preset ratio, mixing the hard example data and the basic data and then dividing them into a training set and a test set, wherein the training set and the test set form the data set.
7. The method according to claim 6, wherein Training an object detection model by using the data set to obtain a trained object detection model, including: Based on the training set, training an initial object detection model by using transfer learning and / or early stopping strategy to obtain a to-be-verified object detection model; Based on the test set, calculating the performance index of the to-be-verified object detection model, and if the performance index reaches a preset standard, determining the to-be-verified object detection model as a trained model.
8. The method according to any one of claims 5 to 7, characterized in that After training an object detection model by using the data set to obtain a trained object detection model, it further includes: Simplifying the trained object detection model by using network pruning and / or quantization technology to obtain a simplified object detection model; Correspondingly, when sending the trained object detection model to the edge device, it further includes: Sending the simplified object detection model to the edge device.
9. An object detection model processing device, characterized in that, Including: A detection result acquisition module, configured to obtain data to be detected and input the data to be detected into an initial object detection model to obtain a model detection result, wherein the data to be detected is a video; A difficult example data acquisition module, which is used to perform object tracking on each frame in the video of the model detection result based on the video stream to obtain a tracking result; if it is detected that the pre-set number of frames before and after any frame in the tracking result all contain detection frames corresponding to the same identifier, and the any frame does not contain the detection frame corresponding to the same identifier, or if it is detected that any frame in the tracking result contains a detection frame corresponding to any identifier, and the pre-set number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, then determine the any frame as a to-be-determined difficult example frame; Obtain a difficult example candidate area in the to-be-determined difficult example frame; extract image feature points of the image in the difficult example candidate area; extract image feature points of the image in the target area; wherein, for the case where the pre-set number of frames before and after any frame all contain detection frames corresponding to the same identifier, and the any frame does not contain the detection frame corresponding to the same identifier, the target area is the detection frames of each frame in the pre-set number of frames before and after the to-be-determined difficult example frame; for the case where any frame contains a detection frame corresponding to any identifier, and the pre-set number of frames before and after the any frame do not contain the detection frame corresponding to the any identifier, the target area is the comparison area of each frame in the pre-set number of frames before and after the to-be-determined difficult example frame, and the comparison area is obtained according to the mapping of the difficult example candidate area in the pre-set number of frames before and after the to-be-determined difficult example frame; match the image feature points of the image in the difficult example candidate area and the image feature points of the image in the target area to obtain a plurality of matching items; if the number of the plurality of matching items is greater than the pre-set number or greater than the pre-set ratio, then determine the minimum bounding rectangle according to the image feature points of the image in the difficult example candidate area; determine the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as difficult example data; A difficult example data uploading module, which is used to upload the obtained difficult example data to the server, so that when the number of the received difficult example data reaches the pre-set standard, the server creates a data set according to the difficult example data and the basic data, wherein the basic data is the data pre-obtained by the server, and uses the data set for model training to obtain a trained object detection model; A detection model receiving module, which is used to receive the trained object detection model sent by the server.
10. An object detection model processing device, characterized in that, Including: A difficult example data receiving module, which is used to receive difficult example data sent by an edge device; A data set creating module, which is used to create a data set according to the difficult example data and the basic data when the number of the received difficult example data reaches the pre-set standard, wherein the basic data is the data pre-obtained by the server; A detection model obtaining module, which is used to perform model training using the data set to obtain a trained object detection model; A detection model sending module, which is used to send the trained object detection model to the edge device; Among them, the difficult example data is obtained by the edge device in the following manner: The edge device obtains a model detection result based on the data to be detected, and performs object tracking on each frame in the video of the model detection result based on the video stream. After obtaining the tracking result; if the edge device detects that a pre-set number of frames before and after any frame in the tracking result contain detection frames corresponding to the same identifier, and the any frame does not contain a detection frame corresponding to the same identifier, or if the edge device detects that any frame in the tracking result contains a detection frame corresponding to any identifier, and a pre-set number of frames before and after the any frame do not contain a detection frame corresponding to the any identifier, then the edge device determines the any frame as a to-be-determined difficult example frame; the edge device acquires a difficult example candidate region in the to-be-determined difficult example frame, extracts image feature points of the image in the difficult example candidate region, and extracts image feature points of the image in the target region; wherein, for the case where a pre-set number of frames before and after the any frame contain detection frames corresponding to the same identifier, and the any frame does not contain a detection frame corresponding to the same identifier, the target region is the detection frames of each frame in a pre-set number of frames before and after the to-be-determined difficult example frame; for the case where any frame contains a detection frame corresponding to any identifier, and a pre-set number of frames before and after the any frame do not contain a detection frame corresponding to the any identifier, the target region is the comparison region of each frame in a pre-set number of frames before and after the to-be-determined difficult example frame, and the comparison region is obtained according to the mapping of the difficult example candidate region in a pre-set number of frames before and after the to-be-determined difficult example frame; the edge device matches the image feature points of the image in the difficult example candidate region with the image feature points of the image in the target region to obtain a plurality of matching items; if the number of the plurality of matching items is greater than a pre-set number or greater than a pre-set ratio, then the edge device determines a minimum bounding rectangle according to the image feature points of the image in the difficult example candidate region, and determines the image content corresponding to the minimum bounding rectangle and the attributes of the minimum bounding rectangle as difficult example data; wherein, the data to be detected is a video.
11. An electronic device, characterized in that, Including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the object detection model processing method according to any one of claims 1 to 4.
12. An electronic device, characterized in that, Including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the object detection model processing method according to any one of claims 5 to 8.
13. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the object detection model processing method according to any one of claims 1 to 4.
14. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the object detection model processing method according to any one of claims 5 to 8.
15. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the object detection model processing method described in any one of claims 1 to 4.
16. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the object detection model processing method described in any one of claims 5 to 8.
Citation Information
Patent Citations
Method for providing AI model, AI platform, computing device and storage medium
CN112529026A
Product appearance detection method based on cloud edge collaborative model optimization and implementation system thereof
CN112788110A