Image sample screening method and device and cleaning robot
By detecting whether the prediction results of the target image output match the image prediction model, determining difficult image samples, and selectively recycling these samples during the training process, the problem of high redundancy in the training sample set data in traditional techniques is solved and the training efficiency is improved.
Patent Information
- Application Number
- CN202311567749.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
When traditional technology is iteratively training image task models, there is a problem of high data redundancy in the collected training image samples, resulting in poor training results.
By acquiring multiple image prediction results output by the image prediction model for the target image, and detecting whether the image prediction results match according to the matching detection rules called by the model identification of each image prediction model. If it does not match, the target image is determined as a difficult example image sample.
When iteratively training the image task model, it is possible to selectively recycle difficult image samples with better training effects, reducing the data redundancy of the training sample set, thereby improving training efficiency.
Smart Images

Figure CN120032146A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to an image sample screening method, device and cleaning robot. Background Art
[0002] With the development of big data technology, various image task models have emerged. Image task models can be used to complete various image prediction tasks, such as target detection tasks, image classification tasks, or image segmentation tasks.
[0003] In traditional technology, a large number of training image samples are usually required to iteratively train image task models. During the iterative training process, the training image samples are usually fully recovered, and the fully recovered training data is added to the training sample set, and the image task model is continued to be iteratively trained, which can improve the utilization rate of the training image samples.
[0004] However, the recovered training image samples are usually not all image samples that contribute to the training effect of the image task model. If the recovered training image samples are directly added to the training sample set, some useless image samples will exist in the training sample set, resulting in high data redundancy in the training sample set. Summary of the invention
[0005] Based on this, it is necessary to provide an image sample screening method, device and cleaning robot that can reduce the data redundancy of the training sample set when iteratively training the image task model in response to the above technical problems.
[0006] In a first aspect, the present application provides an image sample screening method. The method comprises:
[0007] Obtain multiple image prediction results output by the image prediction model for the target image;
[0008] Detecting whether the image prediction results match each other according to the matching detection rules called by the model identifiers of the image prediction models;
[0009] If the image prediction results do not match, the target image is determined as a difficult image sample.
[0010] In one embodiment, the image prediction model includes a target detection model and an image segmentation model, the multiple image prediction results include a target detection result output by the target detection model and a first image segmentation result output by the image segmentation model; and detecting whether each of the image prediction results matches includes:
[0011] Extract the object image position information of the target object in the target detection result and the regional image position information of each segmented image region in the first image segmentation result; and detect whether the target detection result and the first image segmentation result match based on the object image position information and the regional image position information.
[0012] In one embodiment, detecting whether the target detection result matches the first image segmentation result according to the object image position information and the region image position information includes:
[0013] According to the object image position information and the regional image position information, locate the target image area where the target object is located in each of the segmented image areas; obtain the object detection type identifier corresponding to the target object in the target detection result, and the region type identifier corresponding to the target image area; according to the object detection type identifier and the region type identifier, detect whether the target detection result and the first image segmentation result match.
[0014] In one embodiment, detecting whether the target detection result matches the first image segmentation result according to the object image position information and the region image position information includes:
[0015] According to the object image position information and the regional image position information, detect whether the target object in the target image has an intersection with the regional segmentation boundaries between each of the segmented image regions; if there is an intersection, obtain the object image region area of the target object in different segmented image regions; according to each of the object image region areas, detect whether the target detection result and the first image segmentation result match.
[0016] In one embodiment, the image prediction model includes a target detection model and an image classification model, the multiple image prediction results include target detection results output by the target detection model and image classification results output by the image classification model; and detecting whether the image prediction results match includes:
[0017] Determine the object detection type identifier of the target object in the target detection result and the image classification identifier in the image classification result; if the object detection type identifier and the image classification identifier match, determine that the target detection result and the image classification result match; if the object detection type identifier and the image classification identifier do not match, determine that the target detection result and the image classification result do not match.
[0018] In one of the embodiments, the image prediction model includes an image segmentation model, and the image prediction result includes a first image segmentation result output by the image segmentation model; after the matching detection rule called according to the model identifier of each of the image prediction models is detected to see whether each of the image prediction results matches, the method further includes:
[0019] If the prediction results of each of the images match, then determine the neighborhood time frame images corresponding to the target image according to a preset time window, wherein the target image and the neighborhood time frame images are captured by a preset camera device, and the change amplitude of the camera parameters of the preset camera device within the preset time window is less than the preset change amplitude; detect the regional overlap between the segmented image area in the second image segmentation result of the target detection model for each neighborhood time frame image and the segmented image area in the first image segmentation result; and determine whether to use the target image as a difficult image sample based on the overlap of each area.
[0020] In one embodiment, determining whether to use the target image as a difficult image sample according to the overlap of each region includes:
[0021] If the variation range of the overlap between the overlaps of the regions is not less than the preset variation range, the target image is determined as a difficult image sample; if the variation range of the overlap between the overlaps of the regions is less than the preset variation range, the target image is not determined as a difficult image sample.
[0022] In one embodiment, determining whether to use the target image as a difficult image sample according to the overlap of each region includes:
[0023] If the variation range of the overlap between the overlaps of the regions conforms to the preset variation range distribution, the target image is determined as a difficult image sample; if the variation range of the overlap does not conform to the preset variation range distribution, the target image is not determined as a difficult image sample.
[0024] In a second aspect, the present application also provides an image sample screening device. The device comprises:
[0025] An acquisition module, used to acquire multiple image prediction results output by the image prediction model for the target image;
[0026] A matching detection rule determination module, used to determine the matching detection rule commonly corresponding to each of the image prediction results according to the model identifier of each of the image prediction models;
[0027] A matching detection module, used to detect whether the image prediction results match each other according to the matching detection rules called by the model identifier of each image prediction model;
[0028] The difficult image sample screening module is used to determine the target image as a difficult image sample if the image prediction results do not match.
[0029] In a third aspect, the present application further provides a cleaning robot. The cleaning robot comprises a body, a cleaning module arranged on the body, a sensor, a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0030] Acquire multiple image prediction results output by the image prediction model for the target image; detect whether each of the image prediction results matches according to the matching detection rules called by the model identifier of each of the image prediction models; if the image prediction results do not match, determine the target image as a difficult image sample.
[0031] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0032] Acquire multiple image prediction results output by the image prediction model for the target image; detect whether each of the image prediction results matches according to the matching detection rules called by the model identifier of each of the image prediction models; if the image prediction results do not match, determine the target image as a difficult image sample.
[0033] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0034] Acquire multiple image prediction results output by the image prediction model for the target image; detect whether each of the image prediction results matches according to the matching detection rules called by the model identifier of each of the image prediction models; if the image prediction results do not match, determine the target image as a difficult image sample.
[0035] The above-mentioned image sample screening method, device and cleaning robot obtain multiple image prediction results output by the image prediction model for the target image; according to the matching detection rules called by the model identifier of each of the image prediction models, detect whether each of the image prediction results matches. In this way, it is possible to implement targeted matching detection of the image prediction results output by different image prediction models for the target image according to the matching detection rules corresponding to different types of image prediction models. If the image prediction results do not match, it means that there are erroneous prediction results in the image prediction results. At this time, the target image can be determined as a difficult image sample. In this way, when the training image samples are recycled to the training sample set, the difficult image samples with better training effects can be selectively recycled to the training sample set, rather than recycling the full amount of training image samples to the training sample set. Therefore, the data redundancy of the training sample set during the iterative training of the image task model can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic diagram of a flow chart of an image sample screening method in one embodiment;
[0037] Figure 2 A schematic diagram of a process for detecting whether each image prediction result matches in one embodiment;
[0038] Figure 3 is a schematic flow chart of an image sample screening method in another embodiment;
[0039] Figure 4 is a structural block diagram of an image sample screening device in one embodiment;
[0040] Figure 5 FIG. 4 is a diagram showing the internal structure of a cleaning robot in one embodiment. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] In one embodiment, Figure 1 As shown, a method for screening image samples is provided. This embodiment uses the method applied to a terminal as an example. The terminal may be a cleaning robot, such as a mopping cleaning robot or a sweeping cleaning robot. It is understandable that the method may also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0043] Step 102: Obtain multiple image prediction results output by the image prediction model for the target image.
[0044] Among them, the image prediction model can be used to complete different image prediction tasks, so the image prediction model can be an image classification model, an image segmentation model, a target detection model, etc., and the image prediction task can be an image classification task, an image segmentation task, a target detection task, etc.
[0045] As an example, the image segmentation model can be used to segment the ground area in the image, or to segment the foreground and background of the image; the target detection model can be used to detect the type and position of the target object in the image, which target object can be a living target or a non-living target, etc. For example, the living target can be a pedestrian, and the non-living target can be an obstacle.
[0046] Step 104 , detecting whether the prediction results of each image match according to the matching detection rule called by the model identifier of each image prediction model.
[0047] Among them, the matching detection rules between the image prediction results output by different image prediction models are usually different. For example, the image segmentation model and the image classification model can correspond to the matching detection rule A, and the image segmentation model and the target detection model can correspond to the matching detection rule B; the model identifier can be the model type identifier of the image prediction model.
[0048] As an example, the matching detection rule may be a packaged matching detection script, and the corresponding image prediction result may be input into the matching detection script, and the matching detection script may output the corresponding matching detection result.
[0049] As an example, the matching detection rule may be a rule component set based on a rule engine, the corresponding image prediction result is input into the rule component, and then the execution result obtained by executing the rule component is the matching detection result.
[0050] As an example, step 104 includes: obtaining the model identifier of each image prediction model, and calling the matching detection rule corresponding to each image prediction result according to each model identifier; and detecting whether each image prediction result matches by executing the matching detection rule on each image prediction result.
[0051] In one embodiment, the image prediction model includes a target detection model and an image classification model, and the plurality of image prediction results include a target detection result output by the target detection model and an image classification result output by the image classification model; detecting whether the image prediction results match includes:
[0052] Determine the object detection type identifier of the target object in the target detection result and the image classification identifier in the image classification result; if the object detection type identifier and the image classification identifier match, determine that the target detection result and the image classification result match; if the object detection type identifier and the image classification identifier do not match, determine that the target detection result and the image classification result do not match.
[0053] The object detection type identifier is used to identify the object type detected during target detection, and the image classification identifier is used to identify the image category detected during image classification for the entire target image.
[0054] As an example, the image classification model may be a room classification model, the image classification result may be a room classification result, and the image classification identifier may be a room classification identifier. In one example, it is assumed that the target image is a room image. In this case, the image type may be a bedroom or a bathroom, etc., and the object type may be a toilet or a sofa, etc. If the object detection type identifier is used to identify the type of the target object as a toilet, and the image classification identifier is used to identify the type of the target image as a bedroom, and there is a toilet in the bedroom, this is obviously contrary to common sense. Therefore, it can be determined that the target detection result and the image classification result do not match, and there are erroneous image prediction results in the target detection result and the image classification result; if the object detection type identifier is used to identify the type of the target object as a sofa, and the image classification identifier is used to identify the type of the target image as a bathroom, and there is a sofa in the bathroom, this is also obviously contrary to common sense. Therefore, it can be determined that the target detection result and the image classification result do not match, and there are erroneous image prediction results in the target detection result and the image classification result.
[0055] In this way, by detecting whether the object detection type identifier of the target object and the image classification identifier of the target image match, it is possible to accurately determine whether the target detection result and the image classification result are inconsistent, thereby realizing whether the target detection result and the image classification result match, laying the foundation for subsequent detection of whether the target image is a difficult image sample.
[0056] Step 106: If the prediction results of each image do not match, the target image is determined as a difficult image sample.
[0057] Among them, if the image prediction results do not match, it means that there are erroneous prediction results in the image prediction results, and there are models in the image prediction models that cannot accurately predict the target image, so the target image is determined as a difficult image sample; if the image prediction results match, it means that there is no erroneous prediction result in the image prediction results, and each image prediction model can accurately predict the target image, so the target image is not determined as a difficult image sample.
[0058] As an example, when it is determined that the target image is a difficult image sample, the training sample set recovered to the image prediction model can selectively recover difficult image samples that are helpful for training the image task model from the full amount of training image samples, rather than recovering the full amount of training image samples. This can improve the efficiency of screening training image samples for the image task model.
[0059] In the above-mentioned image sample screening method, multiple image prediction results output by the image prediction model for the target image are obtained; and the matching detection rules called according to the model identifier of each image prediction model are used to detect whether each image prediction result matches. In this way, the image prediction results output by different image prediction models for the target image can be matched according to the matching detection rules corresponding to different types of image prediction models. If the image prediction results do not match, it means that there are erroneous prediction results in the image prediction results. At this time, the target image can be determined as a difficult image sample. In this way, when the training image samples are recycled to the training sample set, the difficult image samples with better training effects can be selectively recycled to the training sample set, rather than recycling the full amount of training image samples to the training sample set. Therefore, the data redundancy of the training sample set during the iterative training of the image task model can be reduced.
[0060] In one embodiment, Figure 2 As shown, the image prediction model includes a target detection model and an image segmentation model, and the multiple image prediction results include a target detection result output by the target detection model and a first image segmentation result output by the image segmentation model; detecting whether each image prediction result matches includes:
[0061] Step 202: extracting object image position information of the target object in the target detection result and region image position information of each segmented image region in the first image segmentation result.
[0062] Among them, the target object is an image target for target detection, and target detection can be used to detect the position information and object category of the target object. Therefore, the target detection result may include the object image position information of the target object, and the object image position information is the position information of the target object in the target image, for example, it may be the pixel position coordinates of the target object in the target image; image segmentation can be used to segment the target image into multiple segmented image regions, and therefore the first image segmentation result may include the regional image position information of multiple segmented image regions, and the regional image position information is the position information of the segmented image region in the target image, for example, it may be the pixel position coordinates of the segmented image region in the target image.
[0063] Step 204: Detect whether the target detection result and the first image segmentation result match based on the object image position information and the region image position information.
[0064] When performing image segmentation, in order to ensure the integrity of the segmented image region, the same image object in the target image is usually not segmented into different segmented image regions.
[0065] As an example, step 204 includes: based on the object image position information and the regional image position information, detecting whether there is an intersection between the target object and the regional segmentation boundaries between each segmented image region; if there is an intersection, it means that the target object has been segmented into different segmented image regions, and therefore it is determined that the target detection result and the first image segmentation result do not match; if there is no intersection, it means that the target object has not been segmented into different segmented image regions, and therefore it is determined that the target detection result and the first image segmentation result match.
[0066] As an example, step 204 includes: based on the object image position information and the regional image position information, detecting whether there is an intersection between the target object and the regional segmentation boundaries between each segmented image region; if there is an intersection, it means that the target object has been segmented into different segmented image regions, and therefore it is determined that the target detection result and the first image segmentation result do not match; if there is no intersection, locating the target image region to which the target object belongs in each segmented image region based on the object image position information and the regional image position information; obtaining the object detection type identifier of the target object in the target detection result, and obtaining the regional type identifier of the target image region in the first image segmentation result; and detecting whether the target detection result and the first image segmentation result match based on the object detection type identifier and the regional type identifier.
[0067] As an example, assume that the object detection type identifier of the target object is A, which identifies the object detection type of the target object as an obstacle, and the region type identifier is B, which identifies the region type of the target image area as the ground. If there is an obstacle in an image area, the image area should be segmented into a background area rather than a ground area. At this time, the target detection result and the image segmentation result are contradictory, that is, they do not match. Therefore, it can be preset that the object detection type identifier A and the region type identifier B do not match. Based on this, it can be realized to detect whether the target detection result and the first image segmentation result match according to the object detection type identifier and the region type identifier.
[0068] In one embodiment, detecting whether the target detection result matches the first image segmentation result according to the object image position information and the image position information of each region includes:
[0069] According to the object image position information and the regional image position information, the target image area where the target object is located in each segmented image area is located; the object detection type identifier corresponding to the target object in the target detection result and the region type identifier corresponding to the target image area are obtained; according to the object detection type identifier and the region type identifier, whether the target detection result and the first image segmentation result match.
[0070] Specifically, according to the object image position information and the regional image position information, the target image area where the image position of the target object is located is located in each segmented image area; the object detection type identifier corresponding to the target object in the target detection result and the region type identifier corresponding to the target image area are obtained; if the object detection type identifier and the region type identifier match, it is considered that the target detection result matches the first image segmentation result; if the object detection type identifier and the region type identifier do not match, it is considered that the target detection result does not match the first image segmentation result. In this way, by detecting whether the object detection type of the target object and the image segmentation region type of the image region where the target object is located are compatible, it is possible to accurately detect whether the target object and each segmented image region are compatible in terms of image region type, thereby realizing the detection of whether the target detection result and the first image segmentation result match.
[0071] In one embodiment, detecting whether the target detection result matches the first image segmentation result according to the object image position information and the region image position information includes:
[0072] According to the object image position information and the image position information of each region, detect whether the target object in the target image has an intersection with the regional segmentation boundaries between each segmented image region; if there is an intersection, obtain the object image region area of the target object in different segmented image regions; according to the area of each object image region, detect whether the target detection result and the first image segmentation result match.
[0073] It should be noted that, considering the errors in target detection and image segmentation, if the target object in the target image only has an intersection with the regional segmentation boundaries between the edge areas and the segmented image areas, it should not be considered that the target detection result and the first image segmentation result do not match.
[0074] Specifically, based on the object image position information and the image position information of each region, it is detected whether the target object in the target image has an intersection with the regional segmentation boundaries between each segmented image region; if there is an intersection, the object image region area of the target object in the target image in different segmented image regions is obtained, wherein the object image region area can be represented by the number of pixels; the largest target area is determined in the object image region area, and the ratio between the target area and the total area of the target object in the target image is calculated to obtain the object region area ratio; if the object region area ratio is greater than a preset ratio, it means that in the target image, only the edge region of the target object has an intersection with the regional segmentation boundaries between each segmented image region, and therefore it is considered that the target detection result and the first image segmentation result have matching boundaries; if the object region area ratio is not greater than the preset ratio, it means that in the target image, the non-edge region of the target object has an intersection with the regional segmentation boundaries between each segmented image region, and therefore it is considered that the target detection result and the first image segmentation result do not match.
[0075] In this way, on the basis of detecting whether the target detection result and the first image segmentation result match by detecting whether the target object in the target image intersects with the regional segmentation boundaries between each segmented image area, the object image area area of the target object in different segmented image areas is introduced to further detect whether only the edge area of the target object intersects with the regional segmentation boundaries between each segmented image area. This can eliminate the negative impact of the errors in target detection and image segmentation on the detection of whether the target detection result and the first image segmentation result match, thereby improving the accuracy of detecting whether the target detection result and the first image segmentation result match.
[0076] In this embodiment, the object image position information of the target object in the target detection result and the regional image position information of each segmented image area in the first image segmentation result are extracted; based on the object image position information and the regional image position information, it is detected whether the target detection result and the first image segmentation result match. In this way, it is possible to accurately detect whether the target object and each segmented image area are adapted in image position based on the object image position information of the target detection result and the regional image position information in the first image segmentation result, thereby detecting whether the target detection result and the first image segmentation result match, which lays a foundation for identifying whether the target image is a difficult image sample.
[0077] In one embodiment, Figure 3 As shown, the image prediction model includes an image segmentation model, and the image prediction result includes a first image segmentation result output by the image segmentation model; after detecting whether each image prediction result matches according to a matching detection rule called according to a model identifier of each image prediction model, the method further includes:
[0078] Step 302: If the prediction results of each image match, then determine each neighborhood time frame image corresponding to the target image according to a preset time window, wherein the target image and the neighborhood time frame image are captured by a preset camera device, and the camera parameters of the preset camera device have a change range within the preset time window that is less than a preset change range.
[0079] Among them, the target image can be captured by a preset camera device, and the change amplitude of the camera parameters of the preset camera device within the preset time window is less than the preset change amplitude, that is, the camera parameters of the preset camera device will not produce sudden changes, wherein the camera parameters are parameters that will affect the imaging effect of the final camera image, such as the shooting angle and camera focal length, etc.; the domain time frame image is an image captured within the same preset time window as the target image, for example, assuming the preset time window is 10 milliseconds, then the images captured 10 milliseconds before and 10 milliseconds after the current time of capturing the target image can be used as the neighborhood time frame images of the target image.
[0080] As an example, the preset camera device can be a shooting cleaning robot equipped with a camera. The shooting cleaning robot can be set so that it will not move or rotate over long distances. Therefore, the change range of the camera parameters of the shooting cleaning robot within the preset time window will be smaller than the preset change range.
[0081] Step 304 , detecting the target detection model for each neighborhood time frame image, the degree of regional overlap between the segmented image region in the second image segmentation result and the segmented image region in the first image segmentation result.
[0082] Among them, when performing image segmentation, since the change range of the camera parameters of the preset camera device within the preset time window is less than the preset change range, the image segmentation results of the target image and the neighborhood time frame image will not differ too much. If the difference is too large, it means that there are wrong image segmentation results, and there are difficult image samples in the target image and each neighborhood time frame image.
[0083] As an example, step 304 includes: determining the first region position coordinates of each segmented image region in the first image segmentation result, and determining the second region position coordinates of each segmented image region in the second image segmentation result output by the target detection model for each neighborhood time frame image; based on the first region position coordinates and each second region position coordinates, detecting the region overlap between the segmented image region in each second image segmentation result and the segmented image region in the first image segmentation result.
[0084] Step 306: Determine whether to determine the target image as a difficult image sample according to the overlap degree of each region.
[0085] As an example, step 306 includes: if the overlap of each region is not greater than the preset overlap threshold, it means that there are image segmentation results with prediction errors in the first image segmentation result and each second image segmentation result, and the target image is determined as a difficult image sample; if the overlap of each region is greater than the preset overlap threshold, it means that there is no difficult image sample in the target image and each field time frame image, and the target image is not determined as a difficult image sample.
[0086] In one embodiment, determining whether to determine the target image as a difficult image sample according to the overlap degree of each region includes:
[0087] If the variation range of the overlap between the overlaps of each region is not less than the preset variation range, the target image is determined as a difficult image sample; if the variation range of the overlap between the overlaps of each region is less than the preset variation range, the target image is not determined as a difficult image sample.
[0088] Among them, if the variation range of the overlap between the overlaps of each region is not less than the preset variation range, it means that there are image segmentation results with prediction errors in the first image segmentation result and each second image segmentation result, and the target image is determined as a difficult image sample; if the variation range of the overlap between the overlaps of each region is less than the preset variation range, it means that there are no difficult image samples in the target image and the time frame images of each field, and the target image is not determined as a difficult image sample.
[0089] It should be noted that when the variation range of the overlap between the overlaps of various regions is less than the preset variation range, it means that there are hard-to-solve image samples in the target image and the time frame images of each field. Therefore, the target image is only likely to be a hard-to-solve image sample, but not necessarily a hard-to-solve image sample. If the target image is recycled to the training sample set of the image prediction model at this time, there is a possibility of recycling non-hard-to-solve image samples, and there is still a certain amount of data redundancy in the training sample set.
[0090] In one embodiment, determining whether to determine the target image as a difficult image sample according to the overlap degree of each region includes:
[0091] If the variation range of the overlap between the overlaps of each region meets the preset variation range distribution, the target image is determined as a difficult image sample; if the variation range of each overlap does not meet the preset variation range distribution, the target image is not determined as a difficult image sample.
[0092] Among them, the preset change amplitude distribution is the change amplitude distribution when the target image has the maximum probability of being a difficult image sample. The preset change amplitude distribution can be obtained by analyzing the image segmentation results of multiple image samples taken historically; if the change amplitudes of each overlap degree corresponding to the target image meet the preset change amplitude distribution, it means that the target image is more likely to be a difficult image sample.
[0093] As an example, assuming that the regional overlap is the overlap between the segmented image areas of images taken in adjacent time frames, the preset change amplitude distribution can be a change amplitude distribution in which the overlap variation corresponding to the time frame in which the target image is located on the time axis produces a sudden change. If the overlap variation corresponding to the time frame in which the target image is located produces a sudden change, it means that there are large differences in the image segmentation results of the target image and the neighboring time frames of the previous and next adjacent frames. At this time, the target image is likely to be a difficult sample.
[0094] In the above embodiment, if the variation range of the overlap between the overlaps of various regions conforms to the preset variation range distribution, it means that there is a high possibility that the target image is a difficult image sample, and thus the target image is determined to be a difficult image sample, rather than simply judging whether the target image is a difficult image sample based on the size of the regional overlap. The judgment basis is more reliable, and thus the accuracy of detecting whether the target image is a difficult image sample is higher. Furthermore, by selectively recycling difficult image samples to the training sample set during the iterative training of the image prediction model, it helps to reduce the data redundancy of the training sample set of the image prediction model.
[0095] In one embodiment, first, multiple image prediction results output by the image prediction model for the target image are obtained, and the corresponding matching detection rules are called according to the model identifier of each image prediction model; based on the matching detection rules, the object image position information of the target object in the target detection result and the regional image position information of each segmented image area in the first image segmentation result are extracted.
[0096] Further, according to the object image position information and the image position information of each region, it is detected whether the target object in the target image has an intersection with the regional segmentation boundary between each segmented image region; if there is an intersection, the object image region area of the target object in the target image in different segmented image regions is obtained, wherein the object image region area can be represented by the number of pixels; the largest target area is determined in the object image region area, and the ratio between the target area and the total area of the target object in the target image is calculated to obtain the object region area ratio; if the object region area ratio is greater than the preset ratio, it means that in the target image, only the edge region of the target object has an intersection with the regional segmentation boundary between each segmented image region, so it is considered that the target detection result and the first image segmentation result have matching boundaries; if the object region area ratio is not greater than the preset ratio, it means that in the target image, the non-edge region of the target object has an intersection with the regional segmentation boundary between each segmented image region, so it is considered that the target detection result and the first image segmentation result do not match. In this way, the negative impact of the error of target detection and image segmentation on the detection of whether the target detection result and the first image segmentation result match can be eliminated, so the accuracy of detecting whether the target detection result and the first image segmentation result match can be improved.
[0097] Furthermore, if the target detection result and the first image segmentation result do not match, it means that there are erroneous prediction results in each image prediction result, and there is a model in each image prediction model that cannot accurately predict the target image, so the target image is determined as a difficult image sample; if the image prediction results match, it means that there is no erroneous prediction result in each image prediction result, and each image prediction model can accurately predict the target image, so the target image is not used as a difficult image sample. In this way, when the training image samples are recycled to the training sample set, the difficult image samples with better training effects can be selectively recycled to the training sample set, rather than recycling the full amount of training image samples to the training sample set, thereby reducing the data redundancy of the training sample set when iteratively training the image task model.
[0098] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0099] Based on the same inventive concept, the embodiment of the present application also provides an image sample screening device for implementing the image sample screening method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more image sample screening device embodiments provided below can refer to the limitations of the image sample screening method above, and will not be repeated here.
[0100] In one embodiment, Figure 4 As shown, an image sample screening device is provided, comprising: an acquisition module 402, a matching detection rule determination module 404, a matching detection module 406 and a difficult example image sample screening module 408, wherein:
[0101] An acquisition module, used to acquire multiple image prediction results output by the image prediction model for the target image;
[0102] A matching detection rule determination module, used to determine the matching detection rule commonly corresponding to each of the image prediction results according to the model identifier of each of the image prediction models;
[0103] A matching detection module, used to detect whether the image prediction results match each other according to the matching detection rules called by the model identifier of each image prediction model;
[0104] The difficult image sample screening module is used to determine the target image as a difficult image sample if the image prediction results do not match.
[0105] In one embodiment, the image prediction model includes a target detection model and an image segmentation model, the multiple image prediction results include a target detection result output by the target detection model and a first image segmentation result output by the image segmentation model; the matching detection module is further used to:
[0106] Extract the object image position information of the target object in the target detection result and the regional image position information of each segmented image region in the first image segmentation result; and detect whether the target detection result and the first image segmentation result match based on the object image position information and the regional image position information.
[0107] In one embodiment, the matching detection module is further used for:
[0108] According to the object image position information and the regional image position information, locate the target image area where the target object is located in each of the segmented image areas; obtain the object detection type identifier corresponding to the target object in the target detection result, and the region type identifier corresponding to the target image area; according to the object detection type identifier and the region type identifier, detect whether the target detection result and the first image segmentation result match.
[0109] In one embodiment, the matching detection module is further used for:
[0110] According to the object image position information and the regional image position information, detect whether the target object in the target image has an intersection with the regional segmentation boundaries between the segmented image regions; if there is an intersection, obtain the object image region area of the target object in the different segmented image regions; according to the area of each object image region, detect whether the target detection result and the first image segmentation result match.
[0111] In one embodiment, the image prediction model includes a target detection model and an image classification model, the multiple image prediction results include target detection results output by the target detection model and image classification results output by the image classification model; the matching detection module is further used to:
[0112] Determine the object detection type identifier of the target object in the target detection result and the image classification identifier in the image classification result; if the object detection type identifier and the image classification identifier match, determine that the target detection result and the image classification result match; if the object detection type identifier and the image classification identifier do not match, determine that the target detection result and the image classification result do not match.
[0113] In one embodiment, the image prediction model includes an image segmentation model, and the image prediction result includes a first image segmentation result output by the image segmentation model; the device further includes:
[0114] An image segmentation difficult sample recovery module is used to determine, if the image prediction results match, each neighborhood time frame image corresponding to the target image according to a preset time window, wherein the target image and the neighborhood time frame image are captured by a preset camera device, and the change amplitude of the camera parameters of the preset camera device within the preset time window is less than the preset change amplitude; detect the regional overlap between the segmented image area in the second image segmentation result of the target detection model for each neighborhood time frame image and the segmented image area in the first image segmentation result; and determine whether to use the target image as a difficult image sample based on the overlap of each area.
[0115] In one embodiment, the image segmentation difficult sample recovery module is further used for:
[0116] If the variation range of the overlap between the overlaps of the regions is not less than the preset variation range, the target image is determined as a difficult image sample; if the variation range of the overlap between the overlaps of the regions is less than the preset variation range, the target image is not determined as a difficult image sample.
[0117] In one embodiment, the image segmentation difficult sample recovery module is further used for:
[0118] If the variation range of the overlap between the overlaps of the regions conforms to the preset variation range distribution, the target image is determined as a difficult image sample; if the variation range of the overlap does not conform to the preset variation range distribution, the target image is not determined as a difficult image sample.
[0119] Each module in the above-mentioned image sample screening device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the cleaning robot in the form of hardware, or can be stored in the memory of the cleaning robot in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0120] In one embodiment, a cleaning robot is provided. The cleaning robot may be a sweeping robot or a mopping robot, etc. A sensor and a cleaning module are arranged on the body of the cleaning robot. The sensor may be a camera or a laser radar, etc. The cleaning module may be a mop or a dust removal module, etc. The sensor may be used to obtain a sample image. The cleaning robot may use the above-mentioned image sample screening method to identify and recover difficult sample images in the sample images taken by the camera. The cleaning robot may use the above-mentioned cleaning module to perform cleaning tasks. The internal structure diagram thereof may be shown as follows: Figure 5 As shown. The cleaning robot includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. Among them, the processor of the cleaning robot is used to provide computing and control capabilities. The memory of the cleaning robot includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the cleaning robot is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an image sample screening method is implemented.
[0121] Those skilled in the art will understand that Figure 5 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the cleaning robot to which the scheme of the present application is applied. The specific cleaning robot may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0122] In one embodiment, a cleaning robot is provided, comprising a body, a cleaning module arranged on the body, a sensor, a memory, and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0123] Acquire multiple image prediction results output by the image prediction model for the target image; detect whether each of the image prediction results matches according to the matching detection rules called by the model identifier of each of the image prediction models; if the image prediction results do not match, determine the target image as a difficult image sample.
[0124] In one embodiment, the image prediction model includes a target detection model and an image segmentation model, and the multiple image prediction results include a target detection result output by the target detection model and a first image segmentation result output by the image segmentation model; when the processor executes the computer program, the following steps are also implemented:
[0125] Extract the object image position information of the target object in the target detection result and the regional image position information of each segmented image region in the first image segmentation result; and detect whether the target detection result and the first image segmentation result match based on the object image position information and the regional image position information.
[0126] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0127] According to the object image position information and the regional image position information, locate the target image area where the target object is located in each of the segmented image areas; obtain the object detection type identifier corresponding to the target object in the target detection result, and the region type identifier corresponding to the target image area; according to the object detection type identifier and the region type identifier, detect whether the target detection result and the first image segmentation result match.
[0128] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0129] According to the object image position information and the regional image position information, detect whether the target object in the target image has an intersection with the regional segmentation boundaries between each of the segmented image regions; if there is an intersection, obtain the object image region area of the target object in different segmented image regions; according to each of the object image region areas, detect whether the target detection result and the first image segmentation result match.
[0130] In one embodiment, the image prediction model includes a target detection model and an image classification model, and the multiple image prediction results include target detection results output by the target detection model and image classification results output by the image classification model; when the processor executes the computer program, the following steps are also implemented:
[0131] Determine the object detection type identifier of the target object in the target detection result and the image classification identifier in the image classification result; if the object detection type identifier and the image classification identifier match, determine that the target detection result and the image classification result match; if the object detection type identifier and the image classification identifier do not match, determine that the target detection result and the image classification result do not match.
[0132] In one embodiment, the image prediction model includes an image segmentation model, and the image prediction result includes a first image segmentation result output by the image segmentation model; when the processor executes the computer program, the following steps are also implemented:
[0133] If the prediction results of each of the images match, then determine the neighborhood time frame images corresponding to the target image according to a preset time window, wherein the target image and the neighborhood time frame images are captured by a preset camera device, and the change amplitude of the camera parameters of the preset camera device within the preset time window is less than the preset change amplitude; detect the regional overlap between the segmented image area in the second image segmentation result of the target detection model for each neighborhood time frame image and the segmented image area in the first image segmentation result; and determine whether to use the target image as a difficult image sample based on the overlap of each area.
[0134] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0135] If the variation range of the overlap between the overlaps of the regions is not less than the preset variation range, the target image is determined as a difficult image sample; if the variation range of the overlap between the overlaps of the regions is less than the preset variation range, the target image is not determined as a difficult image sample.
[0136] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0137] If the variation range of the overlap between the overlaps of the regions conforms to the preset variation range distribution, the target image is determined as a difficult image sample; if the variation range of the overlap does not conform to the preset variation range distribution, the target image is not determined as a difficult image sample.
[0138] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0139] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0140] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0141] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0142] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for screening image samples, It is characterized in that The method comprises: Obtain multiple image prediction results output by the image prediction model for the target image; Detecting whether the image prediction results match each other according to the matching detection rules called by the model identifiers of the image prediction models; If the image prediction results do not match, the target image is determined as a difficult image sample.
2. The method according to claim 1, It is characterized in that The image prediction model includes a target detection model and an image segmentation model, and the multiple image prediction results include a target detection result output by the target detection model and a first image segmentation result output by the image segmentation model; The detecting whether the image prediction results match each other includes: Extracting object image position information of the target object in the target detection result and region image position information of each segmented image region in the first image segmentation result; According to the object image position information and the region image position information, it is detected whether the target detection result matches the first image segmentation result.
3. The method according to claim 2, It is characterized in that The detecting, according to the object image position information and the region image position information, whether the target detection result matches the first image segmentation result includes: Locating the target image region where the target object is located in each of the segmented image regions according to the object image position information and the region image position information; Obtaining an object detection type identifier corresponding to the target object in the target detection result and an area type identifier corresponding to the target image area; According to the object detection type identifier and the region type identifier, it is detected whether the target detection result and the first image segmentation result match.
4. The method according to claim 2, It is characterized in that The detecting, according to the object image position information and the region image position information, whether the target detection result matches the first image segmentation result includes: According to the object image position information and the region image position information, detecting whether a target object in the target image intersects with a region segmentation boundary between each of the segmented image regions; If there is an intersection, obtaining the object image region area of the target object in different segmented image regions; According to the area of each object image region, it is detected whether the target detection result and the first image segmentation result match each other.
5. The method according to claim 1, It is characterized in that The image prediction model includes a target detection model and an image classification model, and the multiple image prediction results include target detection results output by the target detection model and image classification results output by the image classification model; The detecting whether the image prediction results match each other includes: Determining an object detection type identifier of a target object in the target detection result and an image classification identifier in the image classification result; If the object detection type identifier and the image classification identifier match, determining that the target detection result and the image classification result match; If the object detection type identifier and the image classification identifier do not match, it is determined that the target detection result and the image classification result do not match.
6. The method according to claim 1, It is characterized in that The image prediction model includes an image segmentation model, and the image prediction result includes a first image segmentation result output by the image segmentation model; after the matching detection rule called according to the model identifier of each of the image prediction models is detected to see whether each of the image prediction results matches, the method further includes: If the image prediction results match, then determine the neighboring time frame images corresponding to the target image according to the preset time window, wherein the target image and the neighboring time frame images are obtained by being photographed by a preset camera device, and the change range of the camera parameters of the preset camera device within the preset time window is less than the preset change range; Detecting the degree of regional overlap between the segmented image area in the second image segmentation result and the segmented image area in the first image segmentation result for each neighborhood time frame image of the target detection model; According to the overlap degree of each region, it is determined whether to use the target image as a difficult image sample.
7. The method according to claim 6, It is characterized in that The step of determining whether to use the target image as a difficult image sample according to the overlap degree of each region includes: If the variation range of the overlap between the overlaps of the regions is not all less than the preset variation range, the target image is determined as a difficult image sample; If the variation range of the overlap between the overlaps of the regions is smaller than the preset variation range, the target image is not determined as a difficult image sample.
8. The method according to claim 6, It is characterized in that The step of determining whether to use the target image as a difficult image sample according to the overlap degree of each region includes: If the variation range of the overlap between the overlaps of the regions meets the preset variation range distribution, the target image is determined as a difficult image sample; If the variation ranges of the overlap degrees do not conform to the preset variation range distribution, the target image is not determined as a difficult image sample.
9. An image sample screening device, It is characterized in that The device comprises: An acquisition module, used to acquire multiple image prediction results output by the image prediction model for the target image; A matching detection rule determination module, used to determine the matching detection rule commonly corresponding to each of the image prediction results according to the model identifier of each of the image prediction models; A matching detection module, used to detect whether the image prediction results match each other according to the matching detection rules called by the model identifier of each image prediction model; The difficult image sample screening module is used to determine the target image as a difficult image sample if the image prediction results do not match.
10. A cleaning robot, comprising a body, a cleaning module arranged on the body, a sensor, a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Target detection-oriented error detection sample screening method
CN121353636A