Verification of physical presence of detected object in image

By combining object detectors and depth information, and using depth changes to verify object presence, the problem of false detection caused by misleading visual cues in autonomous driving is solved, improving detection accuracy and system safety.

CN120871282APending Publication Date: 2025-10-31ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544128.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-04-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In autonomous driving environments, visual cues can lead to misleading object detection, resulting in incorrect avoidance or emergency braking, increasing the risk of traffic accidents.

Method used

By combining object detectors and depth information, the presence of objects is verified by utilizing depth variations. Visual cues are filtered using depth information, and depth maps are processed in particular through derivatives and morphological operators to ensure detection accuracy.

Benefits of technology

It improves the accuracy of object detection, reduces false detections, ensures the safety and reliability of autonomous driving systems, and avoids unnecessary evasive maneuvers or emergency braking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871282A_ABST
    Figure CN120871282A_ABST
Patent Text Reader

Abstract

Verification of a physical presence of a detected object in an image is provided. A method (100) for verifying the presence of an object (3) detected in a scene (1) on the basis of at least one image (2) of the scene (1) comprises the steps of: assigning (110), by a given object detector (4), at least one image region (2a) in the image (2) to the object (3), the image region (2a) indicating the presence of the object (3) in the scene (1); obtaining (120) depth information (5) relating to the image area (2a), said depth information (5) indicating a distance between a sensor for acquiring the image (2) and a scene area (1a) in the scene (1) corresponding to the image area (2a); determining (130) whether the depth information (5) comprises a depth change (5a *) expected in the presence of the object (3) assigned to the image area (2a) by the object detector (4); and if the determination is positive, determining (140) that the object detector (4) assigns the image region (2a) to an effective detection (3 *) that the object (3) is the object (3).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image analysis, and more particularly to detecting objects in a scene based on images taken of the scene. Background Technology

[0002] Autonomous operation of vehicles or robots in corporate settings or even on public roads requires continuous monitoring of their environment. Acquiring and analyzing images of these environments is a crucial part of this monitoring. Detecting objects that the vehicle or / or robot may collide with is particularly important.

[0003] When an object is detected from visual cues, it is important to identify all objects that are physically present in the scene and visible in the image. However, visual cues can be misleading, and objects may be identified even where they do not actually exist. Such false detections during autonomous driving could lead to evasive maneuvers or emergency braking, potentially startling other road users and causing an accident. Summary of the Invention

[0004] This invention provides a method for verifying the presence of a detected object in a scene based on at least one image of the scene. Images can be acquired using any modality (e.g., a still or video camera, or a thermal camera). The image can also be a multimodal image with pixels, whose pixel values ​​are derived from measurements captured using the multimodal method. Specifically, the image can be a color image with pixels, where multiple intensities of basic colors in a given color space are assigned to these pixels. That is, for each pixel, each basic color has an intensity value. Exemplary color spaces are RGB (red, green, blue) and CMYK (cyan, magenta, yellow, and key color = black).

[0005] In this method, at least one image region in the image is assigned to an object, which indicates the object's presence in the scene by a given object detector. That is, the object detector outputs that this region in the image contains an object instance, with or without a specific identifier of the object type. For example, the image region could be a bounding box surrounding the object instance plus a bit of background. However, the image region could also be a more precise outline of the object instance that distinguishes it from the background.

[0006] In the context of this invention, an object is a physical entity having a three-dimensional structure. Specifically, an object can be an entity standing on or otherwise protruding from a flat surface. In the context of autonomous driving, such objects that a vehicle or robot might collide with are primarily relevant objects. For example, a person or animal standing on the road is a relevant object. However, painted road markings or manhole covers that are substantially flush with the road are not considered objects because vehicles or robots can drive right over them.

[0007] Obtain depth information associated with the image region. This depth information indicates the distance between the sensor used to acquire the image and the scene region corresponding to the image region. That is, the depth information is not limited to this distance, but can also be a quantity commensurate with this distance.

[0008] Specifically, depth information may include one or more of the following:

[0009] ● Monocular depth estimation from images;

[0010] ● Depth information from a stereo camera; and / or

[0011] ● Depth information is obtained by measuring the distance to the scene using an electromagnetic interrogation radiation beam.

[0012] In particular, in autonomous driving use cases, sensors used to query radiation for distance measurements using radar or lidar already exist on vehicles. This means that these existing sensors can be reused, reducing the cost of implementing this method on vehicles.

[0013] Therefore, it is advantageous that depth information is obtained from at least one sensor carried by the same vehicle from which images have already been acquired. That is, the vehicle can carry both a sensor for acquiring images and a sensor for acquiring depth information.

[0014] Determine whether the depth information includes the expected depth variations in the presence of an object assigned to the image region by the object detector. In other words, if an object does exist in the scene where the detection indicates it, then it inevitably produces depth variations. This means that if these depth variations are missing from the depth information, then the object cannot exist in the scene, at least not where the detection of object instances in the image region indicates it.

[0015] Therefore, if the expected depth change is determined to exist, then assigning the image region to the object by the object detector is considered effective object detection. Here, determining whether the expected depth change exists can include both qualitatively determining whether the depth change fundamentally exists and quantitatively determining the amount of the depth change.

[0016] For example, a portion of an image related to a distant scene or a part thereof may exhibit only small depth variations. Conversely, a nearby scene or a portion thereof may show much greater depth variations. For instance, a road surface region might show depth variations across every pixel in the region of interest. Therefore, the indication of an object can be associated, for example, with the presence of some local depth variations.

[0017] Depth information has been found to be a particularly useful tool for resolving ambiguities about whether visual cues in an image indicate the presence of an actual object, or whether visual cues merely comprise texture and / or color variations on an object-free surface. Depth information provides geometric cues that offer a more comprehensive view and complement the visual (appearance) cues.

[0018] In particular, in autonomous driving applications, roads often exhibit characteristics that could be mistaken for objects. For example, roads include numerous markings such as lane lines, speed limits, lane-specific directions, instructions regarding right-of-way, instructions regarding who can use a lane, and even unofficial markings such as graffiti. Additionally, there are manhole covers and other devices flush with the road surface. Furthermore, roads themselves may exhibit textural variations due to changes in their surface composition. For instance, a road constructed some time ago might be dug up and then closed with new asphalt gravel of a different texture, or potholes might be repaired with temporary asphalt of a different texture.

[0019] In autonomous driving applications, the sources of potential false object detection are even more numerous. For example, a bus might carry an advertisement depicting a scene different from the actual physical environment, such as a family marveling at a brand-new car. The family members and the car shown in the advertisement should not be considered actual objects, even if they are presented from a perspective that suggests otherwise. Such false detection could lead to incorrect reactions from downstream systems, such as swerves or emergency braking, which are undesirable because other road users did not anticipate them.

[0020] In other words, depth information is particularly well-suited for determining whether visual cues indicating the presence of an object actually belong to the category of "objects" as potential collision targets, or whether they are related to things that can be safely ignored for autonomous driving purposes. Specifically, not everything that differs from the road surface is an object. For example, a flush manhole cover is very different from the road surface, but it is not an object that a vehicle or robot might collide with.

[0021] Filtering using depth information is robust. Depth information becomes inaccurate at long distances. However, this isn't a problem because objects won't be filtered out when there's no change in depth. Therefore, we won't decrease recall. Thus, this filtering mechanism is most effective for near-range objects (though false positives might be more dangerous).

[0022] In a particularly advantageous embodiment, an operator highlighting depth changes is applied to the depth information such that the more dramatic the depth change, the more it is highlighted. This is further used to distinguish between depth changes caused by the presence of objects and stable depth changes caused by the viewpoint between the camera and the ground. Stable depth changes indicate flat surfaces, i.e., drivable surfaces in the case of autonomous driving. More dramatic depth changes indicate the presence of objects protruding from or standing on flat surfaces.

[0023] That is, in another particularly advantageous embodiment, the operator is configured to distinguish between a consistent depth gradient of a flat surface and a more significant depth variation protruding from these surfaces.

[0024] If the depth changes are already more dramatic, examples of operators that further highlight these changes include the derivative operator and the Sobel operator. The Sobel operator is a convolutional edge detection filter that computes the first derivative of the pixel value while smoothing it in a direction perpendicular to the direction in which the derivative is computed. In particular, the result of applying the Sobel operator can include a gradient image that highlights the edges of the original image.

[0025] In another particularly advantageous embodiment, the image is divided into pixels. Depth information includes a depth map. This depth map assigns a value to each pixel in the image region, indicating the distance between the sensor and the location in the scene represented by the corresponding pixel. That is, the value does not need to be exactly this distance; rather, it only needs to be commensurate with this distance. The depth map then adds a third dimension to the original two-dimensional image. It can be perceived as a depth image corresponding to the original image.

[0026] Therefore, image processing operators can be applied to depth maps to improve their quality. In another particularly advantageous embodiment, morphological closure operators are applied to the depth map. For example, such morphological closure operators may include a dilation operation with a kernel of a predetermined size, followed by an erosion operation with a kernel of a predetermined size. Morphological closure operations are used to close any potential holes, ensuring a more consistent and reliable depth representation, especially in stereo depth maps.

[0027] In another particularly advantageous embodiment, the anticipated depth variation includes:

[0028] ● Indicates the sum of scores assigned by the object detector to the total depth variation within the bounding box of an object, and / or

[0029] • The spatial distribution and / or contour of the expected depth variations in the presence of an object.

[0030] These quantities can be easily compared with the corresponding actual depth changes in the depth information, and thus it is easy to determine whether they are consistent, for example, through thresholding.

[0031] In a particularly advantageous embodiment, in response to determining that the proportion of depth changes below a first threshold within the bounding box assigned to an object by the object detector is greater than a second threshold, it is determined that the depth change is expected given the presence of the object. It has been found that the actual presence of an object is associated with the presence of very few local depth changes. Furthermore, since depth changes can only be detected within a certain distance, detection at greater distances has a proportion of small depth changes close to their maximum value and is always maintained. Therefore, filtering primarily affects nearby objects. This is practically significant because closer objects are more relevant to the planning of the next action in the downstream system. This is particularly true for use cases of autonomous driving in vehicles and / or robots.

[0032] In another particularly advantageous embodiment, the method further includes determining a representation of the scene based at least in part on one or more objects whose existence has been verified using depth information. This representation is used by numerous downstream systems to plan the corresponding next action. In particular, this is suitable for use cases in autonomous driving, where the representation is used to plan future trajectories within a predetermined timeframe.

[0033] Therefore, in another particularly advantageous embodiment, an actuation signal is calculated based on a representation of the scene. The actuation signal is used to actuate vehicles, robots, driver assistance systems, quality inspection systems, monitoring systems, and / or medical imaging systems. Because erroneous object detections are filtered out, the probability that the action performed by the corresponding actuated technical system in response to the actuation signal is appropriate in the context represented by the image is increased. In particular, inappropriate actions are not taken in response to erroneous object detections.

[0034] This method can be implemented entirely or partially by a computer and embodied in software. Therefore, the present invention also relates to a computer program having machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method described above. Here, control units for vehicles or robots, and other embedded systems capable of executing machine-readable instructions, are also considered computers. Computing instances include virtual machines, containers, or other execution environments that allow the execution of machine-readable instructions in the cloud.

[0035] Non-transitory storage media and / or downloadable products may include computer programs. A downloadable product is an electronic product that can be sold online and transmitted over a network for immediate fulfillment. One or more computers and / or computing instances may be equipped with the computer programs and / or the non-transitory storage media and / or downloadable products.

[0036] The invention will be described below using accompanying drawings, without any intention to limit the scope of the invention. Attached Figure Description

[0037] Figure 1 An exemplary embodiment of the method 100 for verifying the existence of object 3 in scenario 1;

[0038] Figure 2 An exemplary visualization of the actual geometric cues of object 3 in scene 1 obtained through depth information 5;

[0039] Figure 3 : Image examples that lead to incorrect object detection. Detailed Implementation

[0040] Figure 1 This is a schematic flowchart of an embodiment of method 100, which is used to verify the presence of an object 3 detected in scene 1 based on at least one image 2 of scene 1.

[0041] In step 110, given object detector 4, at least one image region 2a in image 2 is assigned to object 3, image region 2a indicating the presence of object 3 in scene 1.

[0042] In step 120, depth information 5 associated with the image region 2a is obtained. This depth information 5 indicates the distance d between the sensor used to acquire image 2 and the scene region 1a in scene 1 corresponding to image region 2a.

[0043] According to box 121, the operator that highlights the depth change 5a can be applied to the depth information 5, such that the more dramatic the depth change 5a, the more it is highlighted.

[0044] According to box 121a, the operator can be configured to distinguish between a consistent depth gradient of a flat surface and a more significant depth variation protruding from these surfaces.

[0045] According to box 121b, the operator may include the derivative operator and / or the Sobel operator.

[0046] According to box 123, the depth information 5 may include a depth map that assigns a value to each pixel of the image region 2a, the value indicating the distance between the sensor and a position in scene 1 represented by the corresponding pixel.

[0047] According to box 123a, the morphological closure operator can be applied to the depth map.

[0048] According to box 124, depth information 5 may include one or more of the following:

[0049] • Monocular depth estimation from image 2;

[0050] • Depth information from a stereo camera; and / or

[0051] ● Depth information obtained by measuring the distance to the scene using an electromagnetic interrogation radiation beam 5.

[0052] According to box 125, depth information 5 can be obtained from at least one sensor carried by the same vehicle from which image 2 has already been acquired.

[0053] In step 130, it is determined whether the depth information 5 includes a depth change 5a*, which is expected in the presence of an object 3 assigned to image region 2a by object detector 4.

[0054] According to box 131, the expected depth change 5a* may include:

[0055] • The summation score of the total depth variation 5a that the indicator object detector 4 has assigned to the bounding box of object 3, and / or

[0056] ●The spatial distribution and / or contour of the expected depth variation 5a in the presence of object 3.

[0057] According to box 132, it can be determined whether the proportion of depth change 5a below the first threshold in the bounding box assigned to object 3 by object detector 4 is greater than the second threshold. If this is the case (truth value 1), according to box 133, it can be determined that depth change 5a is expected given the presence of object 3.

[0058] In step 140, if the depth information 5 does indeed include the expected depth change (truth value 1), then it is determined that the object detector 4 assigning image region 2a to object 3 is a valid detection 3* of object 3.

[0059] exist Figure 1 In the example shown, in step 150, the representation 1b of scene 1 is determined at least in part based on the existence of one or more objects 3 that have been verified using depth information 5 (i.e., valid detection 3*).

[0060] In step 160, based on representation 1b of scenario 1, the actuation signal 160a is calculated.

[0061] In step 170, the vehicle 50, the driver assistance system 51, the robot 60, the quality inspection system 70, the monitoring system 80 and / or the medical imaging system 90 are actuated using the actuation signal 160a.

[0062] Figure 2This illustrates how depth information 5 can be used to verify the detection of object 3. In all three partial images (a), (b), and (c), the region of interest 2a within the dashed box is shown in the enlarged inset within the solid box.

[0063] Figure 2 Image 2 is a view of road scene 1. Region of interest 2a shows objects that are not protruding from the road surface (as indicated by the attached figure). (shown) road markings, and birds as this prominent object 3. Figure 2 b shows the corresponding Figure 2 The depth information of image 2 shown in figure a is 5. Figure 2 c shows from Figure 2 The depth changes 5a are derived from the depth information 5 shown in b.

[0064] It can be clearly seen that bird 3 produces a corresponding depth change 5a, while road markings No depth change 5a is produced. Therefore, depth change 5a is suitable for matching bird 3 with road markings. Distinguish them. Therefore, it avoids the need for road markings. The object was mistakenly identified as object 3.

[0065] Figure 3 Image 2 shows some examples of road scenes that cause erroneous object detection. The bounding boxes for these erroneous detections are drawn with dashed lines.

[0066] From the Fishyscapes dataset Figure 3 In case a, manhole covers flush with the road surface and graffiti on the road surface cause false detections.

[0067] From the Cityscapes dataset Figure 3 In b, subtle changes in road texture caused by road surface repair lead to false detections.

[0068] From the BDD100K dataset Figure 3 In C, road markings cause error detection.

Claims

1. A method (100) for verifying the presence of an object (3) detected in a scene (1) based on at least one image (2) of a scene (1), comprising the steps of: ●At least one image region (2a) in the image (2) is assigned (110) to the object (3) by a given object detector (4), the image region (2a) indicating the presence of the object (3) in the scene (1); ● Obtain (120) depth information (5) associated with the image region (2a), the depth information (5) indicating the distance between the sensor used to acquire the image (2) and the scene region (1a) in the scene (1) corresponding to the image region (2a); ● Determine whether the depth information (5) includes the expected depth variation (5a*) in the presence of the object (3) assigned to the image region (2a) by the object detector (4); and ●If the determination is positive, then the determination (140) that the image region (2a) is assigned to the object (3) by the object detector (4) is a valid detection (3*) of the object (3).

2. The method (100) according to claim 1, wherein an operator for highlighting depth changes (5a) is applied (121) to the depth information (5) such that the more dramatic the depth changes (5a), the more they are highlighted.

3. The method (100) of claim 2, wherein the operator is configured (121a) to distinguish between a consistent depth gradient of a flat surface and a more significant depth variation protruding from these surfaces.

4. The method (100) according to any one of claims 2 or 3, wherein the operator comprises (121b) a derivative operator and / or a Sobel operator.

5. The method (100) according to any one of claims 1 to 4, wherein the image (2) is divided into pixels and the depth information (5) includes (123) a depth map that assigns a value to each pixel of the image region (2a) that indicates the distance between the sensor and a position in the scene (1) represented by the corresponding pixel.

6. The method (100) according to claim 5 further includes applying (123a) a morphological closure operator to the depth map.

7. The method (100) according to any one of claims 1 to 6, wherein the expected depth change (5a*) comprises (131): • The summation score of the total depth variation (5a) assigned by the object detector (4) to the bounding box of the object (3), and / or • Spatial distribution and / or contour of the expected depth variation (5a) in the presence of object (3).

8. The method (100) according to any one of claims 1 to 7, wherein in response to determining (132) that the proportion of depth change (5a) below a first threshold in the bounding box assigned to object (3) by object detector (4) is greater than a second threshold, it is determined (133) that depth change (5a) is expected in the presence of object (3).

9. The method (100) according to any one of claims 1 to 8, wherein the depth information (5) comprises (124) one or more of the following: ● Monocular depth estimation from image (2); ● Depth information from a stereo camera (5); and / or ● Depth information obtained by measuring the distance to the scene using electromagnetic interrogation radiation beams (5).

10. The method (100) according to any one of claims 1 to 9, wherein depth information (5) is obtained from at least one sensor (125), said at least one sensor being carried by the same vehicle from which the image (2) has already been acquired.

11. The method (100) according to any one of claims 1 to 10, further comprising: The representation (1b) of scene (1) is determined (150) based at least in part on one or more objects (3) that have been verified using depth information (5).

12. The method (100) according to claim 11, further comprising: ●The actuation signal (160a) is determined based on the representation (1b) of scenario (1), and • Actuate (170) a vehicle (50), a driver assistance system (51), a robot (60), a quality inspection system (70), a monitoring system (80), and / or a medical imaging system (90) using an actuation signal (160a).

13. A computer program comprising machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) of any one of claims 1 to 12.

14. A non-transitory machine-readable storage medium and / or downloadable product having the computer program of claim 13.

15. One or more computer and / or computing instances having the computer program of claim 13 and / or the non-transitory machine-readable storage medium and / or downloadable product of claim 14.