Verifying physical presence of object detected in image

The method uses depth information to verify object presence, reducing false detections in autonomous systems by confirming physical objects through depth changes, enhancing system reliability.

JP2025169212APending Publication Date: 2025-11-12ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025074176
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-04-28
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing object detection systems in autonomous vehicles and robots are prone to false positives, leading to unintended evasive maneuvers and potential accidents due to misleading visual cues.

Method used

A method that verifies the presence of detected objects using depth information from sensors like radar or lidar, distinguishing between actual objects and visual cues by analyzing depth changes, and applying operators to enhance depth information for accurate detection.

Benefits of technology

Reduces false object detections by confirming the physical presence of objects, improving the reliability of autonomous systems' actions and preventing unnecessary maneuvers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169212000001_ABST
    Figure 2025169212000001_ABST
Patent Text Reader

Abstract

To provide a method for verifying the presence of objects detected in a scenery based on at least one image of the scenery.SOLUTION: A method (100) includes: a step (110) of associating, by an object detector (4), at least one image region (2a) in an image (2) with an object (3) existing in a scenery (1) by the image region; a step (120) of obtaining depth information (5) relating to the image region, the depth information being indicative of a distance between a sensor used to acquire the image and a scenery region (1a) in the scenery that corresponds to the image region; a step (130) of determining whether the depth information includes depth changes (5a*) that are to be expected when the object associated with the image region by the object detector exists; and a step (140) of determining, if the determination is positive, that the association between the image region and the object by the object detector is a valid detection (3*) of the object.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to image analysis, and in particular to detecting objects in a scene based on captured images of the scene. [Background technology]

[0002] background In order for a vehicle or robot to operate autonomously within a company's premises or even in public road traffic, it is necessary to constantly monitor the surroundings of the vehicle and / or robot. A crucial part of such monitoring is acquiring and analyzing images of such surroundings. It is particularly important to detect objects that the vehicle and / or robot may collide with.

[0003] When detecting objects from visual cues, it is important that all objects that are both physically present in the scene and visible in the image are recognized as objects. However, visual cues can be misleading, and therefore objects may be recognized where none actually exist. Such false detections during automated driving can lead to evasive or emergency braking maneuvers that can take other traffic participants completely by surprise and potentially cause an accident. Summary of the Invention [Problem to be solved by the invention]

[0004] Disclosure of the Invention The present invention provides a method for verifying the presence of an object detected in a scene based on at least one image of the scene. The image may be acquired using any modality, such as a still or video camera, or a thermal camera. The image may be a multimodal image with pixels having pixel values ​​fused from measurements captured using multiple modalities. In particular, the image may be a color image with pixels associated with multiple intensities for the base colors of a given color space. That is, for each pixel, there is one intensity value for each base color. Exemplary color spaces are RGB (red, green, blue) and CMYK (cyan, magenta, yellow, and key=black). [Means for solving the problem]

[0005] During the method, a given object detector associates at least one image region in an image with an object that the image region suggests is present in the scene. That is, the object detector outputs that a region in the image contains an object instance, regardless of whether or not a specific type of object is identified. The image region may, for example, be a bounding box that encloses the object instance plus a portion of the background. However, the image region may also be a more precise outline of the object instance that distinguishes it from the background.

[0006] In the context of the present invention, an object is a physical entity with a three-dimensional structure. In particular, an object may be an entity standing on an otherwise flat surface or an entity protruding from an otherwise flat surface. In the context of autonomous driving, such objects that a vehicle or robot may collide with are predominantly relevant objects. For example, a human or animal standing on a road is a relevant object. However, painted road markings or manhole covers that are substantially flush with the road are not considered to be objects, since a vehicle or robot may simply drive over them.

[0007] Depth information is obtained for the image region, the depth information indicating the distance between the sensor used to capture the image and the scene region in the scene that corresponds to the image region, i.e., the depth information is not limited to this exact distance, but may be a quantity proportional to this distance.

[0008] In particular, the depth information is ·Monocular depth estimation from images, Depth information from a stereo camera, and / or Depth information obtained by measuring the distance to the landscape using a beam of interrogating electromagnetic radiation may include one or more of:

[0009] In particular, in autonomous driving use cases, the vehicle is already equipped with sensors for measuring distance by radar or lidar interrogation radiation, which means that such existing sensors can be reused, thereby reducing the costs of implementing the method in the vehicle.

[0010] Advantageously, therefore, the depth information is obtained from at least one sensor mounted on the same vehicle from which the image was obtained, i.e., this vehicle may be equipped with both the sensor used to obtain the image and the sensor for the depth information.

[0011] It is determined whether the depth information contains depth changes that should be expected if the object associated with the image region by the object detector is present. That is, if an object actually exists in the scene at the location suggested by its detection in the image, then the object will necessarily cause depth changes. This means that if these depth changes are missing from the depth information, then the object cannot be present in the scene, or at least not at the location suggested by the detection of the object instance in the image region.

[0012] Thus, if an expected depth change is determined to exist, then the object detector's association of the image region with the object is determined to be a valid detection of the object. As used herein, determining whether an expected depth change exists may include both a qualitative determination of whether a depth change exists at all and a quantitative determination of the amount of depth change.

[0013] For example, for portions of an image relating to a distant scene or portion thereof, there may be only slight depth variations. For a nearby scene or portion thereof, there may be larger depth variations. For example, a road surface area may show depth variations for all pixels in the region of interest. Thus, an indication of an object may be linked, for example, to the presence of slight local depth variations.

[0014] Depth information has proven to be a particularly advantageous tool for resolving ambiguity as to whether visual cues in an image suggest the presence of actual objects, or whether the visual cues merely comprise texture and / or color variations on an objectless surface. Depth information provides geometric cues that provide a more comprehensive view and complement visual (appearance) cues.

[0015] Particularly in autonomous driving applications, roads often exhibit features that can be mistaken for objects. For example, roads contain many markings, such as lines defining lanes, speed limits, designated driving directions for lanes, indications regarding right-of-way, indications regarding who may use the lanes, and even unofficial markings such as graffiti. There are also manhole covers and other devices that lie flush with the road surface. Furthermore, roads themselves may exhibit texture changes due to changes in the composition of the road surface. For example, a previously constructed road may have been excavated and subsequently filled with new pavement having a different texture, or a depression may have been repaired with temporary asphalt having yet another different texture.

[0016] There are many more sources of potential false object detection in autonomous driving applications. For example, a bus may carry advertisements that show scenes that differ from the actual physical scene, such as a family marveling at their latest car. Neither the family nor the car shown in the advertisement should be recognized as a real object, even if they are presented in a perspective that is reminiscent of real objects. Such false detections can cause erroneous reactions by downstream systems, such as evasion or emergency braking, which are undesirable because they are not anticipated by other traffic participants.

[0017] That is, depth information is particularly well-suited for determining whether visual cues suggesting the presence of an object actually belong to an "object" that is a potential target for collision, or whether the visual cues suggesting the presence of an object relate to something that can be safely ignored for autonomous driving purposes. In particular, not everything that is in some way distinguishable from the road surface is an object. For example, a flush manhole cover, while very different from the road surface, is not an object that a vehicle or robot could collide with.

[0018] Filtering with depth information is robust. Depth information becomes inaccurate at long distances. However, this is not a problem because if there is no depth change, the object will not be filtered out. Therefore, there is no loss of recall in this case. Therefore, this filtering mechanism is most effective for close-distance objects (where false-positive detections would be more dangerous).

[0019] In a particularly advantageous embodiment, a depth change emphasis operator is applied to the depth information, so that the steeper the depth change, the more the depth change is emphasized. This is further used to distinguish between depth changes caused by the presence of an object and steady depth changes caused by the viewpoint between the camera and the ground. Steady depth changes suggest a flat surface, which is a drivable surface in the context of autonomous driving. Furthermore, steep depth changes suggest the presence of an object protruding from the flat surface or standing on the flat surface.

[0020] That is, in a further particularly advantageous embodiment, the operator is configured to distinguish between consistent depth gradients of flat surfaces and more pronounced depth variations protruding from these surfaces.

[0021] Examples of operators that further emphasize depth changes when they are already relatively sharp include derivative operators and Sobel operators. The Sobel operator is a convolution edge detection filter that calculates the first derivative of pixel values ​​while simultaneously smoothing in a direction perpendicular to the direction in which the derivative is calculated. In particular, the result of applying the Sobel operator may include a gradient image that emphasizes edges in the original image.

[0022] In a further particularly advantageous embodiment, the image is divided into a plurality of pixels. The depth information includes a depth map. This depth map associates, with each pixel of the image region, a value indicating the distance between the sensor and the location in the scene represented by that pixel. That is, the value does not have to be exactly this distance, but rather only proportional to this distance. The depth map then adds the concept of a third dimension to the original two-dimensional image. The depth map can be thought of as a depth image corresponding to the original image.

[0023] As a result, image processing operators can be applied to the depth map to improve its quality. In a further particularly advantageous embodiment, a morphological closing operator is applied to the depth map. For example, such a morphological closing operator may include a dilation operation with a kernel of a predetermined size followed by an erosion operation with a kernel of a predetermined size. The morphological closing operation is used to fill any potential holes, thereby ensuring a more consistent and reliable depth representation, especially in stereo depth maps.

[0024] According to a further particularly advantageous embodiment, the expected depth change is a summary score indicating the total amount of depth change within the bounding box that the object detector has associated with the object; and / or The spatial distribution and / or profile of depth changes that should be expected when an object is present Includes.

[0025] These quantities can be easily compared with the respective actual depth changes in the depth information, and therefore it can be easily determined whether they match, for example by thresholding.

[0026] In a particularly advantageous embodiment, a depth change is determined to be expected if an object is present in response to determining that the proportion of depth changes in the bounding box associated with the object by the object detector is below a first threshold and is higher than a second threshold. It has been found that the actual presence of an object is linked to the presence of small local depth changes. Furthermore, since depth changes can only be detected within a certain distance, detections at long distances have a small depth change proportion near its maximum and are always maintained. Therefore, filtering primarily affects nearby objects. This is of practical interest, since closer objects are more relevant for planning the next action in downstream systems. This is particularly true for use cases such as autonomous driving of vehicles and / or robots.

[0027] In a further particularly advantageous embodiment, the method further comprises determining a representation of the scene based at least in part on one or more objects whose presence is verified using the depth information. Such a representation is used by a number of downstream systems to plan their next actions. In particular, this applies to autonomous driving use cases, where the representation is used to plan a future trajectory over a predetermined time period.

[0028] Therefore, in a further particularly advantageous embodiment, an action signal is calculated based on the representation of the scene. A vehicle, a driver assistance system, a robot, a quality inspection system, a surveillance system, and / or a medical imaging system is operated by the action signal. Since false object detections are filtered out, the probability that an action performed by the respective technical system operated in response to the action signal is appropriate in the situation characterized by the image is improved. In particular, inappropriate actions are prevented from being performed in response to false object detections.

[0029] The method may be wholly or partly computer-implemented and may be embodied in software. Accordingly, the present invention also relates to a computer program comprising machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the above-described method. In this specification, control units for vehicles or robots and control units for other embedded systems capable of executing machine-readable instructions should also be considered as computers. Computing instances include virtual machines, containers, or other execution environments that enable the execution of machine-readable instructions in the cloud.

[0030] The non-transitory storage medium and / or downloadable product may include a computer program. A downloadable product is an electronic product that can be sold online for immediate realization and transmitted over a network. One or more computers and / or computing instances may implement such computer program and / or such non-transitory storage medium and / or downloadable product.

[0031] In the following, the invention will be described with the aid of the drawings, without any intention of limiting the scope of the invention. [Brief explanation of the drawings]

[0032] [Figure 1] 1 illustrates an exemplary embodiment of a method 100 for verifying the presence of an object 3 in a scene 1. FIG. [Figure 2] FIG. 1 is an exemplary visualization of using depth information 5 to obtain geometric cues regarding the actual presence of an object 3 in a scene 1. [Figure 3] FIG. 10 is a diagram illustrating an example of an image that causes erroneous object detection. DETAILED DESCRIPTION OF THE INVENTION

[0033] FIG. 1 is a schematic flow chart of an embodiment of a method 100 for verifying the presence of an object 3 detected in a scene 1 based on at least one image 2 of said scene 1 .

[0034] In step 110, a given object detector 4 matches at least one image region 2a in an image 2 with an object 3 that is implied to be present in the scene 1 by that image region 2a.

[0035] In step 120, depth information 5 is obtained for the image region 2a, which indicates the distance d between the sensor used to acquire the image 2 and the scene region 1a in the scene 1 that corresponds to the image region 2a.

[0036] According to block 121, an operator can be applied to the depth information 5 that emphasizes the depth change 5a, so that the steeper the depth change 5a, the more emphasized it is.

[0037] According to block 121a, this operator may be configured to distinguish between consistent depth gradients of flat surfaces and more pronounced depth changes protruding from these surfaces.

[0038] According to block 121b, the operators may include differential operators and / or Sobel operators.

[0039] According to block 123, the depth information 5 may include a depth map that associates, for each pixel of the image region 2a, a value indicating the distance between the sensor and the location in the scene 1 represented by that pixel.

[0040] According to block 123a, a morphological closing operator can be applied to the depth map.

[0041] According to block 124, the depth information 5 is ·Monocular depth estimation from image 2, Depth information from a stereo camera5, and / or Depth information obtained by measuring the distance to the landscape using a beam of interrogating electromagnetic radiation5 may include one or more of:

[0042] According to block 125, the depth information 5 may be obtained from at least one sensor mounted on the same vehicle from which the image 2 was obtained.

[0043] In step 130 it is determined whether the depth information 5 contains a depth change 5a* that is to be expected in the presence of an object 3 associated by the object detector 4 with the image region 2a.

[0044] According to block 131, the expected depth change 5a* is a summary score indicating the total amount of depth change 5a within the bounding box that the object detector 4 has associated with the object 3, and / or the spatial distribution and / or profile of the depth change 5a that should be expected in the presence of the object 3; may include:

[0045] According to block 132, it can be determined whether the proportion of depth changes 5a in the bounding boxes that the object detector 4 has associated with the object 3 that are below a first threshold is higher than a second threshold. If this is the case (truth value 1), then according to block 133 it can be determined that the depth changes 5a are to be expected if the object 3 is present.

[0046] If the depth information 5 contains the expected depth change (truth value 1), then in step 140 it is determined that the association between the image region 2a and the object 3 by the object detector 4 is a valid detection 3* of the object 3.

[0047] In the example shown in Figure 1, in step 150, a representation 1b of the scene 1 is determined based at least in part on one or more objects 3 (i.e., valid detections 3*) whose presence has been verified using depth information 5.

[0048] In step 160, based on the representation 1b of the scene 1, a motion signal 160a is calculated.

[0049] In step 170, the vehicle 50, the driver assistance system 51, the robot 60, the quality inspection system 70, the monitoring system 80, and / or the medical imaging system 90 are operated by the operating signal 160a.

[0050] Figure 2 shows how the detection of an object 3 can be verified using depth information 5. In all three sub-images (a), (b) and (c), the region of interest 2a within the dashed box is shown in the enlarged inset within the solid box.

[0051] FIG. 2(a) is an image 2 of a road scene 1. The region of interest 2a shows a road marking (indicated by reference character ¬3), which is not an object protruding from the road surface, and a bird, which is such a protruding object 3. FIG. 2(b) shows depth information 5 corresponding to image 2 shown in FIG. 2(a). FIG. 2(c) shows depth change 5a derived from depth information 5 shown in FIG. 2(b).

[0052] It can be clearly seen that the bird 3 generates a corresponding depth change 5a, while the road marking ¬3 does not generate a depth change 5a. The depth change 5a is therefore suitable to distinguish between the bird 3 and the road marking ¬3. False detection of the road marking ¬3 as an object 3 can be avoided in this way.

[0053] Figure 3 shows some examples of road scene images 2 that cause false object detections. The bounding boxes for these false detections are shown with dashed lines.

[0054] In Figure 3(a), taken from the Fishyscapes dataset, a manhole cover flush with the road surface and graffiti painted on the road surface cause false positives.

[0055] In Figure 3(b), taken from the Cityscapes dataset, subtle texture changes in the road resulting from road surface repairs cause false positives.

[0056] In Figure 3(c), obtained from the BDD100K dataset, road markings cause false positives.

Claims

1. A method (100) for verifying the presence of an object (3) detected in a scene (1) based on at least one image (2) of said scene (1), comprising: - a step (110) of associating, by a given object detector (4), at least one image region (2 a) in said image (2) with an object (3) suggested by said image region (2 a) to be present in said scene (1); - a step (120) of obtaining depth information (5) relating to said image area (2a), said depth information (5) indicating the distance between the sensor used to obtain said image (2) and a scene area (1a) in said scene (1) corresponding to said image area (2a); - determining (130) whether the depth information (5) contains a depth change (5a*) that should be expected in the presence of the object (3) associated with the image region (2a) by the object detector (4); If the determination is positive, determining (140) that the association of the image region (2 a) with the object (3) by the object detector (4) is a valid detection (3*) of the object (3); A method (100) comprising:

2. - applying (121) an operator to the depth information (5) that emphasizes the depth changes (5a), such that the steeper the depth changes (5a), the more emphasized the depth changes (5a); The method (100) of claim 1.

3. The operator is configured to distinguish between consistent depth gradients of a flat surface and more pronounced depth changes protruding from the surface (121 a); The method (100) of claim 2.

4. the operators include differential operators and / or Sobel operators (121b); The method (100) of claim 2 or 3.

5. The image (2) is divided into a plurality of pixels; the depth information (5) includes a depth map (123); the depth map associating, for each pixel of the image region (2a), a value indicative of the distance between the sensor and the location in the scene (1) represented by said pixel; The method (100) of any one of claims 1 to 4.

6. The method (100) further comprises applying (123a) a morphological closing operator to the depth map. The method (100) of claim 5.

7. The expected depth change (5a*) is a summary score indicating the total amount of depth change (5a) within the bounding box that the object detector (4) has associated with the object (3), and / or the spatial distribution and / or profile of the depth change (5a) that should be expected if said object (3) is present; (131), The method (100) of any one of claims 1 to 6.

8. In response to determining (132) that a proportion of depth changes (5 a) in the bounding boxes associated by the object detector (4) with the object (3) that are below a first threshold is higher than a second threshold, determining (133) that the depth changes (5 a) are to be expected if the object (3) is present. The method (100) of any one of claims 1 to 7.

9. The depth information (5) is - Monocular depth estimation from said image (2); Depth information from a stereo camera (5), and / or Depth information obtained by measuring the distance to the scene using a beam of interrogating electromagnetic radiation (5). (124) including one or more of: The method (100) of any one of claims 1 to 8.

10. The depth information (5) is obtained (125) from at least one sensor mounted on the same vehicle from which the image (2) was obtained. The method (100) of any one of claims 1 to 9.

11. The method (100) comprises: determining (150) a representation (1b) of said scene (1) based at least in part on one or more objects (3) whose presence has been verified using the depth information (5); The method (100) of any one of claims 1 to 10.

12. The method (100) comprises: - determining (160) an action signal (160a) based on said representation (1b) of said scene (1); - operating (170) a vehicle (50), a driver assistance system (51), a robot (60), a quality inspection system (70), a monitoring system (80), and / or a medical imaging system (90) by said operating signal (160a); further comprising: The method (100) of claim 11.

13. 13. A computer program comprising machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) of any one of claims 1 to 12.

14. A non-transitory machine-readable data carrier and / or download product comprising a computer program according to claim 13.

15. One or more computers and / or computing instances comprising a computer program according to claim 13 and / or comprising a non-transitory machine-readable data carrier and / or download product according to claim 14.