Sensor fusion for object avoidance detection
By fusing radar and camera data and utilizing histogram trackers and depth information, the problem of accurate height estimation of stationary objects in the path of vehicles was solved, enabling safe obstacle avoidance decisions in autonomous driving.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APTIV TECHNOLOGIES AG
- Filing Date
- 2022-03-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to accurately estimate the height of stationary objects in a vehicle's path and drivability conditions, and radar or camera methods alone are insufficient to support autonomous or semi-autonomous control.
By fusing radar and camera data, utilizing time-series radar data and image histogram trackers, and combining depth information to estimate height and width, the real-world dimensions of objects are determined, and the speed and direction of vehicles are adjusted based on these estimates to avoid stationary objects.
It enables accurate estimation of the height and width of stationary objects, supporting vehicles to make quick and safe decisions to pass or bypass them, thus improving the reliability of autonomous driving.
Smart Images

Figure CN115144849B_ABST
Abstract
Description
Background Technology
[0001] In some vehicles, sensor fusion systems, or so-called "fusion trackers," combine information from multiple sensors to depict bounding boxes around objects that could obstruct movement. The combined sensor data provides a better estimate of the position of each object within the field of view (FOV) under various conditions. Resizing or repositioning these bounding boxes typically involves using expensive hardware capable of correlating and fusing sensor data quickly enough to support computer-based decision-making for autonomous or semi-autonomous control. Summary of the Invention
[0002] This document describes techniques, apparatus, and systems related to sensor fusion for object avoidance detection. In one example, a method includes: determining radar range detection in a road based on time-series radar data obtained from a vehicle's radar; receiving an image including the road from a camera of the vehicle; determining a histogram tracker based on the radar range detection and the image to identify the size of a stationary object in the road; and determining the height or width of the stationary object relative to the road based on the histogram tracker, the height or width being output from the histogram tracker by applying the radar range to the image-based size of the stationary object. The method further includes: maneuvering the vehicle to avoid or pass drive over the stationary object without collision based on over-drivability conditions derived from the height or width determined according to the histogram tracker.
[0003] In another example, a method includes: receiving a first electromagnetic wave in a first frequency band reflected from an object within the path of a vehicle, receiving an image depicting the object, and adjusting the speed of the vehicle based on the size of the object defined by the first electromagnetic wave and the image.
[0004] In another example, the system includes a processor configured to perform the methods and other methods set forth herein. In yet another example, a system including means for performing the methods and other methods is described. This document also describes a non-transitory computer-readable storage medium having instructions that, when executed, configure the processor to perform the methods summarized above and other methods described herein.
[0005] This invention presents a simplified concept of sensor fusion for object avoidance detection, which is further described in the following detailed description and is illustrated in the accompanying drawings. This invention is not intended to identify essential features of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter. Although primarily described in the context of improving fusion tracker matching algorithms, sensor fusion techniques for object avoidance detection can be applied to other applications that desire to match multiple low-level tracks with high speed and confidence. Attached Figure Description
[0006] The following figures illustrate in detail one or more aspects of sensor fusion for object avoidance detection. The same numbers are used throughout the figures to refer to similar features and components:
[0007] Figure 1-1 An example environment of a vehicle having sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown;
[0008] Figure 1-2 A block diagram of an example vehicle performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown;
[0009] Figure 2-1 An example case of a radio wave sensor of an example vehicle is shown, which performs sensor fusion for object avoidance detection according to one or more implementations of this disclosure;
[0010] Figure 2-2 An example output of a radio wave sensor of an example vehicle is shown, illustrating sensor fusion for object avoidance detection according to one or more implementations of this disclosure.
[0011] Figure 3 Example images captured by the light wave sensor of an example vehicle performing sensor fusion for object avoidance detection, according to one or more implementations of this disclosure;
[0012] Figure 4-1 An example image of an optical sensor of an example vehicle is shown, illustrating the use of a bounding box and sensor fusion for object avoidance detection according to one or more implementations of this disclosure.
[0013] Figure 4-2 Example histograms of images according to one or more implementations of this disclosure are shown;
[0014] Figure 5-1 Example images obtained by the light wave sensor of an example vehicle using a bounding box and performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure are shown.
[0015] Figure 5-2 A convolutional neural network for a vehicle performing sensor fusion for object avoidance detection is shown according to one or more implementations of this disclosure;
[0016] Figure 6 Example methods for sensor fusion for object avoidance detection according to one or more implementations of this disclosure are shown; and
[0017] Figure 7 Another example method for sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. Detailed Implementation
[0018] Overview
[0019] Determining high-level matches between different sensor tracking methods (e.g., radar, vision cameras, lidar) can be challenging. Despite these challenges, sensor fusion systems that can combine multiple perspectives of the environment can be valuable in providing safety. For example, autonomous or semi-autonomous vehicles can benefit from performing accurate estimates of objects on the road. This estimate can then be correlated with control decisions to avoid collisions while driving above or around the object.
[0020] When it comes to certain dimensions, particularly height estimation relative to the ground plane on which a vehicle travels, camera-only or radar-only methods may be insufficient to support autonomous or semi-autonomous control of the vehicle. Radar technology can be very accurate in estimating distances, but typically has poor resolution in determining lateral and longitudinal angles (azimuth and elevation). On the other hand, vision-based camera systems can provide accurate estimates of size based on images, as well as accurate lateral and longitudinal angles (azimuth and elevation). Measuring distances using images can be difficult; poor distance estimation can lead to inaccurate image-based height estimates. Similarly, if radar is used alone, the poor angular resolution from radar can cause it to produce large height estimation errors.
[0021] This document describes techniques, apparatus, and systems for sensor fusion used for object avoidance detection, including stationary object height estimation. The sensor fusion system may include a two-stage pipeline. In the first stage, time-series radar data is passed through a detection model to generate radar distance detections. In the second stage, based on the radar distance detections and camera detections, an estimation model detects over-drivable conditions associated with stationary objects in the vehicle's path. By projecting the radar distance detections onto the pixels of an image, a histogram tracker can be used to determine the pixel-based dimensions of stationary objects and track them across frames. Utilizing depth information, highly accurate pixel-based width and height estimates can be made, and after applying over-drivability thresholds to these estimates, the vehicle can quickly and safely make over-drivability decisions regarding objects in the road.
[0022] One problem overcome by these technologies is estimating the height of stationary objects (such as debris) in a vehicle's path and traversing driving conditions. By fusing camera and radar data, the example sensor fusion system combines the advantages of radar (especially accurate distance detection) with the advantages of camera-based imagery (especially accurate azimuth and elevation determination).
[0023] Example Environment
[0024] The sensors described in this article can be used to measure electromagnetic waves. These sensors can be specifically designed to measure electromagnetic waves within a given wavelength range. As an example, a light receiver can be used as a camera to measure electromagnetic waves within the visible spectrum. As another example, a radio wave receiver can be used as radar to measure electromagnetic waves within the radio wave spectrum.
[0025] For example, consider using radar to sense the distance between objects in the surrounding environment. Radio pulses can be emitted by the radar into the surrounding area. These radio waves can be reflected back from surrounding objects, allowing the determination of the distance and speed of these objects.
[0026] Vehicle safety or control systems can infer not only distance or spacing from objects, but also their dimensions. For example, a vehicle might use the height of an object to determine drive-overs. A vehicle might need the width of an object to determine drive-arounds. Other dimensional or orientation information can provide additional benefits to the vehicle. While radio wave sensors can provide some indication of object size, they may not be a reliable source of this information.
[0027] Consider using a camera or light wave sensor to identify objects in the surrounding environment. Image-based processing techniques can be employed to identify objects in images provided by a light wave sensor. As an example, a histogram of the object and its surroundings can be used to indicate the presence of an object. Neural networks can detect objects in image data provided by a camera. While a camera provides some indication of size, such a light wave sensor may not be a reliable source of this information.
[0028] As discussed above, vehicles can use these and other sensors to provide vehicle autonomy and safety. Sensors can be configured to detect electromagnetic waves or other environmental parameters. Such sensors can individually provide valuable information to the vehicle control system, which, when combined with other sensor information, provides environmental information that could not be directly perceived using a single sensor alone.
[0029] As a simple example, distance information collected by radio wave sensors can be combined with image information collected by light wave sensors to provide environmental information that the sensors cannot directly perceive independently. The physical characteristics of the light wave sensors, along with the distance information, can be used to determine the real-world size of objects. Since pixels within a sensed image from a light wave sensor are assigned values corresponding to the light waves received at a specific location, the number of pixels associated with an object can correspond to the object's real-world size. That is, for a given distance, an object occupying a few pixels in a row across a dimension can correspond to a specific size or dimension (e.g., meters, inches) of an object in front of the vehicle. Therefore, information from multiple sensors can be used to detect objects that affect vehicle autonomy and autonomous vehicle operation.
[0030] One example of this disclosure includes sensor fusion for object detection. In fact, this application, as presented in this disclosure and other examples, enhances our understanding of the environment surrounding a vehicle. These are just a few examples of how the described technologies and devices can be used to provide such environmental information. This document now turns to example operating environments, followed by descriptions of example devices, methods, and systems.
[0031] Figure 1-1 An environment 100 of a vehicle 102 performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. The vehicle 102 can be of various types and implementations. For example, the vehicle 102 can be an autonomous or semi-autonomous car, truck, off-road vehicle, or motorcycle. The vehicle 102 can be semi-autonomous or partially autonomous and includes a safety system capable of autonomously or semi-autonomously controlling motion including the acceleration and direction of the vehicle 102.
[0032] Vehicle 102 is moving in a direction of travel 104. As shown, the direction of travel 104 can be forward along road 106. Road 106 can be associated with pitch 108, which defines the pitch angle offset of vehicle 102 relative to object 110. Road 106 can define the path of vehicle 102 based on the direction of travel 104. The path can be the route taken by the vehicle within road 106.
[0033] Object 110 can be various road obstacles, obstructions, or debris. As an example, object 110 is a child's tricycle. Object 110 is at a distance 112 from vehicle 102. Distance 112 may be unknown to vehicle 102 without the use of radar, lidar, other suitable ranging sensors, or combinations thereof. Distance 112 can be relative to a Cartesian or polar coordinate system. For example, referring to a visual sensor, distance 112 is expressed in units of measurement relative to the x, y, and z coordinates. When inferred from radar, distance 112 is a unit of measurement with amplitude and orientation relative to vehicle 102. For ease of description, unless otherwise stated, distance 112 as used herein is a measurement in a polar coordinate system, which includes amplitude or distance at a specific azimuth and elevation angle from vehicle 102.
[0034] The radar's radio wave transmitter can send radio waves 122 towards object 110. The reflected radio waves 124 can be received by the radar's radio wave receiver 120. Streetlights, vehicle lights, the sun, or other light sources (not shown) can emit light waves reflected by object 110. The reflected light waves 132 can be received by a camera operating as a light wave receiver 130. Although specified as sensors for a specific class of electromagnetic waves, the radio wave receiver 120 and the light wave receiver 130 can be electromagnetic wave sensors of various spectra, including sensor technologies such as ultrasound, lidar, radar, cameras, and infrared.
[0035] Vehicle 102 may include a computer 140 for processing sensor information and communicating with various vehicle controllers. Computer 140 may include one or more processors for calculating sensor data. As an example, computer 140 may operate based on instructions stored on a computer-readable medium 144. Various types of computer-readable media 144 (e.g., random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), non-volatile random access memory (NVRAM), read-only memory (ROM), flash memory) may be used to digitally store data (including sensor data) on computer 140 and provide processing buffers. Data may include operating systems, one or more applications, vehicle data, and multimedia data. Data may include instructions in a computer-readable form, including programs having instructions operable to implement the teachings of this disclosure. Instructions may be any implementation, including field-programmable gate arrays (FPGAs), machine code, assembly code, high-order code (e.g., C), or any combination thereof.
[0036] Instructions may be processor-executable to follow a combination of steps and actions provided in this disclosure. As an example, throughout the discussion of this disclosure, computer 140 may receive radio waves 124 reflected from object 110 within the path of vehicle 102 in a first frequency band. Computer 140 may also receive an image depicting object 110 and adjust the speed of vehicle 102 based on the size of object 110. The size may be defined based on radio waves 124 and the image.
[0037] Figure 1-2 A block diagram of an example vehicle performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. Vehicle 102 includes a camera or light wave receiver 130, a radar or radio wave receiver 120, and a computer 140, which includes a processor 142 and a computer-readable medium 144. Processor 142 executes a speed controller 146 and a steering controller 148 stored in the computer-readable medium 144. The speed controller 146 and steering controller 148 of computer 140 can respectively drive inputs to a speed motor 150 and a steering motor 152.
[0038] A light wave receiver 130 or camera transmits data to a computer 140. The light wave receiver 130 can transmit data to the computer 140 via a Controller Area Network (CAN) bus 151 or another implementation including wired or wireless devices. The data can be image data based on reflected light waves 132. A radio wave receiver 120 or radar can also transmit data to the computer 140 via a CAN bus 151 or another implementation. The data can define objects and their corresponding spacing and speed.
[0039] Computer 140 can execute instructions to calculate dimensional information based on sensor data obtained from light wave receiver 130 and radio wave receiver 120. The dimensional information can be processed by speed controller 146 and steering controller 148 to output drive commands to speed motor 150, steering motor 152 and other components of vehicle 102.
[0040] As an example, if the size associated with object 110 is too large for vehicle 102 to traverse, computer 140, speed controller 146, steering controller 148, or a combination thereof, can use vehicle functions. Vehicle functions can enable vehicle 118 to perform evasive maneuvers. For example, based on size, speed controller 146 can use speed motor 150 to reduce the speed of vehicle 102, or steering controller 148 can use steering motor 152 to steer vehicle 102 to change direction 104 and avoid object 110. Other actions can be taken based on the size associated with object 110, depending on the situation.
[0041] Figure 2-1 An example scenario of a radio wave sensor of an example vehicle performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. For clarity, the environment 100 of vehicle 102 is shown, omitting vehicle 102. Vehicle 102 is traveling on road 106. Emitted radio waves 122 are transmitted from vehicle 102 and reflected by object 110. The reflected radio waves 124 are received by radio wave receiver 120. Radio wave receiver 120 can determine the distance 112 relative to the polar coordinate system to object 110 and the azimuth angle 200 indicating the direction from vehicle 102 to object 110.
[0042] Figure 2-2 An example output of a radio wave sensor of an example vehicle performing sensor fusion for object avoidance detection according to one or more implementations of the present disclosure is shown. The output 202 of a radio wave receiver 120 of a vehicle 102 performing sensor fusion for object detection according to one or more implementations of the present disclosure is shown. Output 202 depicts the distance 112 between the vehicle 102 and the object 110 as an estimated distance 206 according to a time axis 204. As shown, the vehicle 102 is approaching the object 110 and the distance 112 decreases over time. Therefore, the distance 112 can be used to determine the size of the object 110.
[0043] Figure 3Example images captured by an optical sensor of an example vehicle performing sensor fusion for object avoidance detection, according to one or more implementations of this disclosure, are shown. Image 340 captured by an optical receiver 130 of vehicle 102 performing sensor fusion for object detection is shown. As shown, light waves reflected from object 110 are received by optical receiver 130. The distance 112 between object 110 and vehicle 102 is shown as defined by the distance between the lens of object 110 and the lens of optical receiver 130. This distance 112 may be the same as the distance 112 described by radio receiver 120, or an offset depending on how radio receiver 120 and optical receiver 130 are mounted on vehicle 102. The reflected image 320 may be at the same height as object 110 and along an axis 300 perpendicular to the lens. The lens may be an aperture or an optical tool. Along the optical axis, an optical origin 308 or center may be defined according to the lens associated with optical receiver 130. In the example, the light ray 302 of the object is reflected on axis 300 to form the reflected image 320.
[0044] For simplicity, the focal point and associated focal length are not shown. The effective focal length 304 (or spacing) between the lens axis 306 and the reflected image 320 on the image plane is shown. The ray 302 of the reflected light wave 132 passes through the optical origin 308 to reach the corresponding position on the reflected image 320. The y-direction 310, x-direction 312, and z-direction 314 depict the relative position of the object 110 and the reflected image 320 in a Cartesian coordinate system. As shown, the reflected image can be translated into a translated image 330 by camera constants 332 and 334, where camera constant 334 is in the horizontal direction and camera constant 332 is in the vertical direction. Adjusting these camera constants 332 and 334 will give us the received image 340. As shown, the relative position of the object 110 within the received image 340 can be defined in a repeatable manner based on the spacing 112 and the azimuth angle 200, and based on the lens and hardware architecture and hardware of the light wave receiver 130.
[0045] Figure 4-1 An example image of an optical sensor of an example vehicle performing sensor fusion for object avoidance detection using a bounding box, according to one or more implementations of this disclosure, is shown. Image 340 is an example generated by an optical sensor 130 of a vehicle 102 performing sensor fusion for object detection using a bounding box, according to one or more implementations of this disclosure. The received image 340 can be defined by pixels 400, thereby forming an array as the received image 340. For simplicity and clarity, some pixels 400 are omitted.
[0046] Computer 140 can draw a segmentation bounding box 401 (drawn around object 110) and a background bounding box 404 within object bounding box 402, which, when drawn around object 110, specify the region exactly around and outside of segmentation bounding box 401 and object bounding box 402. Segmentation bounding box 401, object bounding box 402, and background bounding box 404 can be located to an estimated position 422 based on estimated spacing 206 and azimuth angle 200. In practice, segmentation bounding box 401, object bounding box 402, and background bounding box 404 can be represented solely by vertices or other modeling methods. The estimated position can be defined relative to the origin 424 of the received image 340. Origin 424 can be defined relative to the intersection of the x-direction 312, y-direction 310, and z-direction 314.
[0047] As an example, a longer estimation interval can place the segmentation bounding box 401, object bounding box 402, and background bounding box 404 at lower pixel positions on the received image 340. A shorter estimation interval can place the segmentation bounding box 401, object bounding box 402, and background bounding box 404 at higher pixel positions on the received image 340.
[0048] Object bounding box 402 can generally define object 110, and segmentation bounding box 401 can specifically define the size of object 110. As an example, object 110 can be defined as having a pixel height 406 of six pixels and a pixel width 408 of ten pixels, and is segmented as defined by segmentation bounding box 401. As shown, the maximum pixel height 428 and the minimum pixel height 426 are defined based on the segmentation of object 110 within segmentation bounding box 401.
[0049] Computer 140 can use the maximum pixel height 428 and the minimum pixel height 426 to define the real-world height of object 110. Therefore, computer 140 can send this information to speed controller 146 and steering controller 148 to adjust the acceleration or orientation of vehicle 102.
[0050] Figure 4-2 Example histograms of images according to one or more implementations of this disclosure are shown. Example histogram representation 410 of a received image 340 according to one or more implementations of this disclosure is shown. The histogram representation is defined by the frequency 414 of color bins 412 within the object bounding box 402 and the background bounding box 404.
[0051] Computer 140 can output a background bounding box 404 corresponding to histogram 416. Histogram 416 defines the baseline color frequency at a specific location within the received image 340 based on the background bounding box 404. Similarly, histogram 418 defines the color frequency at a specific location within the background bounding box 404 based on the object bounding box 402.
[0052] A comparison between histograms 416 and 418 can be performed to identify the histogram portion 420 associated with object 110. A threshold can be set to define an acceptable range for histogram portion 420, and the size of object bounding box 402 can be continuously adjusted to maintain histogram portion 420 or its ratio, and the size of segmentation bounding box 401 within object bounding box 402 is designed based on object 110. Thus, pixel height 406 and pixel width 408 can be defined based on segmentation bounding box 401. In this way, vehicle 102 with light wave receiver 130, radio wave receiver 120, and computer 140 is fused to determine the size of object 110. As an example, computer-readable medium 144 and processor 142 can respectively store and execute instructions to operate speed controller 146 and steering controller 148 based on size.
[0053] Figure 5-1 Example images obtained by an optical sensor of an example vehicle are shown, using bounding boxes and performing sensor fusion for object avoidance detection according to one or more implementations of this disclosure. Image 340 illustrates an example case using a segmented bounding box 401 (which is a segmentation of an object bounding box 402 (not shown)) and an optical receiver 130 of a vehicle 102 for performing sensor fusion for object detection. In this example, a convolutional neural network (CNN) can be used to determine pixel height 406 and pixel width 408.
[0054] The CNN is configured to detect object 110 within the received image 340. CNN 500 can be used to determine the maximum pixel height 528 and the minimum pixel height 526. The dimensions can be based on a segmentation bounding box 401, which can be defined by CNN 500 or defined as the absolute maximum and minimum pixel values associated with the object in the semantic object recognition implementation.
[0055] Figure 5-2 The CNN 500 is described in more detail. The CNN 500 receives a received image 340 as input from the light wave receiver 130 of the vehicle 102. The CNN 500 can receive the received image 340 as a byte array, where each byte relates to the value (e.g., color) of a corresponding pixel in the received image 340.
[0056] CNN 500 may include layers 502, 504, and 506 leading to segmented bounding boxes 401. Layers can be added or removed in CNN 500 to improve the identification of segmented bounding boxes 401. CNN 500 may provide additional methods, or methods in conjunction with sensor fusion of radio wave receiver 120 and light wave receiver 130, for defining the segmentation of object bounding boxes 402, thereby obtaining the dimensions of segmented bounding boxes 401 and the corresponding dimensions of object 110.
[0057] Figure 6 An example method 600 for sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. Figure 6 The process illustrated includes a set of boxes that specify the operations and steps to be performed, but is not necessarily limited to the order or combination of operations shown for performing the respective boxes. Furthermore, any operation among one or more operations can be repeated, combined, reorganized, omitted, or linked to provide a wide array of additional and / or alternative methods. In the various sections of the following discussion, reference may be made to the examples in the foregoing figures, which are for illustrative purposes only. This technique is not limited to being performed by one or more entities operating on a single device.
[0058] In block 602, (e.g., using radar) radio waves 122 reflected from object 110 within the vehicle path and distributed in a first frequency band are received. The first frequency band may be a radio wave band. As an example, computer 140 or a portion thereof may perform one or more steps of method 600. As an example, the first frequency band may be between a few hertz and 300 gigahertz or have a wavelength from 100,000 km to 1.0 mm.
[0059] In block 604, light waves 132 reflected from object 110 within a second frequency band are received (e.g., using a camera). The second frequency band may be the visible light band. As an example, computer 140 or a portion thereof may perform one or more steps of method 600. As an example, the second frequency band may be a wavelength between a few terahertz and a few petahertz, or between 100 μm and a few nanometers. A sensor may be used to detect the reflected electromagnetic waves. The sensor may include processing and other implementations to determine the frequency of the received wave and convert such wave into quantitative information, images, other representations, and other combinations thereof.
[0060] In box 606, parameters of the vehicle 102 can be adjusted. These parameters can be the speed or direction of the vehicle 102, controlled by the speed motor 150 or the steering motor 152. For example, the steering parameter can be changed to allow the vehicle 102 to maneuver around the object 110. The computer 140 can change the speed parameter to decelerate or accelerate the vehicle, thereby avoiding the object 110. Other parameters of the vehicle 102 can also be adjusted.
[0061] The parameters of vehicle 102 can be adjusted according to the size of object 110. For example, the size of object 110 can be either its height or its width. The size of object 110 can be defined based on the number of pixels 400 in the received image 340.
[0062] Radar detections, defined in a polar coordinate system as distance, azimuth, and elevation relative to vehicle 102, can be obtained and converted to a Cartesian coordinate system used by the camera, such as the X, Y, and Z coordinate system, where X is the lateral or horizontal distance from vehicle 102, Y is the vertical distance relative to the ground plane of vehicle 102, and Z is the longitudinal or depth distance from vehicle 102. Based on Equation 1 and the estimated position 422 of the defined pixel coordinates (u, v), one or more of the segmentation bounding box 401, object bounding box 402, and background bounding box 404 can be placed on the received image 340.
[0063]
[0064] Where u is the horizontal component of the pixel coordinate, f i It is the x-direction 312 component of the effective focal length of 304, and o i The camera constant is 334. This assumes that the radar estimated range (e.g., spacing 112) is r and the azimuth angle (e.g., azimuth angle 200) is θ, then X is equal to a rotation and translation of r*sinθ and Z is equal to a rotation and translation of r*cosθ. The object bounding box 402 and the background bounding box 404 can be placed on the received image 340 based on Equation 2.
[0065]
[0066] Where v is another component of the pixel coordinates in the vertical direction; when estimating a spacing of 206, f j It is the y-direction 310 component of the effective focal length of 304; and o j It is the vertical camera constant 332. One or more of the components u or v can be adjusted to compensate for pitch 108, and the compensation may include an estimated spacing 112 in combination with pitch 108.
[0067] The estimated positions 422 of the segmentation bounding box 401, the object bounding box 402, and the background bounding box 404 are now known, and the size of the object 110 is determined. The object 110 can be detected based on the histogram portion 420 with the corresponding target point on the received image 340. Thus, the size of the object bounding box 402 can be adjusted based on the spacing between the boundary of the segmentation box 401 and the object bounding box 402. The size can be estimated based on Equations 3 and 4.
[0068]
[0069] On top, O w It is a width estimate of the actual width of object 110, and w p This refers to the pixel width 408 of the segmentation bounding box 401. The pixel width 408 can be based on the segmentation bounding box 401 or on the object 110 detected in the received image 340. As an example, a CNN 500 or a segmentation algorithm can be used to determine the actual pixels associated with object 110. In this case, the pixel width 408 can be defined as the maximum and minimum pixel coordinates of object 110 on the horizontal or x-axis. For example, object 110 could have a maximum pixel at position (2, 300) and a minimum pixel at position (5, 600), indicating that the object's maximum pixel width is 5 pixels, even if the vertical coordinates are not equal.
[0070] Continue, r w It could be a light wave receiver 130 with a receiver height of 154, f i It can be the x-direction 312 component of the effective focal length 304, and i w This is the total image width in pixels. As an example, i w It can be 2 pixels, where object 110 has a height of 300 pixels. Therefore, the width estimate O w It is the dimension associated with object 110, which computer 140 can then use to operate vehicle 102 autonomously.
[0071]
[0072] On top, O h It is a height estimate of the actual height of object 110, and h pThis is the pixel height 406 of the segmentation bounding box 401. The pixel height 406 can be based on the segmentation bounding box 401 or on the object 110 detected in the received image 340. As an example, a CNN 500 or a segmentation algorithm can be used to determine the actual pixels associated with object 110. In this case, the pixel height 406 can be defined as the maximum and minimum pixel coordinates of object 110 on the vertical or height axis. For example, object 110 may have a maximum pixel at position (40, 300) and a minimum pixel at position (150, 310), indicating that the maximum pixel height of the object is 300 pixels, even if the horizontal coordinates are not equal. Continuing, f j It can be the y-direction 310 component of the effective focal length 304, and i h This is the total image height in pixels. As an example, i h It can be 600 pixels, where object 110 has a height of 50 pixels. Therefore, the height estimate O h It is the dimension associated with object 110, which computer 140 can then use to operate vehicle 102 autonomously.
[0073] As cited herein, adjectives, including first, second, and others, are used only to provide clarity and specification of elements. As an example, first row 210 and second row 220 are interchangeable and used only for clarity when referring to the current accompanying drawing.
[0074] Figure 7 Another example method 700 for sensor fusion for object avoidance detection according to one or more implementations of this disclosure is shown. Figure 7 The process illustrated includes a set of boxes that specify the operations and steps to be performed, but is not necessarily limited to the order or combination of operations shown for performing the respective boxes. Furthermore, any operation among one or more operations can be repeated, combined, reorganized, omitted, or linked to provide a wide array of additional and / or alternative methods. In the various sections of the following discussion, reference may be made to the examples in the foregoing figures, which are for illustrative purposes only. This technique is not limited to being performed by one or more entities operating on a single device.
[0075] At point 702, radar distance detection in the road is determined based on time-series radar data obtained from the radar of the vehicle. For example, during the first stage of the sensor fusion pipeline, radar distance is determined by receiving radar data from the radar of vehicle 102.
[0076] At point 704, an image including the road is received from the vehicle's camera. For example, during the second stage of the sensor fusion pipeline, an image including a two-dimensional or three-dimensional pixel grid is received from the vehicle's camera 102.
[0077] At point 706, a histogram tracker is determined based on radar distance detection and imagery to identify the size of stationary objects in the road. For example, outputs from two stages of an assembly line are combined to form the histogram tracker. The computer 140 of vehicle 102 executes the histogram tracker, which generates histogram comparisons to correlate the radar distance detection with objects in the image data.
[0078] At 708, the height or width of a stationary object relative to the road is determined from a histogram tracker, which can apply radar distance as depth information to the image-based dimensions of the stationary object. For example, as indicated above, computer 140 determines the relative height above the ground where vehicle 102 is traveling at a measured radar distance, given a height expressed in pixels.
[0079] At 710, the vehicle maneuvers based on overtaking drivability conditions derived from the height or width determined by the histogram tracker to avoid or pass over a stationary object without collision. For example, an estimated height based on radar distance and camera data can indicate that the object is too high for the vehicle 102 to drive over. Similarly, an estimated width based on radar distance and camera data can indicate that the object is too wide for the vehicle 102 to avoid without changing lanes or maneuvering around the object.
[0080] Determining high-level matches between different sensor tracking methods (e.g., radar, vision cameras, lidar) can be challenging. For certain sizes, camera-only or radar-only methods may be insufficient to support autonomous or semi-autonomous vehicle control for height estimation relative to the ground plane in which the vehicle travels. By projecting radar distance detection onto the pixels of an image, histogram trackers, as described in this paper, can be used to discern pixel-based dimensions of stationary objects and track them across frames. Utilizing depth information applied to the image, highly accurate pixel-based width and height estimations can be made, enabling the vehicle to make rapid and safe overtaking drivability decisions regarding objects in the road after applying a drivability threshold. These techniques overcome problems that other sensor fusion systems or radar-only or camera-only systems may have when estimating the height of stationary objects (such as debris) and overtaking drivability conditions in a vehicle's path.
[0081] Example
[0082] Example 1. A method comprising receiving a first electromagnetic wave in a first frequency band reflected from an object within the path of a vehicle. The method further comprises receiving an image depicting the object. The method further comprises adjusting the speed of the vehicle based on the size of the object defined by the first electromagnetic wave and the image.
[0083] Example 2. A method in any of these examples further includes: determining a size based on a pixel array configured to form an image. The size is further based on the position of an object within the pixel array according to an estimated spacing. The estimated spacing is defined based on a first electromagnetic wave.
[0084] Example 3. A method in any of these examples, wherein the size is defined based on a pixel array configured to form an image, the size being further based on the position of an object within the pixel array according to an estimated spacing, the estimated spacing being defined based on a first electromagnetic wave.
[0085] Example 4. A method in any of these examples, where the position is defined by pixel coordinates relative to the origin of the pixel array.
[0086] Example 5. A method in any of these examples, wherein the pixel coordinates include a first component based on the estimated spacing and an angle defined according to the first electromagnetic wave and the direction of travel of the vehicle.
[0087] Example 6. A method in any of these examples, wherein the first component is further based on the product of the following: the focal length defined by the second receiver of the image, and a first ratio between the estimated spacing and the angle-based azimuth spacing.
[0088] Example 7. A method in any of these examples where the first component is offset by a first camera constant, which is defined relative to the second receiver and the origin.
[0089] Example 8. A method in any of these examples, wherein the pixel coordinates include a second component based on the estimated spacing and the receiver height of a second receiver of the image defined relative to a first receiver of the first electromagnetic wave.
[0090] Example 9. A method in any of these examples, wherein the second component is further based on the product of the following: the focal length defined by the second receiver, and a second ratio between the estimated spacing and the receiver height.
[0091] Example 10. A method in any of these examples where the second component is offset by a second camera constant, which is defined relative to the second receiver and the origin.
[0092] Example 11. Any of these examples may further include: a vehicle-based pitch adjustment second component.
[0093] Example 12. A method for any of these examples, wherein the size is based on a ratio of a first product to a second product, the first product including an estimated spacing, a pixel count between the maximum and minimum pixel coordinates of the object on the axis of the pixel array, and a receiver height of a second receiver of the image defined by a first receiver of the first electromagnetic wave, the second product including a focal length defined by the second receiver and the maximum pixel position associated with the axis.
[0094] Example 13. The method of any of these examples further includes: recognizing an object using a convolutional neural network configured to receive an image as input; locating the object based on an estimated spacing and an azimuth angle defined according to a first electromagnetic wave; and defining maximum and minimum pixel coordinates according to the convolutional neural network.
[0095] Example 14. The method of any of these examples further includes: identifying an object by comparing a first histogram of the pixel array with a second histogram of the pixel array at a position defined by the estimated spacing and an azimuth angle defined according to a first electromagnetic wave; and defining a maximum pixel coordinate and a minimum pixel coordinate based on the first histogram and the second histogram.
[0096] Example 15. A method in any of these examples where a first histogram is compared with a second histogram to define a target level to associate a bounding box with an object.
[0097] Example 16. A method for any of these examples, where the axis is the vertical axis.
[0098] Example 17. A method for any of these examples, where the first frequency band is a radio frequency band.
[0099] Example 18. A method in any of these examples, where the size is the height of the object.
[0100] Example 19. A method in any of these examples, where the size is the maximum vertical spacing of the objects.
[0101] Example 20. A method, either alone or in combination with any of these examples, comprising: determining radar distance detection in a road based on time-series radar data obtained from a vehicle's radar; receiving an image including the road from a camera of the vehicle; determining a histogram tracker based on the radar distance detection and the image for identifying the size of a stationary object in the road; determining, according to the histogram tracker, the height or width of the stationary object relative to the road, the height or width being output from the histogram tracker by applying radar distance to an image-based size of the stationary object; and maneuvering the vehicle to avoid or drive over the stationary object without collision based on a driveability condition derived from the height or width determined according to the histogram tracker.
[0102] Example 21. A method for any of these examples, wherein the image-based dimensions include a height or width defined by the pixel array forming the image, the height or width being further based on the position of the stationary object within the pixel array and defined by an estimated distance from the vehicle to the stationary object, the estimated distance being defined by radar distance detection.
[0103] Example 22. A computer-readable medium including instructions that, when executed by a processor, configure the processor to perform any of these methods.
[0104] Example 23. A system comprising means for performing any of these methods.
[0105] Conclusion
[0106] While various embodiments of the present disclosure have been described in the foregoing description and illustrated in the accompanying drawings, it should be understood that the present disclosure is not limited thereto, but can be practiced in various ways within the scope of the following claims. It will be apparent from the foregoing description that various modifications can be made without departing from the spirit and scope of the present disclosure as defined by the following claims.
[0107] Unless the context clearly specifies otherwise, the use of "or" and grammatically related terms indicates an unrestricted, non-exclusive alternative. As used herein, the phrase referring to "at least one" of a list of items means any combination of those items, including a single member. For example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc, as well as any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).
Claims
1. A method executed on at least one processor of a vehicle, the method comprising: Radar data is obtained from the radar of the vehicle, the radar data including distance detection in the vicinity of the ground plane including the road; Based on the distance detection and the direction of travel on the road, a stationary object is determined to be at a radar distance from the vehicle and at an azimuth angle relative to the vehicle in the path of the vehicle, the azimuth angle indicating the direction of the stationary object relative to the vehicle; Images of the road and the stationary object are received from a camera focused on the lens axis by forming a pixel array of pixels. The images have a total image height and a total image width, as well as a focal length between the lens axis and the reflection of the image on an image plane, which has an x-direction component on the horizontal axis and a y-direction component on the vertical axis. Based on the distance detection and the image, a histogram tracker is determined to identify the size of the stationary object, and the number of pixels in the pixel array associated with the stationary object is identified. Determine the pixel height of the stationary object, the pixel height indicating the number of pixels in the pixel array associated with the stationary object on the vertical axis within the image; Determine the pixel width of the stationary object, the pixel width indicating the number of pixels in the pixel array associated with the stationary object on the horizontal axis within the image; The actual height of the stationary object relative to the road is estimated based on the pixel height of the stationary object, the total image height of the image, the y-direction component of the focal length, the azimuth angle, and the radar distance. The actual width of the stationary object is estimated based on the pixel width of the stationary object, the total image width of the image, the x-direction component of the focal length, the azimuth angle, and the radar distance. as well as The vehicle is maneuvered based on the estimated actual height and estimated actual width of the stationary object to drive over or around at least one of the stationary objects without collision.
2. The method of claim 1, further comprising: The speed or direction of travel is adjusted based on the estimated actual height of the stationary object relative to the road to avoid collisions.
3. The method according to claim 1 or claim 2, characterized in that, The determination of the pixel height and pixel width of the static object is performed using a convolutional neural network.
4. A system for a vehicle, the system comprising a processor communicating with a vehicle radar and a vehicle camera, the processor being configured to perform the method of any one of claims 1-3.
5. The system as described in claim 4, characterized in that, The position is defined by the pixel coordinates relative to the origin of the pixel array.
6. The system as described in claim 5, characterized in that, The vehicle radar and the vehicle camera include a forward view from the vehicle.
7. A computer-readable medium comprising instructions that, when executed by a processor, configure the processor to: Perform the method according to any one of claims 1-3.
Citation Information
Patent Citations
Radar guided vision system for vehicle validation and vehicle motion characterization
US20090067675A1
Object recognition system
US6590521B1