Object detection system based on an ultra-wideband sensor network and a non-stereo camera system
The integration of a non-stereo camera system with a UWB sensor network and Bayesian filtering for vehicle object detection addresses inaccuracies in depth estimation and complexity, achieving precise and efficient object localization with reduced power usage.
Patent Information
- Application Number
- DE102024102345
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-01-28
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2044-01-28
AI Technical Summary
Current object detection systems for vehicles face challenges with non-stereo camera systems providing inaccurate depth estimates and UWB sensor networks requiring target objects to be equipped with sensors, leading to limited detection range and high computational complexity.
An object detection system that combines a non-stereo camera system with a UWB sensor network, using a Bayesian filter to merge camera-based and UWB-based positions, and incorporates an initial calibration procedure to account for camera lens distortion, employing rotated bounding box algorithms and Kalman filters for precise estimation.
The system provides accurate and efficient object detection with reduced computational complexity and power consumption, improving depth and angle estimation by calibrating camera depth using UWB-derived real depth, and integrating sensor data effectively.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
introduction
[0001] The present disclosure relates to an object detection system for a vehicle that fuses data from an ultra-wideband (UWB) sensor network and a non-stereo camera system to estimate a location of a target object located in an environment around the vehicle.
[0002] An autonomous vehicle performs various tasks, including, but not limited to, perception, localization, mapping, path planning, decision-making, and motion control. For example, an autonomous vehicle may incorporate perception sensors such as a stereo camera system, LiDAR, and radar to collect perception data regarding the vehicle's surroundings. The perception data collected by the stereo camera system, LiDAR, and radar sensors can be used for object detection and range measurement.
[0003] It should be acknowledged that the current approach to performing object detection and ranging based on perception data collected by LiDAR sensors is complex and computationally intensive, and radar sensors tend to consume relatively large amounts of power. One approach to reduce complexity and computational requirements could involve replacing LiDAR and radar sensors with an ultra-wideband (UWB) sensor network. However, UWB sensor networks also have drawbacks. Specifically, UWB sensor networks offer only a limited detection range and require the target object to be equipped with a UWB sensor. Furthermore, many vehicles are equipped with only a single or non-stereo camera system, which may not be able to provide an accurate depth estimate of the target object.
[0004] Thus, although current object detection systems serve their intended purpose, there is a need in the art for an improved approach to object detection and ranging based on image data acquired by a non-stereo camera system.
[0005] Mingyang, Guan; Sagar, Krishna; Ning, Xu: GPS-Denied UAV-UGV Relative Positioning System via Vision-UWB Fusion. In: 2023 IEEE 18th Conference on Industrial Electronics and Applications (ICIEA). IEEE, publication date in IEEE Xplore 11.09.2023. pp. 227-232 discloses a device for localizing a drone using an unmanned vehicle. The drone is localized in two steps: performing a state estimation using UWB signals and image information, and determining the final position of the drone by merging the state estimates based on the UWB signals and image information. Zeng, Qingxi; Liu, Dehui; Lv, Chade: UWB / binocular VO fusion algorithm based on adaptive kalman filter. Sensors, 2019, Vol. 19, No. 18, p. 4044, discloses a method for determining the position of a vehicle based on visual odometry data and UWB sensor data. Barker, Allen L.; Brown, Donald E.; Martin, Worthy N.: Bayesian estimation and the Kalman filter. Computers & Mathematics with Applications, 1995, Vol. 30, No. 10, pp. 55–77, discloses a Bayesian derivation of a state estimation result for discrete-time Markov process models. Kumar, Ashutosh, et al.: Citywide reconstruction of traffic flow using the vehicle-mounted moving camera in the CARLA driving simulator. In: 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2022. pp. 2292–2299 discloses a method for estimating vehicle traffic flow based on vehicle-mounted cameras. Summary
[0006] According to several aspects, an object detection system for a vehicle is disclosed that estimates a position of a target object located in an environment surrounding the vehicle. The object detection system includes an ultra-wideband (UWB) sensor network including three or more vehicle-mounted anchors in wireless communication with a tag mounted on the target object, each anchor transmitting and receiving sensor signals indicating real-time distances between each anchor and the tag. The object detection system includes a non-stereo camera system that acquires image data representing the target object located in the environment surrounding the vehicle, and one or more controllers in electronic communication with the UWB sensor network and the non-stereo camera system.The one or more controllers include one or more processors that execute instructions to estimate a camera-based position of the target object based on the image data, wherein the camera-based position is adjusted to account for a calibrated estimated camera depth determined during an initial calibration procedure. The one or more controllers estimate a UWB-based position of the target object by executing one or more range-based localization algorithms that analyze the sensor signals. The one or more controllers fuse the camera-based position of the target object and the UWB-based position of the target object together using a Bayesian filter to estimate the position of the target object.The calibrated estimated camera depth represents a coarse camera depth of the target object determined based on the image data acquired by the non-stereo camera system, which is calibrated based on a real depth of the target object, where the real depth of the target object is determined based on the sensor signals of the UWB sensor network.
[0007] In another aspect, the Bayesian filter is a Kalman filter.
[0008] In yet another aspect, a process model of the Kalman filter predicts a plurality of state vectors of the vehicle.
[0009] In one aspect, a measurement model of the Kalman filter performs an update of the plurality of state vectors of the vehicle determined by the process model based on an observation vector.
[0010] In one aspect, the initial calibration procedure includes executing one or more rotated object detection algorithms that determine a rotated bounding box that identifies the target object located within a corresponding image frame of the image data.
[0011] In another aspect, the rotated object detection algorithm is a Darknet 53 “You only look once” (YOLO) algorithm with a recurrent neural network (RNN).
[0012] In yet another aspect, the initial calibration procedure comprises determining a plurality of position parameters of the rotated bounding box, wherein the position parameters of the rotated bounding box include an x-axis position coordinate, a y-axis pixel coordinate, a width of the rotated bounding box, a height of the rotated bounding box, and an angular orientation of the rotated bounding box with respect to the horizontal axis of the corresponding image frame.
[0013] In one aspect, the initial calibration procedure includes determining a coarse camera depth based on: dcam=fcam*hrealhb, where d cam represents the coarse camera depth, h b the height of the rotated bounding box h b represents, f cam represents a focal length of a camera that is part of the non-stereo camera system, and d realrepresents a real depth of the target object, which is determined based on the sensor signals received from the three or more anchors.
[0014] In another aspect, a relationship between the coarse camera depth and the center of the corresponding image is expressed by a straight line equation: y=βx+γ, where β represents a gradient of the line and γ represents a y-axis intercept of the line.
[0015] In yet another aspect, the initial calibration procedure comprises solving a linear regression model representing a relationship between the coarse camera depth, the real depth of the target object, an error introduced by image distortion of the camera, the gradient, the y-axis intercept, and a center of the corresponding image frame, and wherein the linear regression model is expressed as: e=dcam−dreal=βdcen+γ, where e represents the error and d cen represents the center of the corresponding image.
[0016] In one aspect, the initial calibration procedure includes solving for the calibrated estimated camera depth based on a difference between the coarse camera depth and the error, expressed as: Dcali=dcam−e, where D cali represents the calibrated estimated camera depth.
[0017] In another aspect, the calibrated estimated camera depth takes into account an error introduced by a lens distortion of a camera that is part of the non-stereo camera system, and wherein the error is linearly related to a center of a corresponding image frame of the image data.
[0018] In yet another aspect, the one or more processors of the one or more controllers execute instructions to create a vector map based on an attractive field strength between the vehicle and a target position, repulsive field strengths between the vehicle and the target object, and the vehicle and one or more other objects located in the environment, and a total field strength at a current position of the vehicle.
[0019] In one aspect, the target position represents a destination location of the vehicle.
[0020] In another aspect, an object detection system for a vehicle is disclosed that estimates a position of a target object located in an environment surrounding the vehicle. The object detection system includes a UWB sensor network including three or more vehicle-mounted anchors in wireless communication with a tag mounted on the target object, each anchor transmitting and receiving sensor signals indicating real-time distances between each anchor and the tag. The object detection system includes a non-stereo camera system including a camera that captures image data representing the target object located in the environment surrounding the vehicle and one or more controllers in electronic communication with the UWB sensor network and the non-stereo camera system.The one or more controllers include one or more processors that execute instructions to estimate a camera-based position of the target object based on the image data, wherein the camera-based position is adjusted to account for a calibrated estimated camera depth determined during an initial calibration procedure, and wherein the calibrated estimated camera depth accounts for an error introduced by lens distortion of the camera, and the error is linearly related to a center of a corresponding image frame of the image data.The one or more controllers estimate a UWB-based position of the target object by executing one or more range-based localization algorithms that analyze the sensor signals, and fuse the camera-based position of the target object and the UWB-based position of the target object together using a Kalman filter to estimate the position of the target object.
[0021] In another aspect, the initial calibration procedure includes executing one or more rotated object detection algorithms that determine a rotated bounding box defining the target object located within a corresponding image frame of the image data.
[0022] In yet another aspect, the rotated object detection algorithm is a Darknet 53 “You only look once” (YOLO) algorithm with a recurrent neural network (RNN).
[0023] In one aspect, an object detection system for a vehicle is disclosed that estimates a position of a target object located in an environment surrounding the vehicle. The object detection system includes a UWB sensor network including three or more vehicle-mounted anchors in wireless communication with a tag mounted on the target object, each anchor transmitting and receiving sensor signals indicating real-time distances between each anchor and the tag. The object detection system includes a non-stereo camera system including a camera that captures image data representing the target object located in the environment surrounding the vehicle and one or more controllers in electronic communication with the UWB sensor network and the non-stereo camera system.The one or more controllers comprise one or more processors that execute instructions to estimate a camera-based position of the target object based on the image data, wherein the camera-based position is adjusted to account for a calibrated estimated camera depth determined during an initial calibration procedure, and wherein the calibrated estimated camera depth accounts for an error introduced by lens distortion of the camera, and the error is linearly related to a center of a corresponding image frame of the image data, and wherein the initial calibration procedure comprises executing one or more rotated object detection algorithms that determine a rotated bounding box defining the target object located within a corresponding image frame of the image data.The controllers estimate a UWB-based position of the target object by executing one or more range-based localization algorithms that analyze the sensor signals and fuse the camera-based position of the target object and the UWB-based position of the target object together using a Kalman filter to estimate the position of the target object.
[0024] Further areas of applicability will become apparent from the description provided herein. It should be understood that the description and specific examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Brief description of the drawings
[0025] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Fig. 1 illustrates a schematic diagram of a vehicle incorporating the disclosed object detection system including an ultra-wideband (UWB) sensor network and a non-stereo camera system in electronic communication with one or more controllers, according to an exemplary embodiment; Fig. 2 is a block diagram illustrating the software architecture of the one or more Fig. 1 according to an exemplary embodiment; and Fig. 3 illustrates by means of the Fig. 1 depicts image data captured by the non-stereo camera system representing an environment around the vehicle, according to an exemplary embodiment. Detailed description
[0026] The following description is merely exemplary in nature and is not intended to limit the disclosure, application, or uses.
[0027] Referring to Fig. 1, a vehicle 10 is illustrated that includes the disclosed object detection system 12. As explained below, the object detection system 12 estimates the position of one or more target objects 14 located in an environment 16 around the vehicle 10 by fusing data collected by an ultra-wideband (UWB) sensor network 22 and a non-stereo camera system 24. In the non-limiting embodiment shown in the figures, the target object 14 located in the environment 16 is a secondary vehicle, and the vehicle 10 represents the ego vehicle. However, it should be understood that Fig. 1 is merely exemplary in nature. In fact, the target object 14 may be any type of stationary or moving object located in the environment 16, such as a pedestrian, a bicycle, an animal, a light pole, or a traffic sign. It is also understood that the vehicle 10 may be any type of vehicle, such as, but not limited to, a sedan, a truck, an SUV, a van, or a motorhome.
[0028] The object detection system 12 includes one or more controllers 20 in electronic communication with the UWB sensor network 22 and the non-stereo camera system 24. The non-stereo camera system 24 acquires image data representing the target object 14 located in the environment 16 surrounding the vehicle 12. Although a non-stereo camera system 24 is described, it should be understood that a stereo camera system may also be used. The UWB sensor network 22 includes three or more anchors 30 in wireless communication with a tag 32. In the non-limiting embodiment as shown in Fig. 1, the UWB sensor network 22 includes four anchors 30 mounted on the vehicle 12 and the tag 32 is mounted on the target object 14.
[0029] The non-stereo camera system 24 includes a single camera 40 mounted on the vehicle 12 that captures image data indicative of the environment 16 around the vehicle 10. It should be understood that although Fig. 1 illustrates a single camera 40, in embodiments the vehicle 10 may also include more than one camera 40. The anchors 30 of the UWB sensor network 22 are mounted on the vehicle 12, while the tag 32 of the UWB sensor network 22 is mounted on the target object 14 located in the environment 16 around the vehicle 10. The tag 32 is a mobile sensor that is movable away from the vehicle 10 and sends and receives sensor signals. Each anchor 30 of the UWB sensor network 22 is in wireless communication with the tag 32 to send and receive the sensor signals for tracking the position of the tag 32. The sensor signals indicate real-time distances between each anchor 30 mounted on the vehicle 12 and the tag 32.
[0030] Fig. 2 is a block diagram illustrating the software architecture of the one or more Fig. 1. The one or more controllers 20 include an object detection module 50, a depth and angle module 52, a range-based localization module 54, a calibration module 56, a sensor fusion module 58, and a path planning and navigation module 60. As explained below, the object detection module 50, the depth and angle module 52, the range-based localization module 54, and the calibration module 56 estimate a calibrated estimated camera depth D during an initial calibration procedure performed offline. cali . The calibrated estimated camera depth D cali is stored in a memory of the one or more controllers 20. The calibrated estimated camera depth D cali represents a coarse camera depth d camof the target object 14, which is determined on the basis of the image data acquired by the non-stereo camera system 24, which is based on a real depth d real of the target object 14, which is determined based on the sensor signals from the UWB sensor network 22. It should be understood that the real depth d determined based on the sensor signals from the UWB sensor network 22 is real of the target object 14 is the real depth of the target object 14 and the object detection system 12 calibrates the image data acquired by the non-stereo camera system 24 based on the sensor signals from the UWB sensor network 22.
[0031] It should be understood that the initial calibration procedure is performed offline and the calibrated estimated camera depth D calistored in a memory of the one or more controllers 20. The initial calibration procedure will now be described. Referring to both Fig. 2 and Fig. 3, the object detection module 50 of the one or more controllers 20 receives the image data acquired by the non-stereo camera system 24, wherein the image data indicates the environment 16 around the vehicle 12. The object detection module 50 executes one or more algorithms for detecting rotated objects that have a (in Fig. 3) that identifies the target object 14 located within a corresponding image frame 72 of the image data. It should be understood that the target object 14 may have any orientation that does not coincide with a horizontal axis 74 of a corresponding image frame 72 ( Fig. 3) and therefore requires the rotated bounding box 70. An example of a rotated object detection algorithm is the Darknet-53 "You only look once" (YOLO) algorithm using a recurrent neural network (RNN); however, it should be understood that other types of rotated object detection algorithms may also be used.
[0032] The object detection module 50 of the one or more controllers 20 determines a plurality of position parameters of the rotated bounding box 70, wherein the position parameters indicate a location or position of the rotated bounding box 70 relative to the image frame 72 and dimensions of the rotated bounding box 70. In particular, in one embodiment, the position parameters of the rotated bounding box 70 include an x-axis position coordinate (expressed in pixels), x b , a y-axis pixel coordinate y (expressed in pixels) b, a width of the rotated bounding box w b , a height of the rotated bounding box h b and an angular orientation of the rotated bounding box θ b with respect to the horizontal axis 74 of the image frame 72.
[0033] The depth and angle module 52 of the one or more controllers 20 receives the plurality of position parameters of the rotated bounding box 70 and estimates a coarse camera depth d cam and a coarse angle θ of the target object 14 with respect to the camera 40 based on the plurality of position parameters of the rotated bounding box 70, the size of the image frame 72, and the dimensions of the vehicle 12. Specifically, the size of the image frame 72 includes a width W and a height H of the image frame, and the dimensions of the vehicle 12 include a real vehicle height h real and a real vehicle width w real . The coarse camera depth d camis determined based on equation 1, which is as follows: dcam=fcam−hrealhb, where f cam represents a focal length of the camera 40, wherein the focal length f cam is calibrated offline.
[0034] The coarse angle θ of the target object 14 with respect to the camera 40 is determined based on the coarse camera depth d cam determined and expressed in equations 2-4, which are as follows: x'=W2−(xb+wb2), offset=hreal∗x'hb, θ=arcsin(offsetdcam), where x' is a distance from a vertically oriented center M of the image frame 72 ( Fig. 3) to a center of the rotated bounding box 70.
[0035] The distance-based localization module 54 of the one or more controllers 20 receives the sensor signals from the anchors 30 of the UWB sensor network 22, which indicate the real-time distances between each anchor 30 mounted on the vehicle 12 and the tag 32, where the tag 32 is mounted on the target object 14. The distance-based localization module 54 executes one or more distance-based localization algorithms to determine a real-world depth d real of the target object 14 based on the real-time distances between each anchor 30 and the tag 32 of the UWB sensor network 22 based on the sensor signals received from the anchors 30. As will be explained below, in concrete terms, in order to determine the real depth d real of the target object 14, any distance-based triangulation localization algorithm, such as the least squares algorithm, may be used.
[0036] The calibration module 56 of the one or more controllers 20 receives the coarse camera depth d cam and the coarse angle θ of the target object 14 with respect to the camera 40 from the depth and angle module 52 and the real depth d real of the target object 14, which is determined by means of the UWB sensor network 22, from the distance-based localization module 54 and determines the calibrated estimated camera depth D cali . The calibrated estimated camera depth D cali takes into account an error e introduced by a lens distortion of the camera 40, the error e having a center d cen of the picture frame 72 (shown in Fig. 3) is linearly related.
[0037] It should be understood that a linear relationship exists between the coarse camera depth d determined by the object detection module 50 cam and the center of cen of the picture frame 72 (shown in Fig. 3). Specifically, in one embodiment, the relationship between the coarse camera depth d cam and the center of cen of the image frame 72 is expressed by a straight line equation or equation of a line or y = βx + γ, where β represents a gradient of the line and y represents a y-axis intercept of the line. During the initial calibration procedure, the calibration module 56 of the one or more controllers 30 solves a linear regression model that establishes a relationship between the coarse camera depth d cam , the real depth of real , the error e, the gradient β, the intersection point y and the center d cen of the image frame 72 and is expressed in equation 5: e=dcam−dreal=β∗dcen+γ, where the center d cen of the image frame 72 based on the position parameters of the rotated bounding box 70 (in Fig. 3). In particular, in one embodiment, the center d cen of the image frame 72 based on the x-axis position coordinate x b , the width of the rotated bounding box w b and the height of the rotated bounding box h b calculated and is expressed in equation 6 as: dcen=(xb−wb2)+(yb−hb2)2
[0038] The gradient β is based on a partial derivative of the center d cen of the picture frame 72 (in Fig. 3) and a partial derivative of the coarse camera depth d cam and is expressed in equation 7 as: β=δcamδcen
[0039] The calibration module 56 of the one or more controllers 20 resolves for the intersection point γ by positioning the target object 14 at the center d cen of the image frame 72 and a distance difference between the center of the d cenof the image frame 72 and the coarse camera depth d cam is determined, and this is expressed in equation 8 as: γ=dcam−dcen
[0040] The calibration module 56 of the one or more controllers 20 triggers, according to the calibrated estimated camera depth D cali based on a difference between the coarse camera depth d cam and the error e, and this is expressed in equation 9 as: Dcali=dcam−e
[0041] Referring to Fig. 2, the sensor fusion module 58 of the one or more controllers 20 receives the image data from the non-stereo camera system 24 and estimates a camera-based position LxCam, LyCam of the target object 14 based on the image data, wherein the camera-based position LxCam, LyCam is adjusted to the calibrated estimated camera depth D calidetermined by the initial calibration procedure described above. The sensor fusion module 58 of the one or more controllers 20 determines the camera-based position LxCam, LyCam of the target object 14 by executing one or more rotated object detection algorithms to determine the (in Fig. 3) rotated boundary box 70 which defines the target object 14 located in the environment 16 around the vehicle 12 ( Fig. 1), the position parameters of the rotated bounding box 70 (x b , y b , w b , h b , θ b ) and the coarse camera depth d cam is estimated based on the equation 1 (shown above), where the coarse camera depth d cam the position LxCam, LyCam of the target object 14. The sensor fusion module 58 then calibrates the position LxCam, LyCam of the target object 14 based on the calibrated estimated camera depth D cali and the coarse angle θ of the target object 14 based on equation 10, which is: [LxCam,LyCam]=[sinθ∗Dcali,cosθ∗Dcali]
[0042] The sensor fusion module 58 of the one or more controllers 20 receives the sensor signals from the anchors 30 of the UWB sensor network 22, which indicate the real-time distances between each anchor 30 mounted on the vehicle 12 and the tag 32. The sensor fusion module 58 of the one or more controllers estimates a UWB-based position LxUWB,LyUWB of the target object 14 by executing one or more range-based localization algorithms that analyze the sensor signals. In one embodiment, the sensor fusion module 58 executes a range-based triangulation localization algorithm, such as the least squares algorithm, to determine the UWB-based position LxUWB,LyUWB of the target object 14, which is expressed in equation 11 as: [LxUWB,LyUWB]=fLSE(d1UWB,d2UWB,…,dnUWB,o1UWB,o2UWB,…,onUWB), where f LSE represents the least squares algorithm, d1UWB,d2UWB,…,dnUWB represent a distance range between a respective anchor 30 and the tag 32, wherein the UWB sensor network 22 comprises an n number of anchors 30 and o1UWB,o2UWB,…,onUWB represent the mounting positions of each anchor 30 on the vehicle 12 ( Fig. 1).
[0043] The sensor fusion module 58 of the one or more controllers 20 fuses the camera-based position LxCam, LyCam of the target object 14 and the UWB-based position LxUWB,LyUWB of the target object 14 with each other using a Bayesian filter to estimate the position of the target object 14 ( Fig. 1). Some examples of Bayesian filters include, but are not limited to, a particle filter and a Kalman filter. In the example described, a Kalman filter is used to estimate the position of the target object 14. It should be understood that other types of filters may also be used. For example, in another embodiment, a linear or nonlinear filter may be used instead.
[0044] In one embodiment, a process model of the Kalman filter predicts a plurality of state vectors x kof the vehicle 12 and a measurement model of the Kalman filter performs an update of the plurality of state vectors x k of the vehicle 12, which are determined by means of the process model, based on an observation vector z k through. The multitude of state vectors x k of the vehicle 12 comprises a current position L expressed in x and y coordinates x , L y of the vehicle 12 and a current speed V x , V y of the vehicle 12, expressed in x and y coordinates at a current timestamp k.
[0045] The process model of the Kalman filter is expressed in equation 12 as: xk=Axk−1+wk, where A represents a state transition matrix at a time k - 1, x k-1 represents the state vectors at a previous time k - 1 and w k represents normally distributed system noise.
[0046] The measurement model of the Kalman filter is expressed in equation 13 as: zk=Hxk+vk, where z k an observation vector to a current (k th ) timestamp, H represents the observation matrix of either the camera 40 or the UWB sensor network 22 and v k Observation noise, where the observation noise is white Gaussian noise. The observation matrix H of the camera 40 is expressed as H=[LxCam,0,0,00,LyCam,0,0], and the observation matrix H of the UWB sensor network 22 is expressed as H=[LxUWB,0,0,00,LyUWB,0,0], The sensor fusion module 58 of the one or more controllers 20 then estimates the position of the target object 14 based on the plurality of state vectors x k of the vehicle 12 to the previous timestamp and the observation vector z kto the current timestamp based on equation 14, which is: xk=Axk−1+zk+wk
[0047] Referring to Fig.2, the path planning and navigation module 60 of the one or more controllers 20 receives a target position T of the vehicle 12, wherein the target position T represents a destination location of the vehicle 12. In one non-limiting embodiment, the target position is, for example, a parking space. The path planning and navigation module 60 creates a vector map 90 based on field strengths of the target object 14, other objects located in the environment 16 around the vehicle 12, and the target position T. In particular, the vector map 90 is created based on an attractive field strength from the target position T, repulsive field strengths from the target position 14 and the other objects located in the environment 16, and a total field strength from a current position of the vehicle 12.It should be understood that the current position of the vehicle 12 is constantly changing during the travel of the vehicle 12, and therefore the current position of the vehicle 12 can be represented by each pixel that is part of the vector map 90. The attractive field strength at the target position T is expressed in Equation 15, and an angle of the attractive field strength is expressed in Equation 16 as:. PA=c|x−xg|2+|y−yg|2 θA=tan−1(yg−yxg−x), where P A represents the attractive field strength, (x g , y g ) represents the coordinates of the target position T, c represents a constant, θ Arepresents the angle of the attractive field strength and x, y represent the x- and y-coordinates of the target position T. The repulsive field strengths of the target object 14 and the other target objects located in the environment 16 are determined by determining a distance between the vehicle 12 and a corresponding object, and this is expressed in equation 17 as: d=|x−xk|2+|y−yk|2, where d represents the distance between the vehicle 12 and a k-th object and x k , y k represent the x and y coordinates of the kth object.
[0048] The path planning and navigation module 60 then categorizes each object by class, where the object class indicates a specific type of object. Some examples of the specific object type include, but are not limited to, a human, a passenger car, a truck, a public bus, an animal such as a dog or cat, a traffic sign, a bicycle, and a piece of furniture such as a table or chair. The class of each object indicates the object's overall dimensions (e.g., height and width) and the object's hazard level, where the hazard level indicates the degree of impact the object may have on the vehicle 12 in the event of a collision. The hazard level is based on an estimated mass of the object, where the estimated mass is determined based on the object's overall dimensions. An object with a greater mass poses a greater hazard to the vehicle 12.The path planning and navigation module 60 can then assign a size s to each object. k , a distance parameter r k and a force-mass parameter m k based on the class of the object. In particular, the size s k assigned based on the overall dimensions, the distance parameter r k assigned on the basis of the degree of danger, with a higher degree of danger increasing the distance parameter r k increases, and the force-mass parameter m k assigned based on the estimated mass of the object. If the size s k of the object is less than the distance d between the vehicle 12 and the object and if the distance d is less than the distance parameter r k , then the path planning and navigation module 60 solves for the repulsive field strength from the object in equation 18 and an angle of the repulsive field strength in equation 19 as: PRk=mk∗|x−xk|2+|y−yk|2 θRk=tan−1(yk−yxk−x), where PRk represents the repulsive field strength from the object and θ R represents the angle of the repulsive field strength.
[0049] The route planning and navigation module 60 determines the total field strength from the current position of the vehicle 12 based on the attractive field strength P A , the angle of the attractive field strength θ A , the repulsive field strength from the object PRk and the angle of the repulsive field strength θRk and is expressed in equation 20 as: P=(PA,θA)+∑k=1n(PRk,θRk), where P represents the total field strength.
[0050] Referring to the figures, the object detection system provides various technical effects and advantages. In particular, the object detection system provides an approach for calibrating the coarse camera depth of a target object, calculated using image data acquired by the non-stereo camera system, based on a real depth of the target object determined by the UWB sensor system, resulting in reduced complexity and computational requirements compared to an approach that utilizes perception data collected by LiDAR sensors. The disclosed approach also consumes less power compared to an approach that utilizes perception data collected by radar sensors.It should also be understood that identifying the target object based on rotated object detection algorithms that specify a rotated bounding box provides improved depth and angle estimation compared to object detection algorithms that use a bounding box aligned with the horizontal axis of the image.
[0051] The controllers may refer to or be part of an electronic circuit, a combinational logic circuit, a field-programmable gate array (FPGA), a processor (shared, dedicated, or grouped) that executes code, or a combination of some or all of the above elements, such as in a system-on-chip. Furthermore, the controllers may be based on or be part of a microprocessor, such as a computer having at least one processor, memory (RAM and / or ROM), and associated input and output buses. The processor may operate under the control of an operating system located in memory. The operating system may manage computer resources such that computer program code embodied as one or more computer software applications, such as an application located in memory, may include instructions executed by the processor.In an alternative embodiment, the processor may execute the application directly, in which case the operating system may be omitted.
[0052] The description of the present disclosure is merely exemplary in nature, and variations that do not depart from the gist of the present disclosure are intended to be considered within the scope of the present disclosure. Such variations are not to be considered a departure from the spirit and scope of the present disclosure.
Claims
[1] An object detection system (12) for a vehicle (10) that estimates the position of a target object (14) located in an environment (16) around the vehicle (10), the object detection system (12) comprising: an ultra-wideband (UWB) sensor network (22) comprising three or more anchors (30) mounted on the vehicle (10) in wireless communication with a tag (32) mounted on the target object (14), each anchor (30) transmitting and receiving sensor signals indicative of real-time distances between each anchor (30) and the tag (32); a non-stereo camera system (24) that acquires image data representing the target object (14) located in the environment (16) around the vehicle (10); and one or more controllers (20) in electronic communication with the UWB sensor network (22) and the non-stereo camera system (24), the one or more controllers (20) including one or more processors executing instructions to: estimate a camera-based position of the target object (14) based on the image data, wherein the camera-based position is adjusted to account for a calibrated estimated camera depth determined during an initial calibration procedure; to estimate a UWB-based position of the target object (14) by executing one or more distance-based localization algorithms that analyze the sensor signals; and fusing the camera-based position of the target object (14) and the UWB-based position of the target object (14) using a Bayesian filter to estimate the position of the target object (14), wherein the calibrated estimated camera depth represents a coarse camera depth of the target object (14) determined based on the image data acquired by the non-stereo camera system (24) calibrated based on a real depth of the target object (14), wherein the real depth of the target object (14) is determined based on the sensor signals of the UWB sensor network (22). [2] The object detection system (12) of claim 1, wherein the Bayesian filter is a Kalman filter. [3] The object detection system (12) of claim 2, wherein a process model of the Kalman filter predicts a plurality of state vectors of the vehicle (10). [4] The object detection system (12) of claim 3, wherein a measurement model of the Kalman filter performs an update of the plurality of state vectors of the vehicle (10) determined by the process model based on an observation vector. [5] The object detection system (12) of claim 1, wherein the initial calibration procedure comprises: Executing one or more rotated object detection algorithms that determine a rotated bounding box (70) that identifies the target object (14) located within a corresponding image frame (72) of the image data. [6] The object detection system (12) of claim 5, wherein the algorithm for detecting rotated objects is the Darknet 53 "You only look once" (YOLO) algorithm with a recurrent neural network (RNN). [7] The object detection system (12) of claim 5, wherein the initial calibration procedure comprises: Determining a plurality of position parameters of the rotated bounding box (70), wherein the position parameters of the rotated bounding box (70) include an x-axis position coordinate, a y-axis pixel coordinate, a width of the rotated bounding box (70), a height of the rotated bounding box (70), and an angular orientation of the rotated bounding box (70) with respect to the horizontal axis (74) of the corresponding image frame (72). [8] The object detection system (12) of claim 7, wherein the initial calibration procedure comprises: Determine a coarse camera depth based on: dcam=fcam*hrealhb, where d cam represents the coarse camera depth, h b the height of the rotated boundary box (70) h b represents, f cam represents a focal length of a camera (40) that is part of the non-stereo camera system (24), and d realrepresents a real depth of the target object (14) determined on the basis of the sensor signals received from the three or more anchors (30).