Target detection positioning and visual cost map generation method based on depth camera
Through the depth camera-based target detection and positioning and visual cost map generation method, the problem that a single low-cost radar is difficult to perceive a dynamic environment is solved, and more realistic dynamic obstacle perception and path planning are achieved, avoiding the occurrence of dangerous behaviors.
Patent Information
- Application Number
- CN202311685121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-03
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to achieve comprehensive and reliable dynamic environment perception through a single low-cost radar, and the point cloud information processing of depth cameras is large in amount and hardware computing power requirements are high.
The object detection and positioning and visual cost map generation method based on depth cameras are used to realize the functions of each node and communication between nodes through the ROS system, and the Bresenham algorithm is used to draw the occupied area of dynamic obstacles on the two-dimensional grid map.
It effectively makes up for the lack of low-performance radar's perception ability of dynamic obstacles, and more realistically reflects the occupancy of dynamic obstacles on the grid map through visual cost maps, avoiding the path passing through dangerous areas.
Smart Images

Figure BSA0000297594600000041 
Figure BSA0000297594600000043 
Figure BSA0000297594600000051
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of environmental perception and autonomous navigation of unmanned vehicles, and particularly relates to a method for target detection and positioning based on a depth camera and generation of a visual cost map. Background Art
[0002] In the field of autonomous navigation of unmanned vehicles, the perception module senses the surrounding environment through sensors and reflects the information on the cost map, which is the basis for subsequent path planning and motion control. The current perception method relying on radar can better sense the static environment, but limited by sensor performance and interference factors, it is difficult to rely on a single low-cost radar to complete comprehensive and reliable dynamic environment perception. Directly enhancing the environment perception ability by improving radar performance will cause a sharp increase in the overall vehicle cost. As one of the most basic sensors on unmanned vehicles, the camera can capture rich road information and has an absolute advantage in detecting and recognizing typical dynamic targets such as pedestrians and vehicles. However, directly processing the point cloud information of the depth camera without discrimination has a large amount of calculation and high requirements for hardware computing power. Therefore, how to reasonably utilize the image and depth information of the depth camera to improve the perception ability of unmanned vehicles for dynamic environments under the condition of limited radar performance has become a research hotspot in the academic and industrial fields. Summary of the Invention
[0003] The technical problem to be solved by the present invention is how to sense typical dynamic obstacles through a depth camera and project them onto a two-dimensional grid cost map. To solve the above problem, the present invention provides a method for dynamic obstacle detection, positioning and cost map development based on a depth camera, which realizes the functions of each node and communication between nodes under the ROS system. The method for target detection and positioning based on a depth camera and generation of a visual cost map includes the following steps:
[0004] Step 1: Start the unmanned vehicle chassis and the depth camera in an environment where the static obstacle mapping has been completed, and publish the RGB image and the depth image respectively;
[0005] Step 2: The target detection node subscribes to the RGB image, identifies typical dynamic targets such as pedestrians, cars, bicycles, motorcycles, electric bicycles, electric motorcycles, trucks in the image and publishes the detection frame;
[0006] Step 3: The target positioning node subscribes to the detection frame and reads the depth information at the corresponding position of the depth image according to the pixel position of the center of the detection frame;
[0007] Step 4: Calculate the coordinates of the target center in the camera coordinate system;
[0008] Step 5: Calculate the coordinates of the target center point in the vehicle chassis coordinate system;
[0009] Step 6: Calculate the coordinates of the lower left and lower right corners of the detection box in the chassis coordinate system, and calculate the distance between the two points as the target width;
[0010] Step 7: Calculate the coordinates of the target center point in the world coordinate system, and publish the coordinate values of the target in the world coordinate system and the target width together;
[0011] Step 8: Develop the cost map layer of the vision layer, subscribe to the target coordinates and width on the vision layer, and convert the coordinates into grid indices and the width into the number of cells;
[0012] Step 9: Use the Bresenham algorithm to draw a circle with the target center as the center and the target width as the diameter on the grid map, and mark the grids inside the circle as obstacles to achieve the superposition of visual perception information.
[0013] Advantages of the present invention:
[0014] The present invention can make up for the problem that low-performance radars have limited perception ability for typical dynamic obstacles such as pedestrians and vehicles due to the limitation of their installation positions. As a supplement to the local map, the cost map of the vision layer can more realistically reflect the occupancy of dynamic obstacles on the grid map, and avoid dangerous behaviors such as the planned path passing through the gaps between the limbs of pedestrians or the gaps between the tires of vehicles. Description of the drawings
[0015] Figure 1 It is a relationship diagram of each module of the present invention;
[0016] Figure 2 It is a schematic diagram of the camera imaging principle of the present invention;
[0017] Figure 3 It is a schematic diagram of drawing a circle by the Bresenham algorithm of the present invention;
[0018] Figure 4 It is a schematic diagram of the effect of the present invention. Detailed implementation manners
[0019] The present invention will be further described below in conjunction with the drawings and simulation examples in Gazebo, so that those skilled in the art can better understand the present invention and be able to implement it, but the examples given are not intended to limit the present invention.
[0020] As Figure 1As shown in the figure, the depth camera 1 publishes the RGB image 2 and the depth image 3. The object detection node 4 based on YOLOv5 subscribes to the RGB image. After completing the object detection, it publishes the object detection box 5. The object localization node 6 calculates the object center coordinates and width 7 based on the position of the object detection box and the depth information and publishes them. The cost map 8 of the visual information layer subscribes to the object center position and width and uses the Bresenham algorithm to draw a circle, which together with the radar layer 9 and the static layer 10 forms a multi-layer cost map 11 to complete the superposition of visual perception information.
[0021] The object detection, localization and visual cost map generation method based on a depth camera according to the present invention comprises the following steps:
[0022] Step 1: Start the unmanned vehicle in a simulation scenario where the static obstacles have been mapped. Load the original static map and the visual layer cost map plugin, complete the correction of the vehicle's initial pose, start the depth camera, and publish the RGB image and the depth image respectively. Place a pedestrian model directly in front of the unmanned vehicle.
[0023] Step 2: Start the object detection node based on Yolo_v5 to subscribe to the RGB image of the depth camera, identify the pedestrian target in the image and publish the detection box.
[0024] Step 3: Start the object localization node, subscribe to the detection box and according to the pixel position [u, v] of the center of the detection box T , read the depth information z at the corresponding position of the depth image c . Wherein, u is the abscissa of the center of the detection box in the pixel coordinates, and v is the ordinate of the center of the detection box in the pixel coordinates;
[0025] Step 4: The camera imaging principle is as Figure 2 shown. O-x-y-z is the camera coordinate system, O′-x′-y′ is the image coordinate system, o-u-v is the pixel coordinate system, f is the camera focal length, P is a point in the physical world, z c is the depth value of point P, and P′ is the projection of P on the image plane. Calculate the coordinates of the object center in the camera coordinate system from the following formula;
[0026]
[0027] Wherein f x , f y are the normalized focal lengths in the x and y directions respectively, c x , c y is the offset of the pixel coordinate origin relative to the image coordinate origin. The camera internal parameter matrix K is given at the factory. x c , y c , z crespectively represent the components of the target center on the x-axis, y-axis, and z-axis in the camera coordinate system;
[0028] Step 5: Calculate the coordinates of the target center point in the vehicle chassis coordinate system using the following formula;
[0029]
[0030] where the orthogonal rotation matrix R and the translation matrix T describe the installation position of the camera on the vehicle chassis and are obtained through camera calibration, and x b , y b , z b respectively represent the components of the target center on the x-axis, y-axis, and z-axis in the vehicle chassis coordinate system;
[0031] Step 6: For the positions of the lower left and lower right corners of the detection box, repeat Steps 3 - 5 to calculate their coordinates in the chassis coordinate system, and calculate the distance between the two points as the target width d;
[0032] Step 7: Calculate the coordinates of the target center point in the world coordinate system [x w , y w , z w T , and publish [x w , y w , z w T and d, where x c , y c , z c respectively represent the components of the target center on the x-axis, y-axis, and z-axis in the world coordinate system;
[0033] Step 8: Subscribe to [x w , y w , z w T and d on the vision layer plugin, and convert the projection point of [x w , y w , z w T on the two-dimensional grid map into grid indices [x m , y m T , and convert d into the number of grid cells d m :
[0034]
[0035] where x m , y m respectively represent the row and column index values of the target center projected on the two-dimensional grid map, and x minRepresents the minimum value of the abscissa of the map in the world coordinate system, y min Represents the minimum value of the ordinate of the map in the world coordinate system. Resulution is the map resolution, indicating the length in the real world represented by the side length of a grid. The function floor(·) represents taking the integer part of the input parameter;
[0036] Step 9: As Figure 3 shown, use the Bresenham algorithm to draw a circle on the grid map with [x m , y m T as the center and d m as the diameter, and utilize the "eight-fold symmetry" of the circle to draw the circle on the grid map. As Figure 4 shown, set the cost value of the grids inside the circle to the obstacle threshold to achieve the superposition of visual perception information. Effect comparison: The perception effect relying only on the radar is 12, and the perception effect after superimposing the visual cost map is 13.
Claims
1. Method for target detection and positioning based on depth camera and generation of visual cost map, Characterized in that, It includes the following steps: Step 1: Start the unmanned vehicle in the simulation scenario where static obstacle mapping has been completed, load the original static map and the visual layer cost map plugin, complete the initial pose correction of the vehicle, start the depth camera, publish RGB images and depth images respectively, and place a pedestrian model directly in front of the unmanned vehicle; Step 2: Start the target detection node based on Yolo_v5 to subscribe to the RGB images of the depth camera, identify the pedestrian target in the image and publish the detection box; Step 3: Start the target positioning node, subscribe to the detection box, and based on the pixel position [u, v] at the center of the detection box T , read the depth information z at the corresponding position of the depth image c , where u is the abscissa of the center of the detection box in pixel coordinates, and v is the ordinate of the center of the detection box in pixel coordinates; Step 4: Calculate the coordinates of the target center in the camera coordinate system by the following formula; Among them f x and f y are the normalized focal lengths in the x and y directions respectively, c x and c y are the offsets of the pixel coordinate origin relative to the image coordinate origin. The camera internal parameter matrix K is given at the time of factory. x c and y c and z c represent the components of the target center on the x-axis, y-axis, and z-axis in the camera coordinate system respectively; Step 5: Calculate the coordinates of the target center point in the vehicle chassis coordinate system by the following formula; Among them, the orthogonal rotation matrix R and the translation matrix T describe the installation position of the camera on the vehicle chassis, which are obtained through camera calibration. x b , y b , z b respectively represent the components of the target center on the x-axis, y-axis, and z-axis in the vehicle chassis coordinate system; Step 6: For the positions of the lower left corner and the lower right corner of the detection box, repeat Steps 3 - 5 to calculate their coordinates in the chassis coordinate system, and calculate the distance between the two points as the target width d; Step 7: Calculate the coordinates [x w , y w , z w of the target center point in the world coordinate system by TF transformation in ROS T , and publish [x w , y w , z w T and d, where x c , y c , z c represent the components of the target center on the x-axis, y-axis, and z-axis in the world coordinate system respectively; Step 8: Subscribe to [x w , y w , z w T and d on the vision layer plugin, and convert the projection point of [x w , y w , z w T on the two-dimensional grid map into grid indices [x m , y m , T and convert d to the number of grids d m : Among them, x m , y m respectively represent the row and column index values of the target central projection on the two-dimensional grid map. x min represents the minimum value of the map abscissa in the world coordinate system, and y min represents the minimum value of the map ordinate in the world coordinate system. The resolution is the map resolution, indicating that the side length of a grid represents the length in the real world. The function floor(·) represents taking the integer part of the input parameter; Step 9: Use the Bresenham algorithm to draw a circle on the grid map with [x m , y m T as the center and d m as the diameter. Utilize the "eight-fold symmetry" of the circle to draw the circle on the grid map, and set the cost value of the grids inside the circle to the obstacle threshold to achieve the superposition of visual perception information.
2. The method for target detection and positioning based on depth camera and generation of visual cost map according to Claim 1, Characterized in that, Typical dynamic obstacles that can be recognized by the target detection module based on Yolo_v5 in Step 2 include but are not limited to: pedestrians, cars, bicycles, motorcycles, electric bicycles, electric motorcycles, trucks.
3. The method for dynamic target detection and positioning based on depth camera and generation of visual cost map according to Claim 1, Characterized in that, The calculation of the target width in Step 6 is based on the pixel positions of the left and right borders of the width detection box, and the selection of the left and right border pixels is not limited to the lower left corner pixel and the lower right corner pixel.