Sensor information fusion method based on single-line laser radar and camera

By projecting lidar points onto the camera image and using YOLOv8 target detection algorithm to correct lidar data, the problem of limited field of view combined with lidar and incomplete access to obstacle information when autonomous mobile robots navigate in complex environments is solved, achieving more comprehensive environmental perception and seamless simulation and reality adaptation.

CN120122111APending Publication Date: 2025-06-10BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122176.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When autonomous mobile robots navigate in complex environments, the combination of single-line lidar and monocular color cameras has the problem of limited field of view and incomplete access to obstacle information.

Method used

By projecting lidar points onto the image acquired by the camera, the YOLOv8 target detection algorithm is used to obtain obstacle information, correct lidar data, so that it can express more comprehensive obstacle profile and volume, and replace the original lidar data within the camera's field of view.

Benefits of technology

The robot's perception of the environment is improved, the lidar's ability to identify obstacles is enhanced, and the method is seamlessly adapted to simulation and real environments, reducing the coupling of sensor data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122111A_ABST
    Figure CN120122111A_ABST
Patent Text Reader

Abstract

The invention discloses a sensor information fusion method based on a single-line laser radar and a camera, and the method comprises the steps: 1, obtaining the information of an obstacle in front of a robot according to a monocular color camera carried on the robot, and 2, carrying out the projection transformation and information fusion between the laser radar and an image, firstly, a laser radar point is projected to an image acquired by a camera, the corresponding laser radar point is corrected according to obstacle information on the image, and then the corrected laser radar point is re-projected to a coordinate system of an original laser radar point, so that the laser radar can express the maximum contour and the longitudinal volume of an obstacle which cannot be sensed originally. And finally, the original laser radar data in the view field range of the camera is replaced by the corrected laser radar data. According to the method, coupling does not exist between two kinds of sensor data, so that the effect in a simulation environment and the effect in a real environment do not have difference, and the method has high robustness and is easy to implement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for fusing sensor information based on a single-line lidar and a camera, and belongs to the field of sensor information fusion for autonomous mobile robots. Background Art

[0002] With the rapid development of the robotics industry, the application of autonomous mobile robots has seen unprecedented growth in various fields. For example, mobile robots can assist humans in completing various tasks in scenarios such as hospitals, supermarkets, and warehouses. The basis for completing these tasks is that mobile robots need to have perfect navigation capabilities. When a mobile robot navigates in a complex environment, one of the main challenges it faces is how to more comprehensively perceive the surrounding environment. On mobile robots, lidar and cameras are commonly used environmental perception sensors. For economic reasons, autonomous mobile robots usually use single-line two-dimensional lidar and monocular color cameras. However, both of these sensors have inherent limitations in environmental sensing. The field of view of the camera is limited, and the ability of the single-line lidar to detect obstacles with longitudinal shape changes and obtain comprehensive information about obstacles with irregular shapes is greatly limited. Therefore, the present invention proposes a method for fusing sensor information of a single-line lidar and a color camera. This method uses the obstacle information in the image to improve the ability of the lidar to identify obstacles, and does not depend on the environment in which the image is located, and there is no coupling between the image and the lidar data. Therefore, this method can be seamlessly transferred from a virtual environment to real-world applications, such as being used in a navigation framework for reinforcement learning. Summary of the Invention

[0003] In view of the above problems in the prior art, the object of the present invention is to propose a method for fusing sensor information of a single-line lidar and a camera. This method realizes a more comprehensive and accurate perception of the environment by efficiently integrating the precise distance data of the lidar and the rich information of the camera image, and can be applied to mobile robots in reality. Considering the convenience during application and the actual situation of the robot, the Robot Operating System (ROS) is selected to obtain and transmit sensor information. The overall idea of this method is to project lidar points onto the image obtained by the camera, and correct the lidar data based on the obstacle position information obtained from the camera image, so that the lidar can represent the maximum contour and longitudinal volume of obstacles that could not be perceived originally. Finally, the corrected lidar data is used to replace the original lidar data within the camera's field of view. Among them, the relatively mature YOLOv8 network is selected for the image target detection method, which has the characteristics of high speed and high integration, and users can conveniently create and train custom datasets. The multi-sensor information fusion method of the present invention mainly considers its application in the reinforcement learning navigation framework, so it attaches importance to seamless adaptability in simulation scenarios and real-world scenarios, that is, it can have the same performance in simulation and reality without modification.

[0004] The sensor information fusion method based on a single-line lidar and a camera of the present invention mainly includes the following two parts:

[0005] The first part: Obtaining obstacle information based on deep learning object detection

[0006] Obtain obstacle information in front of the robot according to the monocular color camera mounted on the robot;

[0007] The second part: Projection transformation and information fusion between the lidar and the image

[0008] First, project lidar points onto the image obtained by the camera, correct the corresponding lidar points according to the obstacle information on the image, and then re-project the corrected lidar points back to the coordinate system of the original lidar points,

[0009] so that the lidar can represent the maximum contour and longitudinal volume of obstacles that could not be perceived originally. Finally, the corrected lidar data is used to replace the original lidar data within the camera's field of view.

[0010] Furthermore, the obtaining of the obstacle information in front of the robot includes:

[0011] Step 1: Camera image acquisition and processing, including: The camera on the robot acquires and publishes a video stream, reduces the image while keeping the aspect ratio unchanged to a lower resolution, and modifies the color channel of the reduced image from BGR to RGB;

[0012] Step 2: Training and application of the YOLOv8 network model, including: Inputting the processed image into the pre-trained YOLOv8 network model to obtain the target information required in the input image, and publishing this information as a ROS topic as the output of the YOLOv8 network for subsequent use.

[0013] Among them, the target information output by the YOLOv8 network includes confidence, the boundary of the target in the input image, and the target type.

[0014] Among them, the pre-trained YOLOv8 network model is trained through the following steps: (1) Determine the training dataset according to the specific usage scenario of the robot; (2) To enhance the accuracy of model training, the training dataset needs to incorporate data points in special situations, that is, data under various non-standard conditions, including: extremely close obstacles, distant and blurred targets, obstacles blocking each other, etc.

[0015] Further, the second part is the joint calibration of the single-line lidar and the camera, including the following steps:

[0016] Step 1: Joint calibration of the camera and the lidar: First, default the known internal parameters of the camera and arrange the calibration scenario; then, start the robot through the ROS system and publish the lidar and camera sensor data to the ROS topic, and display the two-dimensional point cloud data obtained by the lidar; Record the lidar scan point information at the landmark positions on the front obstacle; Identify the positions where the above lidar scan points should fall in the image, and record the pixel positions on the image corresponding to these lidar scan points; Finally, calculate the rotation matrix and translation vector for transforming from the world coordinate system with the lidar sensor installation position as the origin to the camera coordinate system, that is, the external parameters of the camera in the current layout.

[0017] Step 2: Project the lidar scan points onto the image through the transformation between coordinate systems. After determining the external parameters of the camera, this projection process is represented by the equation of camera calibration; (1) Through the external parameters of the camera, the coordinates of the lidar scan points in the camera coordinate system can be obtained; (2) Then, through the focal length parameter in the internal parameters of the camera, the three-dimensional coordinates in the camera coordinate system are converted into two-dimensional coordinates in the image coordinate system; (3) Finally, the points on the image coordinate system are converted into the pixel coordinate system; (4) Combining the above three results of (1)-(3), the transformation from the world coordinate system to the pixel coordinate system is obtained; (5) Performing the transformation on each lidar scan point can project all the lidar points within the camera's field of view onto the camera image.

[0018] Step 3: Correction of lidar data. Correct all the lidar points that fall within the obstacle range and record the correction information for them.

[0019] Step 4: Reprojection from image to lidar. If the corrected lidar data is to be converted into usable sensor data, these pixel points also need to be reprojected back to the lidar points in the world coordinate system; then, according to the correction information recorded in Step 3, the reprojected lidar data is corrected; after the corrected lidar data is published, it is used to replace the original lidar data within the field of view of the camera directly in front of the robot.

[0020] Among them, the transformation formula from the world coordinate system to the pixel coordinate system is:

[0021]

[0022] Among them, f x , f y , u 0 , v 0 are the internal parameters of the camera, representing the focal length and the optical center coordinates of the camera; R, T are the external parameters of the camera, representing the rotation matrix and the translation vector.

[0023] Among them, for reprojecting the pixel points back to the lidar points in the world coordinate system, the formula needs to be derived from another perspective: First, organize formula (4) and abbreviate it as:

[0024]

[0025] Among them, K represents the internal parameter matrix of the camera; further, placing the points (X w , Y w , Z w ) in the world coordinate system required on the left side, formula (6) can be transformed into:

[0026]

[0027] To further simplify and calculate the above formula (7), two matrices Mat 1 and Mat 2 are constructed. Let:

[0028] Mat 2 = R -1 T

[0029] where all the elements that make up Mat 1 and Mat 2 are known; since what needs to be calculated is Z c , only the terms in the third row of formula (7) need to be calculated, so there is:

[0030]

[0031] where Z w is the height value of a point in the world coordinate system. Since the world coordinate system takes the installation position of the lidar sensor on the robot as the origin, and the single-line lidar only scans a plane at a fixed height, the value of Z w must be 0. After calculating the value of Z c , substituting it into formula (7) can calculate the corresponding point in the world coordinate according to the pixel coordinates obtained from the image.

[0032] Advantages and Efficacy

[0033] The present invention proposes a sensor information fusion method based on a single-line lidar and a camera, aiming to improve the ability of the single-line lidar to identify obstacles through the obstacle information in the image. This method relies on the YOLOv8 object detection algorithm and significantly improves the ability of the robot to perceive the surrounding environment through the lidar. Its main advantages include:

[0034] Low requirements: The sensors only need a single-line lidar and a monocular color camera, and the robot only needs to be able to run or remotely connect to the ROS system and the YOLOv8 object detection model to use this method;

[0035] High robustness: It can operate stably in complex and dynamic environments. If it is unable to obtain or process image information due to unexpected situations, the original lidar data can also be used;

[0036] Seamless adaptability: The proposed method does not directly use image data for sensor information fusion, and there is no coupling between the two sensor data. Therefore, there is no difference in the effect between the simulation environment and the real environment, which is conducive to the promotion from the simulation environment to the real environment.

[0037] With these advantages, the present invention provides an innovative implementation solution for the environmental perception of autonomous mobile robots equipped with single-line lidar and cameras. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a schematic flowchart of the sensor information fusion method based on single-line lidar and camera.

[0040] Figure 2 It is a partially manually obtained and manually labeled tagged dataset.

[0041] Figure 3 It is the application effect diagram of the trained YOLOv8 network model in the simulation environment.

[0042] Figure 4 It is a diagram showing the coordinate system related to the robot sensor.

[0043] Figure 5 It is the result diagram of projecting the lidar points onto the image.

[0044] Figure 6 It is the application effect diagram of the method proposed by the present invention in the simulation environment.

[0045] Figure 7 It is the application effect diagram of the method proposed by the present invention in the real environment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0047] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0048] The overall process schematic diagram of the sensor information fusion method based on a single-line lidar and a camera of the present invention is shown in Figure 1 , and mainly includes the following two parts:

[0049] The first part: Obtaining obstacle information based on deep learning object detection

[0050] This part mainly obtains the obstacle information in front of the robot according to the monocular color camera mounted on the robot. The possible obstacles in the usage scenario of the autonomous mobile robot expected by the present invention include pedestrians, desks and chairs, traffic cones, etc. The specific process of this part is as follows:

[0051] Step 1: Camera image acquisition and processing

[0052] The process of collecting images and processing the images includes: the camera on the robot collects and publishes a video stream, reduces the image to a lower resolution without changing the aspect ratio of the length and width to speed up the processing speed, and modifies the color channel of the reduced image from BGR to RGB to adapt to the image processing of YOLOv8. Specifically as follows:

[0053] ⑴ Connect to the ROS system and publish the camera topic: Start from the robot or remotely access the ROS system started on the host, check through the topic printing command to ensure that the terminal robot is successfully connected to the ROS system, and start the camera node to publish the image topic to obtain a real-time video stream;

[0054] ⑵ Image scaling: Scale the image subscribed from the image topic proportionally to a resolution of 640 by 480. This process can be implemented using fast matrix multiplication in the opencv library, specifically by using the resize function and setting the image size to a width of 640 pixels and a height of 480 pixels. This is used to improve the processing speed and enhance real-time performance;

[0055] ⑶ Color channel modification: After the image is scaled, adjust the color channels. Only the bit order of the color vector part needs to be changed, which can be achieved using the cvtColor function in the OpenCV library, and the encoding form is adjusted to COLOR_BGR2RGB.

[0056] Step 2: Train and apply the YOLOv8 network model

[0057] After obtaining the input image through processing, the next step is to input the input image into the pre-trained YOLOv8 network model to obtain the target information required in the input image, and publish this information as a ROS topic for subsequent use. The target information output by the YOLOv8 network includes confidence, the boundary of the target in the input image, and the target type.

[0058] The pre-trained YOLOv8 network model is trained through the following steps: Make a dataset containing class labels, with a total of 1000 pictures. Figure 2 Some examples of the dataset are shown. The dataset contains label information for pedestrians, tables, chairs, cone obstacles, and other obstacles. When training, select a pre-trained model with a smaller scale and train through the following steps:

[0059] ⑴ Determine the training dataset according to the specific usage scenario of the robot. The envisioned application scenario of the autonomous mobile robot in this invention is an indoor scenario. In this scenario, the main obstacles are pedestrians, tables, and chairs. Considering the experimental test requirements, there are also cone obstacles. In addition, since the robot's perspective is relatively low, a dataset needs to be specifically made from the robot's perspective. Therefore, 4 types of target objects, namely tables, chairs, pedestrians, and cones, are collected in the simulation environment, with a total of 1000 labeled training data. The training data includes the target image, its corresponding bounding box, and class label.

[0060] ⑵ To enhance the accuracy of model training, the training dataset should incorporate data points under special circumstances. Generally speaking, the breadth of the training data is directly proportional to the accuracy of the model's target matching; if the training set is limited to conventional data, the model will perform well in simple scenarios with clear targets, but will struggle in the face of complex situations. Especially when dealing with edge cases, such as individuals with significantly deviated features among similar targets, and data containing noise (this noise may stem from malfunctions of image acquisition devices, signal interruptions, or image or target blurring caused by harsh natural environments, as well as incomplete data situations, such as transmission loss or the target not fully entering the shooting field of view, etc.), the model will have difficulty detecting target information and is prone to losing the target. Therefore, the training set created in this invention covers various types of data under non-standard conditions, including extremely close-range situations where only the feet of pedestrians can be captured, situations where distant obstacles are blurred, situations where obstacles are blocked by walls or other obstacles, situations of environmental light changes, etc. Considering training efficiency and recognition speed, the YOLOv8s model with a smaller number of parameters is selected as the pre-training weight, and the number of training iterations is set to 2000 times to ensure the effective convergence of the model. Figure 3 Shows the recognition effect of the model in the simulation environment after training.

[0061] After the custom YOLOv8 network model is trained, its recognition results need to be published as ROS topics for use. For this purpose, a dedicated ROS message type needs to be defined. The recognition result message defined in this invention is as follows:

[0062] M result ={[Probability,x min ,y min ,x max ,y max ,Class],…}

[0063] Among them, Probability is the confidence level, x min ,y min ,x max ,y max are the diagonal coordinates of the recognition box in the image, and Class is the target type. All the results of one recognition will be published in one message.

[0064] Part Two: Projection Transformation and Information Fusion between LiDAR and Images

[0065] The main purpose of this part is to project LiDAR points onto the image obtained by the camera, correct the corresponding LiDAR points according to the obstacle information on the image, and re-project the corrected LiDAR points back to the coordinate system of the original LiDAR points. The coordinate system relationships related to the sensors on the robot are shown in Figure 4 , note Figure 4The origin O of the middle world coordinate system w It should be at the installation position of the lidar. Since on a mobile robot, the installation positions of the lidar sensor and the camera sensor are fixed, the positional relationship between them can be represented by a rotation matrix and a translation vector, that is, the extrinsic matrix of the camera. The projection between the lidar and the image is a process of obtaining the extrinsic matrix through calibration and then calculating based on the lidar point data and the intrinsic and extrinsic matrices of the camera. The specific steps are as follows:

[0066] Step 1: Joint calibration of the camera and the lidar

[0067] The joint calibration of the lidar and the camera is to unify the data of the two sensors into the same coordinate system, so as to achieve more accurate environmental perception. The lidar provides accurate distance information, while the camera captures rich information within the field of view. By joint calibration, the advantages of both can be combined to make up for the deficiencies of a single sensor. After the relative positions of the camera and the lidar are fixed, what needs to be done is to correspond the points in the world coordinate system scanned by the lidar one by one with the pixel points at the corresponding positions on the image, and calculate the rotation matrix and the translation vector through these two sets of points. Currently, most of the existing lidar-camera joint calibration methods are for multi-line lidar and camera joint calibration, with a relatively complex process and the need for a special calibration board, while the single-line lidar and camera joint calibration method designed in the present invention can greatly simplify the process.

[0068] First, the intrinsic parameters of the camera are assumed to be known, and the calibration scene is arranged. Place the robot equipped with the camera sensor and the lidar sensor on an open flat ground, and record the height h of the lidar scanning plane from the ground at this time. Then, place several obstacles in front of the robot at different distances, which are higher than the installation position of the lidar sensor on the robot, and ensure that these obstacles can be photographed by the camera carried on the robot. The obstacles used need to have regular shapes, and a line is drawn as a mark at the height of h from the ground. After arranging the calibration scene, start the robot through the ROS system and publish the lidar and camera sensor data to the ROS topic. Then, display the two-dimensional point cloud data obtained by the lidar through the rviz interface provided by ROS. At this time, these lidar points should be able to clearly represent the obstacles placed in front of the robot. After confirming that the lidar data is correct, record the lidar scanning point information at the landmark positions (center, edge, quarter, etc.) on the front obstacles, and record four to five points for each obstacle on average. In this way, a total of about 25 reference points scanned by the lidar in the world coordinate system are recorded:

[0069] P lidar = [[x 1 ,y 1 ,[x2 , y 2 , … [x 25 , y 25

[0070] After that, identify the positions in the image where the above 25 lidar scan points should fall, and record the pixel positions on the image corresponding to these lidar scan points:

[0071] P camera = [[h 1 , w 1 , [h 2 , w 2 , … [h 25 , w 25

[0072] Finally, input the two sets of points P lidar and P camera along with the internal parameters of the camera into the cv2.solvePnP function to calculate the rotation matrix R and translation vector T for the transformation from the world coordinate system with the lidar sensor installation position on the robot as the origin to the camera coordinate system, which are the external parameters of the camera in the current layout.

[0073] Step 2: Projection from Lidar to Image

[0074] Lidar scan points can be projected onto the image through coordinate system conversions. After determining the external parameters of the camera, this projection process can be represented using the equations of camera calibration. The world coordinate system where the lidar scan points are located has the installation position of the lidar sensor on the robot as the origin and the forward direction of the robot as the y-axis, following the right-hand rule. Then, a certain lidar scan point P w in this coordinate system can be represented in Cartesian coordinates as (X w , Y w , Z w ). Through the external parameters of the camera, the coordinates of point P w in the camera coordinate system, P c , can be obtained, and its Cartesian coordinates are represented as (X c , Y c , Z c ):

[0075]

[0076] After that, through the focal length parameter f in the camera internal parameters, the three-dimensional coordinates P c in the camera coordinate system can be converted to the two-dimensional coordinates P p in the image coordinate system, and its coordinates are (x, y):

[0077] ​​

[0078] Finally, the points on the image coordinate system need to be converted to the pixel coordinate system. The pixel coordinate system and the image coordinate system of the camera are both on the imaging plane, but their respective origins and measurement units are different. The origin of the image coordinate system is the intersection point of the camera optical axis and the imaging plane, usually the midpoint of the imaging plane, while the origin of the pixel coordinate system is generally in the upper left of the midpoint of the imaging plane. In addition, the unit of the image coordinate system is millimeters, which is a physical unit, while the unit of the pixel coordinate system is pixels, which are discrete points arranged in rows and columns. Therefore, the conversion of point P p to the point p with coordinates (u, v) in the pixel coordinate system is as follows:

[0079]

[0080] where dx and dy represent the horizontal and vertical widths corresponding to one pixel, and u 0 and v 0 represent the optical center coordinates of the camera in the pixel plane. Combining the above three formulas, the conversion formula from the world coordinate system to the pixel coordinate system can be obtained:

[0081]

[0082] By using formula (4) to convert each lidar scan point, all the lidar points within the camera's field of view can be projected onto the camera image. Figure 5 The effect of lidar point projection is shown. The green dots in the figure are the projection positions of the lidar data points.

[0083] Step 3: Correction of lidar data

[0084] For the lidar points projected onto the image, if they fall within the recognition box of an obstacle in the image, it is considered that they fall within the range that can represent the obstacle. For all the lidar points that fall within the same recognition box, it is considered that they can represent the same obstacle. After that, it is necessary to correct all the lidar points that fall within the obstacle range and record the correction information for them. Specifically as follows:

[0085] ⑴ For the lidar points L obstacle that fall on the same obstacle, record their serial numbers among all the lidar points on the image, and uniformly correct them to the minimum distance value among these lidar points;

[0086] ⑵ After the above processing, multiply all the lidar points that are unified to the minimum value within the obstacle range by a coefficient k to obtain the final corrected lidar data L correct . The overall formula is as follows:

[0087] Lcorrect = k * min(L obstacle ) ⑸

[0088] Among them, the coefficient k is set according to different obstacle categories. The easier it is to collide, the smaller the value. For example, it can be set to 0.9 for a table or a standing person, and 0.8 for a cone or a pedestrian, etc.

[0089] Step 4: Reprojection from Image to LiDAR

[0090] After determining which LiDAR data need to be corrected, although the obstacle information at the LiDAR scanning position has been obtained, the pixel points corresponding to the LiDAR data that need to be corrected are currently obtained. To convert the corrected LiDAR data into usable sensor data, these pixel points need to be reprojected back to the LiDAR points in the world coordinate system. Although a coordinate point P in three dimensions can be determined through Step 2 w corresponding to a pixel point p on an image, but conversely, the corresponding coordinate point in three dimensions cannot be directly determined from the pixel point on the image. This is because when performing reverse calculation on the formula, the Z on the left side c is unknown. To solve this problem, the formula needs to be derived from another perspective. First, organize formula (4) and abbreviate it as:

[0091]

[0092] Among them, K represents the internal parameter matrix of the camera. Further, place the point (X w , Y w , Z w ) in the world coordinate system to the left side, then formula (6) can be transformed into:

[0093]

[0094] To further simplify and calculate the above formula (7), construct two matrices Mat 1 and Mat 2 , and let:

[0095] Mat 2 = R -1 T

[0096] Among them, all elements constituting Mat 1 and Mat 2 are known. Since what needs to be calculated is Z c , so only the terms in the third row of formula (7) need to be calculated, then there is:

[0097]

[0098] Among them, Z w is the height value of a point in the world coordinate system. Since the world coordinate system takes the installation position of the lidar sensor on the robot as the origin, and a single-line lidar only scans a plane at a fixed height, so Z w must be 0. After calculating the value of Z c , substituting it into formula (7), the corresponding points in the world coordinates can be calculated based on the pixel coordinates obtained from the image. Then, the lidar data after reprojection is corrected according to the correction information recorded in step three. After publishing the corrected lidar data, it is used to replace the original lidar data within the field of view of the camera directly in front of the robot.

[0099] So far, all the content of the sensor fusion method for the single-line lidar and the color camera is completed. It enhances the ability of the lidar to express obstacles through image information and has good practical applicability.

[0100] The following is a feasible actual deployment plan to demonstrate the method of the present invention through a specific embodiment:

[0101] In this embodiment, after deeply understanding various types of obstacles that may appear in the indoor application scenario of the robot, a training data set with annotation data is made. The content of this data set comes from the pictures collected and annotated in the simulation scenario and the real scenario. For the indoor scenario, the obstacles are selected as common pedestrians, tables, chairs, and cones in the indoor environment, and the annotation data set is manually annotated after collection. Set the number of training batches to 64 and the number of training epochs to 2000. Then set the pre-trained model to start training the YOLOv8 network model. During the training process, monitor various indicators of the model and keep training until the model converges. After the model converges, use a variety of evaluation indicators and methods to evaluate the training effect, such as accuracy, recall rate, etc., to determine whether the model meets the expected requirements. If the accuracy and other data of the model cannot meet the expected requirements, it is necessary to modify the number of training epochs, batches, or data set and retrain. After training, connect the robot to the ROS system and publish the sensor topic, write a script to subscribe to the image topic and use the trained YOLOv8 network model for object detection after processing, and publish the detection results through the ROS topic, and print and check whether the topic content is correct.

[0102] After the object detection part is completed, a calibration scene is arranged for the joint calibration of the camera and the lidar. 25 groups of lidar point data in the world coordinate system within the camera's field of view and the corresponding pixel point positions in the image are collected. This set of points and the internal parameters of the camera are passed into the cv2.solvePnP function to calculate the external parameters from the camera to the lidar. Then, using the internal and external parameters of the camera, the lidar data is projected onto the image, and the image with the lidar point projection superimposed is displayed through the cv2.imshow function to check whether the projection result meets the expectations and whether it can represent obstacles at different distances. Then, according to the subscribed detection result topic, the lidar point data falling within the object detection box of the obstacle is corrected. For all the lidar points within the same detection box, they are all corrected to the point closest to the robot among these points, and a coefficient representing the urgency of the obstacle or the maximum volume occupied longitudinally is multiplied according to the type of the obstacle. Finally, after correcting the lidar points falling within the obstacle range, all the lidar point data on the image is reprojected back to the world coordinate system. The original lidar data and the corrected lidar data are simultaneously displayed through rviz to check whether the corrected lidar data is correctly corrected according to the obstacle information. Figure 6 Shows the application effect of the method proposed in the present invention in the simulation environment. In each subfigure, (a) and (b) are pedestrians and table obstacles in the simulation environment, (c) and (d) are the recognition effects of the trained YOLOv8 model, and (e) and (f) show the lidar data before and after correction. The red point cloud is the original lidar data, and the white point cloud is the corrected lidar data. Figure 7 The application effect of the method proposed in the present invention in the real environment. In each subfigure, (a) and (b) are the recognition effects of the trained YOLOv8 model on pedestrians at close range in the real environment, and (c) and (d) show the lidar data before and after correction for pedestrian obstacles in the real environment. The red point cloud is the original lidar data, and the white point cloud is the corrected lidar data. It can be seen that the method proposed in the present invention has basically the same effect in the simulation environment and the real environment and can be conveniently migrated.

[0103] In summary, the sensor information fusion method based on a single-line lidar and a camera can better adapt to various scenarios and improve the ability of an autonomous mobile robot to perceive obstacles using the lidar.

Claims

1. A sensor information fusion method based on a single-line laser radar and a camera, characterized in that: Part 1: Obstacle Information Acquisition Based on Deep Learning Object Detection Obstacle information in front of the robot is obtained based on the monocular color camera on the robot. Part II: Projection transformation and information fusion between LiDAR and images First, the lidar points are projected onto the image acquired by the camera, and the corresponding lidar points are corrected according to the obstacle information on the image. Secondly, the corrected lidar points are reprojected back to the coordinate system of the original lidar points, so that the lidar can express the maximum outline and longitudinal volume of the obstacles that were originally imperceptible. Finally, the corrected lidar data is used to replace the original lidar data within the camera's field of view.

2. The method according to claim 1, characterized in that: The obtaining of obstacle information in front of the robot includes: Step 1: Camera image acquisition and processing, including: the camera on the robot acquires and publishes a video stream, shrinks the image to a lower resolution without changing its aspect ratio, and changes the color channel of the shrunken image from BGR to RGB; Step 2: training and applying the YOLOv8 network model, including: inputting the image obtained by processing into the pre-trained YOLOv8 network model to obtain the target information required in the input image, and publishing this information into a ROS topic as the YOLOv8 network output for subsequent use.

3. The method according to claim 2, characterized in that: The target information output by the YOLOv8 network includes confidence, the boundary of the target in the input image, and the target type.

4. The method according to claim 2, characterized in that: The pre-trained YOLOv8 network model is trained through the following steps: (1) determining the training data set according to the specific usage scenario of the robot; (2) in order to enhance the accuracy of model training, the training data set needs to include data points under special circumstances, that is, data under various non-standard conditions, including: extremely close obstacles, long-distance blurred targets, obstacles blocking each other, etc.

5. The method according to claim 1, characterized in that: The second part is the joint calibration of the single-line laser radar and the camera, which includes the following steps: Step 1: Joint calibration of camera and lidar: First, assume the intrinsic parameters of the camera are known and arrange the calibration scene; then, start the robot through the ROS system and publish the lidar and camera sensor data to the ROS topic, and display the two-dimensional point cloud data obtained by the lidar; record the lidar scanning point information that falls on the landmark position of the obstacle in front; identify the position where the above lidar scanning points should fall in the image, and record the pixel positions on the image corresponding to these lidar scanning points; finally, calculate the rotation matrix and translation vector from the world coordinate system with the lidar sensor installation position as the origin to the camera coordinate system, that is, the external parameters of the camera under the current layout; Step 2: Project the laser radar scanning points onto the image through the conversion between coordinate systems. After determining the external parameters of the camera, the projection process is expressed using the camera calibration equation; (1) The coordinates of the laser radar scanning points in the camera coordinate system can be obtained through the external parameters of the camera; (2) Then, the three-dimensional coordinates in the camera coordinate system are converted to the two-dimensional coordinates in the image coordinate system through the focal length parameters in the camera intrinsic parameters; (3) Finally, the points in the image coordinate system are converted to the pixel coordinate system; (4) Combine the above three results (1)-(3) to obtain the conversion from the world coordinate system to the pixel coordinate system; (5) After converting each laser radar scanning point, all laser radar points within the camera field of view can be projected onto the camera image; Step 3: Correct the LiDAR data, correct all LiDAR points that fall within the obstacle range, and record their correction information; Step 4: Reprojection of image to lidar. If the corrected lidar data is to be converted into usable sensor data, it is also necessary to reproject these pixel points back to lidar points in the world coordinate system. Then, the reprojected lidar data is corrected according to the correction information recorded in step 3. After the corrected lidar data is released, it is used to replace the original lidar data within the field of view of the camera directly in front of the robot.

6. The method according to claim 5, characterized in that: The conversion formula from the world coordinate system to the pixel coordinate system is: Among them, f x ,f y ,u0,v0 are the intrinsic parameters of the camera, representing the focal length and optical center coordinates of the camera; R,T are the extrinsic parameters of the camera, representing the rotation matrix and translation vector.

7. The method according to claim 5, characterized in that: To reproject the pixel points back to the laser radar points in the world coordinate system, it is necessary to derive the formula from another perspective: First, sort out formula (4) and abbreviate it as: Where K represents the intrinsic parameter matrix of the camera; further, the point (X w ,Y w ,Z w ) is placed on the left side, then formula (6) can be transformed into: In order to further simplify and calculate the above formula (7), two matrices Mat1 and Mat2 are constructed, and the following is made: Among them, all the elements that make up Mat1 and Mat2 are known; since what needs to be calculated is Z c , so we only need to calculate the items in the third row of formula (7), then: Among them, Z w is the height of the point in the world coordinate system. Since the world coordinate system is based on the installation position of the laser radar sensor on the robot, and the single-line laser radar only scans a plane at a fixed height, Z w The value of must be 0. c After obtaining the value of , substitute it into formula (7) and the corresponding point in world coordinates can be calculated based on the pixel coordinates obtained from the image.