Target detection method and system based on monocular camera and triangulation ranging radar IOU matching fusion

By fusing information between a monocular camera and a 3D triangular range-finding radar and combining the YOLOv4-tiny algorithm, the problem of traditional sensors being difficult to accurately detect targets in complex environments is solved, and the precise detection and positioning of targets is achieved, and the detection accuracy and robustness are improved.

CN120047666APending Publication Date: 2025-05-27HEBEI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081333.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In complex living environments, traditional vision sensors and radar sensors find it difficult to accurately detect targets in indoor environments, especially when there are many pedestrians and debris, and it is difficult to accurately locate the point clouds associated with the targets.

Method used

By fusing information between a monocular camera and a 2D triangular rangefinder radar, performing point cloud projection and clustering, and combining the lightweight YOLOv4-tiny algorithm, accurate detection of the target is achieved.

Benefits of technology

This method can achieve accurate detection and positioning of goals in complex environments, improve the positive detection rate of point cloud extraction, effectively remove interference caused by debris in the environment, and has high accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047666A_ABST
    Figure CN120047666A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and system based on monocular camera and triangulation ranging radar IOU matching fusion, and the method comprises the following steps: S1, obtaining camera coordinate system parameters, and converting the camera coordinate system parameters into pixel coordinate system parameters; s2, associating the laser radar point cloud with the pixel coordinate system, and performing external parameter calibration; s3, carrying out point cloud clustering and projection on the point cloud calibrated by the external parameters; and S4, performing IOU (Input / Output Unit) matching fusion on the projected point cloud and the YOLOv4-tiny recognition frame information. The method has the advantages that interference of sundries in the environment can be effectively reduced by adding a clustering method through a multi-sensor fusion mode, the method has the advantages of being high in adaptability and stable in effect, the target point cloud is effectively extracted from the environment point cloud, and target detection is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mobile robot positioning detection technology, and more specifically to a target detection method and system based on IOU matching and fusion of a monocular camera and a triangulated ranging radar. Background Art

[0002] At present, with the rapid development of robot technology, the use of robots is increasing day by day. The application fields of robots include but are not limited to industry, service industry, medical industry, and smart agriculture. Therefore, it is increasingly important to improve the robot's ability to detect targets in the environment. Traditional visual sensors can help identify targets, but cannot provide distance information from the target. A single radar can only obtain distance information from the surrounding environment and cannot distinguish specific targets, which contains certain hidden dangers in human-computer interaction and collaboration. More specifically, relying solely on lidar or ultrasonic sensors to obtain distance information has the problem of lack of visual information and target distinction, which has great limitations. However, the existing fusion detection technology cannot better implement accurate detection in complex and mobile environments. More specifically, in indoor environments with many pedestrians and debris, it is difficult to accurately locate the point cloud associated with the target.

[0003] Therefore, in view of the above problems, how to provide a method that can be applied to complex living environments and detect the target position more accurately is an issue that technical personnel in this field urgently need to solve. Summary of the invention

[0004] In view of this, the present invention provides a target detection method and system based on IOU matching fusion of a monocular camera and a triangulated ranging radar. The present invention fuses information of a monocular camera and a 2D triangulated ranging radar, and after point cloud projection and clustering, combined with a lightweight YOLOv4-tiny algorithm, can quickly and accurately find the target object in the surrounding environment and achieve accurate detection of the target. The advantage of this method and system is that it integrates multiple environmental perception sensors, making the final information obtained richer and more accurate. At the same time, the integration of the clustering algorithm can well eliminate the interference caused by debris in the environment. Combined with the YOLOv4-tiny algorithm, by training different models, accurate detection and positioning of any target in a complex environment can be achieved.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: A target detection method based on IOU matching and fusion of a monocular camera and a triangulated ranging radar specifically comprises the following steps: S1, obtaining camera coordinate system parameters, converting the camera coordinate system parameters into image coordinate system parameters, translating the image coordinate system parameters and converting the world coordinate system parameters into pixel coordinate system parameters; S2. Associate the lidar point cloud with the pixel coordinate system parameters for extrinsic parameter calibration; S3. Cluster and project the point cloud after the extrinsic parameter calibration; S4. Perform IOU matching and fusion on the projected point cloud and the YOLOv4-tiny recognition box information.

[0006] Preferably, the specific operation of step S1 is as follows: S11. Place the calibration board in the camera's field of view and take images at multiple angle positions and locations; S12. Convert the world coordinate system to the camera coordinate system parameters through

[0007] where 、 、 are the coordinates in the camera coordinate system; R is the rotation matrix; t is the translation matrix; S13. Convert the camera coordinate system to the image pixel coordinate system parameters ; S14. The pixel coordinate system parameters .

[0008] Furthermore, use the camera_calibration calibration package for calculation to obtain the internal parameter data of the camera; the internal parameter data of the camera is obtained using the Zhang Zhengyou calibration method; the S13 formula is obtained from the S12 formula through the similar triangle theorem; the image coordinate system parameters are translated to pixel coordinate system parameters

[0009] and the pixel coordinate system parameter formula is combined with formula

[0010] to obtain the pixel coordinate system parameters.

[0011] Furthermore, the relative positions of the camera and the lidar remain fixed, and solvePnP is solved through the solvePnP function in OpenCV; the image coordinate system parameters are a two-dimensional rectangular coordinate system with the origin on the optical axis of the lens and the distance to the optical center equal to the focal length of the camera; the pixel coordinate system parameters are a two-dimensional rectangular coordinate system with the origin located at the upper left corner of the image and the two axes parallel to the image coordinate system.

[0012] Preferably, the lidar publishes in polar coordinate form

[0013] where Represents the i-th point, which is the distance of this point relative to the radar, represents the radian of this point relative to the front of the radar; through the formula

[0014] the lidar point cloud information in the Cartesian coordinate system is obtained, where and are the X-axis and Y-axis coordinate system parameters respectively.

[0015] Furthermore, the camera coordinate system parameters are a three-dimensional rectangular coordinate system with the optical center of the camera as the origin and the optical axis of the camera lens as the z-axis; the image coordinate system parameters are a two-dimensional rectangular coordinate system, with the origin on the optical axis of the lens and the distance to the optical center equal to the focal length of the camera; the pixel coordinate system is a two-dimensional rectangular coordinate system, with the origin at the upper left corner of the image and the two axes parallel to the image coordinate system.

[0016] Preferably, the lidar point cloud in S2 is data in a two-dimensional coordinate system with the center of the radar as the origin, and the conversion process is real-time data publishing.

[0017] Preferably, after the lidar point cloud is obtained, ten groups of data are calibrated using the camera_calibration calibration package.

[0018] Furthermore, the ten groups of data and the camera internal parameter matrix are input into the solvePnP function for solution to obtain the rotation matrix R and the translation matrix t.

[0019] Preferably, the specific operation of step S2 is as follows: S21. Set a point-taking flag, and the point-taking flag is at the same height as the radar; S22. Use a visualization tool to subscribe to the topic of the radar; S23. Call the OpenCV library to subscribe to the image topic; S24. Use the imshow function to display the RGB image.

[0020] Preferably, the specific operation of S3 is as follows: S31. Create two lists LIST1 and LIST2 to save the point cloud after clustering and the point cloud for real-time processing respectively; S32. Judge the distance between point clouds and compare according to a set threshold; S33. If it is less than the threshold, save it to LIST2. If it is greater than the threshold, save LIST2 to LIST1 and clear LIST2.

[0021] Preferably, the specific operation of step S4 is as follows: S41. Synchronize the time of camera data and radar data; S42. Cluster the target point cloud data by inter-point distance and output the target point cloud data; S43. Compare the camera coordinate system parameters with the radar point cloud data and extract the target point cloud information.

[0022] Furthermore, use the ApproximateTime synchronization policy of the sync_policies module.

[0023] Preferably, the monocular camera finds the target through the YOLOv4-tiny algorithm, and the index value only extracts the target point cloud of the relevant point cloud.

[0024] Preferably, it includes a visual information acquisition module, a radar information acquisition module, a positioning module, and an information fusion module; The visual information acquisition module (1) is a monocular camera. The visual information acquisition module (1) acquires data and outputs 2D recognition box information through YOLOv4-tiny (2); The radar information acquisition module (7) is a triangulation radar. The radar information acquisition module (7) acquires radar point cloud information and outputs coordinate transformation data through point cloud clustering (8). The point cloud clustering creates two lists, LIST1 and LIST2, to save the clustered point cloud and the real-time processed point cloud respectively; The positioning module (4) subscribes to the relative position through the tf2 software package, and the relative position information is transformed into the 2D map coordinate system through the 3D point cloud information; The information fusion module (5) synchronizes the 2D recognition box information and the coordinate transformation data in time for data fusion, projects the point cloud after the point cloud clustering into the image coordinate system, and performs matching fusion through two-dimensional IOU to output the target position (9).

[0025] Furthermore, the / scan topic published by the radar information acquisition module and the / bounding_boxes topic published by the darknet-ros software package are jointly input into sync_policies to achieve time synchronization.

[0026] Furthermore, The threshold is

[0027] where is a constant for suppressing noise, and r i is the distance from the i-th point to the sensor.

[0028] Furthermore, the data of the radar information acquisition module is compared with the visual information acquisition module to obtain the specific target point cloud.

[0029] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a target detection method and system based on IOU matching and fusion of a monocular camera and a triangulation ranging radar, which can be applied to complex living and working environments. The advantages of this method and system are that multiple environmental perception sensors are fused, making the finally obtained information richer and more accurate. At the same time, a clustering algorithm is fused, which can effectively eliminate the interference caused by sundries in the environment. At the same time, it can effectively reduce costs in the wide application of future robots and support national construction.

[0030] Secondly, the invention can accurately extract the point cloud associated with the target. Compared with the detection method without clustering and matching, the positive detection rate of point cloud extraction has increased by 35.3%. It can effectively remove the interference caused by sundries in the environment, has high accuracy and robustness, meets the needs of the unmanned vehicle to accurately find the target, and provides a strong guarantee for the subsequent real-time following of the unmanned vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0032] Figure 1 The drawings are schematic diagrams of the target detection method based on IOU matching and fusion of a monocular camera and a triangulation ranging radar provided by the present invention; Figure 2 The drawings are schematic diagrams of the environmental point cloud information of the 2D lidar provided by the present invention; Figure 3 The drawings are schematic diagrams of the joint calibration provided by the present invention; Figure 4 The drawings are schematic diagrams of the quantitative analysis results provided by the present invention; Figure 5 The drawings are schematic diagrams of the target detection system based on IOU matching and fusion of a monocular camera and a triangulation ranging radar provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] Embodiment 1 The embodiment of the present invention discloses a target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar. Figure 1 First, obtain the camera coordinate system parameters, convert the camera coordinate system parameters into image coordinate system parameters, translate the image coordinate system parameters and convert the world coordinate system parameters into pixel coordinate system parameters. The radar coordinate system of the 2D radar is a two-dimensional coordinate system with the center of the radar as the origin. The point cloud data in the polar coordinate system of the topic released by the laser radar is converted. The more specific operation is: place the calibration plate in the camera's field of view and change the angle position and position multiple times to shoot the image; then convert the world coordinate system into the camera coordinate system parameters, through the formula

[0035] in , , is the coordinate in the camera coordinate system; R is the rotation matrix; t is the translation matrix; then the camera coordinate system is converted to the image coordinate system parameter expression

[0036] Finally, the pixel coordinate system parameters are obtained as follows:

[0037] After transformation, the radar point cloud information in the Cartesian coordinate system is obtained to prepare for calibration.

[0038] Secondly, the laser radar point cloud is associated with the pixel coordinate system parameters, external parameter calibration is performed, and the data is aligned so that the data of both can be processed and analyzed in the same coordinate system, such as Figure 3 More specifically, the operations are: setting a point flag, the point flag is at the same height as the radar; using a visualization tool to subscribe to the topic of the radar; calling the OpenCV library to subscribe to the image topic; and using the imshow function to display the RGB image.

[0039] Next, the point cloud calibrated by the external parameters is clustered and projected. More specifically, two lists LIST1 and LIST2 are created to store the clustered point cloud and the real-time processed point cloud respectively; the distance between the point clouds is determined and compared according to the set threshold. If it is less than the threshold, it is stored in LIST2; if it is greater than the threshold, LIST2 is stored in LIST1 and LIST2 is cleared.

[0040] Finally, the projected point cloud is fused with the YOLOv4-tiny recognition box information by IOU matching, and YOLOv4-tiny is used to detect the target. Figure 4After the algorithm identifies the target, it processes the transmitted detection results to obtain information such as the coordinates, size, confidence level, and category of the target recognition box, which is used for subsequent combination with the point cloud information.

[0041] Furthermore, a mobile robot is used as the platform, with a CPU being a quad-core ARM Cortex_A57 MPCore processor. The platform uses a monocular camera with a frame rate of 30 Hz and a focal length of 2.8 mm, a triangular ranging radar with a scanning frequency of 10 Hz and a ranging range of 16 meters, and is equipped with an IMU inertial measurement unit, etc. For the selected experimental scenario, there are many sundries in this environment, which will generate a large number of small point cloud clusters for interference, and can better evaluate the stability of this algorithm. Pedestrians are selected as the recognition target in the indoor environment, and the visualization tool rviz is started for observation of the detected point cloud distribution. Record the point cloud information and camera information with a duration of one minute as the experimental data set, with a frame rate of 10 frames, and a total of 600 frames of information are extracted. The recognition success rate of the algorithm in this paper under this data set is statistically calculated. Three indicators, namely the positive detection rate, false detection rate, and missed detection rate, are used to evaluate the performance of this algorithm, and the calculation formulas are as follows:

[0042] Among them, is the total number of frames in which pedestrians appear during the experiment; is the number of frames in which the pedestrian point cloud is correctly extracted; is the number of frames in which the extraction result is misjudged; is the number of frames in which pedestrians are not detected.

[0043] It can be concluded that in an open environment, the difference between algorithm A and algorithm B is small. However, when it comes to a complex environment with more sundries, the positive detection rate of algorithm A drops significantly, while algorithm B can better remove the point cloud generated by the sundries closer to the target and improve the accuracy. The comparison results are shown in the following table: Table 1 Comparison between algorithm A and algorithm B From the data in the table, it can be seen that the positive detection rate of algorithm B proposed in this paper is about 89.4%, which is 35.3% higher than that of algorithm A without using clustering and the two-dimensional IOU method. The accuracy of pedestrian point cloud extraction has been significantly improved, and the average time consumption for processing each frame of data has a small difference and is within an acceptable range.

[0044] Example 2 This embodiment of the present invention discloses a target detection system based on IOU matching and fusion of a monocular camera and a triangular ranging radar. The target detection method based on IOU matching and fusion of a monocular camera and a triangular ranging radar in Example 1 can be applied to this system. As Figure 5, the system includes a visual information acquisition module, a radar information acquisition module, a positioning module, and an information fusion module; the visual information acquisition module (1) is a monocular camera, and the visual information acquisition module (1) acquires data and outputs 2D recognition box information through YOLOv4-tiny (2). The radar information acquisition module (7) is a triangulation radar. The radar information acquisition module (7) acquires radar point cloud information and outputs coordinate transformation data through point cloud clustering (8). Point cloud clustering creates two lists, LIST1 and LIST2, to save the clustered point cloud and the point cloud being processed in real time respectively; the positioning module (4) subscribes to the relative position through the tf2 software package, and the relative position information is transformed into the 2D map coordinate system through the 3D point cloud information; the information fusion module (5) synchronizes the 2D recognition box information and the coordinate transformation data in time for data fusion, projects the point cloud after point cloud clustering into the image coordinate system, and performs matching fusion through 2D IOU to output the target position (9).

[0045] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description of the method part for related parts.

[0046] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target detection method based on IOU matching fusion of monocular camera and triangulated ranging radar, characterized in that: The specific steps include: S1, obtaining camera coordinate system parameters, converting the camera coordinate system parameters into image coordinate system parameters, translating the image coordinate system parameters and converting the world coordinate system parameters into pixel coordinate system parameters; S2, associating the laser radar point cloud with the pixel coordinate system parameters to perform external parameter calibration; S3, clustering and projecting the point cloud calibrated by the external parameters; S4. Perform IOU matching and fusion on the projected point cloud and YOLOv4-tiny recognition frame information.

2. The target detection method based on IOU matching fusion of monocular camera and triangulated ranging radar according to claim 1 is characterized in that: The specific operations of step S1 are: S11, placing the calibration plate in the camera field of view and changing the angle position and position multiple times to capture images; S12, converting the world coordinate system into the camera coordinate system parameters, in , , is the coordinate in the camera coordinate system; R is the rotation matrix; t is the translation matrix; S13, the camera coordinate system is converted into the image coordinate system parameters ; S14, obtaining the pixel coordinate system parameters 。 3. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 2 is characterized in that: After the lidar point cloud is acquired, ten sets of data are calibrated using the camera_calibration package.

4. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 1 is characterized in that: Laser radar releases polar coordinate system format in represents the i-th point, is the distance of the point relative to the radar, Indicates the arc angle of the point relative to the front of the radar; Obtain the laser radar point cloud information in a Cartesian coordinate system, where and They are the X-axis and Y-axis coordinate system parameters respectively.

5. The target detection method based on IOU matching fusion of monocular camera and triangulated ranging radar according to claim 1 is characterized in that: The laser radar point cloud in S2 is two-dimensional coordinate system data with the center of the radar as the origin, and the conversion process is real-time data release.

6. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 1 is characterized in that: The specific operations of step S2 are: S21, setting a point mark, wherein the point mark is at the same height as the radar; S22. Subscribe to the topic of the radar using a visualization tool; S23, call OpenCV library to subscribe to the image topic; S24. Use the imshow function to display the RGB image.

7. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 1 is characterized in that: The specific operations of S3 are: S31, create two lists LIST1 and LIST2 to respectively store the point cloud completed by clustering and the point cloud processed in real time; S32, determining the distance between point clouds and performing comparison according to a set threshold; S33, if it is less than the threshold, store it in LIST2; if it is greater than the threshold, store LIST2 in LIST1 and clear LIST2.

8. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 1, characterized in that: The specific operations of step S4 are: S41, time synchronization of camera data and radar data; S42, clustering the target point cloud data based on the distance between points and outputting the target point cloud data; S43, comparing the camera coordinate system parameters with the radar point cloud number and extracting target point cloud information.

9. The target detection method based on IOU matching fusion of a monocular camera and a triangulated ranging radar according to claim 8 is characterized in that: The monocular camera searches for the target through the YOLOv4-tiny algorithm, and the index value only extracts the target point cloud of the relevant point cloud.

10. A target detection system based on IOU matching and fusion of monocular camera and triangulated ranging radar, characterized in that: It includes visual information acquisition module, radar information acquisition module, positioning module and information fusion module; The visual information acquisition module (1) is a monocular camera, and the visual information acquisition module (1) acquires data and outputs 2D recognition frame information through YOLOv4-tiny (2); The radar information acquisition module (7) is a triangulation ranging radar, the radar information acquisition module (7) acquires radar point cloud information, outputs coordinate transformation data through point cloud clustering (8), and the point cloud clustering creates two lists LIST1 and LIST2 to respectively store the clustered point cloud and the real-time processed point cloud; The positioning module (4) subscribes to the relative position through the tf2 software package, and the relative position information is transformed into a 2D map coordinate system through 3D point cloud information; The information fusion module (5) synchronizes the time of data fusion of the 2D recognition frame information and the coordinate transformation data, projects the point cloud after point cloud clustering to the image coordinate system, and outputs the target position (9) through matching and fusion through two-dimensional IOU.