Visual target position detection method and apparatus

By combining camera and LiDAR data, and utilizing point cloud-image single mapping matrix and point cloud clustering, the problem of insufficient 3D depth estimation accuracy in visual target detection is solved, enabling more accurate obstacle recognition and avoidance, and improving the safety and reliability of autonomous driving.

CN121564329BActive Publication Date: 2026-05-26CHINA COAL CONSTR GRP CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA COAL CONSTR GRP CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing visual object detection methods only operate in a two-dimensional image plane, resulting in poor accuracy in three-dimensional object depth estimation and an inability to accurately calculate the distance to objects, which affects the safety and reliability of autonomous driving.

Method used

By combining two-dimensional images captured by cameras on the vehicle and three-dimensional laser point clouds captured by lidar, the three-dimensional laser point clouds are projected onto the two-dimensional images through a calibrated point cloud-image single mapping matrix. The three-dimensional position of the visual target is then calculated by combining point cloud clustering and motion velocity estimation.

Benefits of technology

It improves the accuracy of visual target position detection, ensuring that the autonomous driving system can more accurately identify and avoid obstacles, thereby improving driving safety and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564329B_ABST
    Figure CN121564329B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for visual target position detection, belonging to the field of autonomous driving technology. The method includes: S1: acquiring a two-dimensional image captured by a camera on a vehicle and a corresponding three-dimensional laser point cloud captured by a lidar sensor on the vehicle; S2: detecting visual targets on the two-dimensional image using a visual detection algorithm to obtain a two-dimensional target region corresponding to each visual target; S3: projecting each three-dimensional laser point from the three-dimensional laser point cloud onto the two-dimensional image using a calibrated point cloud-image single mapping matrix to obtain two-dimensional projection points; S4: for each two-dimensional target region, determining the set of corresponding three-dimensional laser points based on the two-dimensional projection points within the two-dimensional target region to obtain a set of three-dimensional laser points corresponding to the two-dimensional target region; S5: comprehensively calculating the position of the visual target based on the two-dimensional target region of the visual target and its corresponding set of three-dimensional laser points. This invention improves the accuracy of visual target position detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a visual target position detection method and apparatus. Background Technology

[0002] Cameras can capture high-resolution images, providing rich color, texture, and shape information. Visual target detection analyzes the image or video data captured by the camera to identify, locate, and classify targets.

[0003] Visual object location estimation plays a crucial role in autonomous driving, helping vehicles perceive their surroundings and make decisions. In autonomous driving, vehicles need real-time awareness of their environment, including roads, other vehicles, pedestrians, and buildings. Knowing the distances to these objects is essential, as it determines the vehicle's speed and path. Visual object depth estimation technology infers object distances by analyzing camera images. This technology helps vehicles identify obstacles ahead and take timely avoidance measures, ensuring driving safety and stability. Furthermore, depth estimation can help vehicles plan routes to avoid obstacles and terrain changes. Visual object depth estimation is vital in autonomous driving, enabling vehicles to drive accurately, stably, and efficiently, thereby improving the safety and reliability of autonomous driving.

[0004] Therefore, the technology of calculating the world position of a target based on two-dimensional visual target information is particularly important. However, visual detection is only performed on the image plane, and the detected image target only has two-dimensional pixel information, which results in poor accuracy for three-dimensional target depth estimation. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a visual target position detection method and apparatus, which improves the accuracy of visual target position detection.

[0006] The technical solution provided by this invention is as follows:

[0007] A method for visual target location detection, the method comprising:

[0008] S1: During vehicle operation, acquire two-dimensional images captured by cameras on the vehicle and corresponding three-dimensional laser point clouds captured by lidar on the vehicle.

[0009] S2: Visual targets on the two-dimensional image are detected using a visual detection algorithm to obtain the two-dimensional target region corresponding to each visual target;

[0010] S3: Using the calibrated point cloud-image single mapping matrix, project each three-dimensional laser point in the three-dimensional laser point cloud onto the two-dimensional image to obtain two-dimensional projection points;

[0011] In this process, the point cloud-image single mapping matrix is ​​obtained by calculating the correspondence between the feature corner points of the 3D laser point cloud in the set calibration scene and the corresponding image corner points of the 2D image, and then calibrating the point cloud-image single mapping matrix.

[0012] S4: For each two-dimensional target region, determine the set of corresponding three-dimensional laser points based on the two-dimensional projection points within the two-dimensional target region to obtain the set of three-dimensional laser points corresponding to the two-dimensional target region;

[0013] S5: Calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set.

[0014] Furthermore, S4 includes:

[0015] S41: Retain the two-dimensional projection points that fall within the two-dimensional target area to obtain the projection point set;

[0016] S42: Obtain the three-dimensional laser point corresponding to each two-dimensional projection point of the projection point set, and perform point cloud clustering to obtain the three-dimensional laser point set.

[0017] Furthermore, S4 includes:

[0018] S41': Retain the two-dimensional projection points that fall within the two-dimensional target area to obtain a projection point set, and obtain a mapping point pair consisting of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point;

[0019] S42': Obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair;

[0020] S43': Using the initial mapping point pair as a reference, for each visual target, calculate the average scaling ratio based on the mapping point pair corresponding to the visual target;

[0021] S44': Remove mapping point pairs that deviate from the set scaling average and recalculate the scaling average.

[0022] S45': Based on the recalculated average scaling ratio, and the length and height of the two-dimensional target region, obtain the length and height of the target region corresponding to the visual target;

[0023] S46': Obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set;

[0024] S47': Based on the length, height, and width of the target area, obtain the 3D bounding box of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the 3D bounding box, and perform point cloud clustering to obtain the 3D laser point set.

[0025] Furthermore, prior to S4, the following steps are also included:

[0026] S31: Obtain images from several frames preceding the two-dimensional image as preceding images, and detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target;

[0027] S32: Calculate the motion speed of the visual target on the two-dimensional image plane based on the preceding target region on the preceding image, the two-dimensional target region on the two-dimensional image, and the time difference between the preceding image and the two-dimensional image;

[0028] S33: Based on the first moment corresponding to the two-dimensional image, the second moment corresponding to the three-dimensional laser point cloud, and the motion speed, infer the predicted position of the visual target on the two-dimensional image at the second moment, and set the visual target on the two-dimensional image to the predicted position.

[0029] Furthermore, point cloud clustering is performed in the following manner:

[0030] Each 3D laser point to be clustered is projected onto the xy plane, and the xy plane is rasterized.

[0031] The number of laser points projected into each grid is counted. If the number is greater than a set threshold, the grid is assigned a value of 255; otherwise, the grid is assigned a value of 0.

[0032] Connectivity search is performed on the assigned raster image to achieve point cloud clustering.

[0033] A visual target position detection device, the device comprising:

[0034] The data acquisition module is used to acquire two-dimensional images captured by the camera on the vehicle and corresponding three-dimensional laser point clouds captured by the lidar on the vehicle during the vehicle's operation.

[0035] The visual detection module is used to detect visual targets on the two-dimensional image using a visual detection algorithm, and obtain the two-dimensional target region corresponding to each visual target;

[0036] The point cloud mapping module is used to project each three-dimensional laser point of the three-dimensional laser point cloud onto a two-dimensional image using a calibrated point cloud-image single mapping matrix, so as to obtain two-dimensional projection points.

[0037] In this process, the point cloud-image single mapping matrix is ​​obtained by calculating the correspondence between the feature corner points of the 3D laser point cloud in the set calibration scene and the corresponding image corner points of the 2D image, and then calibrating the point cloud-image single mapping matrix.

[0038] The point cloud location determination module is used to determine the set of corresponding three-dimensional laser points for each two-dimensional target region based on the two-dimensional projection points within the two-dimensional target region, thereby obtaining the three-dimensional laser point set corresponding to the two-dimensional target region.

[0039] The position detection module is used to calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set.

[0040] Furthermore, the point cloud location determination module includes:

[0041] The projection point set determination unit is used to retain two-dimensional projection points that fall within the two-dimensional target area to obtain a projection point set;

[0042] A three-dimensional laser point set determination unit is used to obtain the three-dimensional laser point corresponding to each two-dimensional projection point of the projection point set, and to perform point cloud clustering to obtain the three-dimensional laser point set.

[0043] Furthermore, the point cloud location determination module includes:

[0044] The mapping point pair determination unit is used to retain the two-dimensional projection points that fall within the two-dimensional target area, obtain the projection point set, and obtain the mapping point pair composed of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point;

[0045] The mapping point pair reference determination unit is used to obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair;

[0046] The first calculation unit is used to calculate the average scaling ratio for each visual target based on the initial mapping point pair as a reference, according to the mapping point pair corresponding to the visual target.

[0047] The second calculation unit is used to remove mapping point pairs that deviate from the set scaling ratio average and recalculate the scaling ratio average.

[0048] The target region length and height acquisition unit is used to obtain the length and height of the target region corresponding to the visual target based on the recalculated average scaling ratio and the length and height of the two-dimensional target region.

[0049] The target area width acquisition unit is used to obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set;

[0050] The 3D bounding box determination unit is used to obtain the 3D bounding box of the target area based on the length, height and width of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the 3D bounding box, and perform point cloud clustering to obtain the 3D laser point set.

[0051] Furthermore, the device also includes:

[0052] The preceding target region acquisition module is used to acquire images of several frames preceding the two-dimensional image as preceding images, and to detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target.

[0053] The motion speed calculation module is used to calculate the motion speed of the visual target on the two-dimensional image plane based on the preceding target region on the preceding image, the two-dimensional target region on the two-dimensional image, and the time difference between the preceding image and the two-dimensional image.

[0054] The position prediction module is used to infer the predicted position of the visual target on the two-dimensional image at the second time based on the first time corresponding to the two-dimensional image, the second time corresponding to the three-dimensional laser point cloud, and the motion speed, and set the visual target on the two-dimensional image to the predicted position.

[0055] Furthermore, point cloud clustering is performed in the following manner:

[0056] Each 3D laser point to be clustered is projected onto the xy plane, and the xy plane is rasterized.

[0057] The number of laser points projected into each grid is counted. If the number is greater than a set threshold, the grid is assigned a value of 255; otherwise, the grid is assigned a value of 0.

[0058] Connectivity search is performed on the assigned raster image to achieve point cloud clustering.

[0059] The present invention has the following beneficial effects:

[0060] This invention combines two-dimensional images with three-dimensional laser point clouds to comprehensively calculate the position of visual targets, thereby improving the accuracy of visual target position detection. Attached Figure Description

[0061] Figure 1 This is a flowchart of the visual target position detection method of the present invention;

[0062] Figure 2 This is a schematic diagram of the visual target position detection device of the present invention. Detailed Implementation

[0063] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0064] Example 1:

[0065] This invention provides a method for visual target location detection, such as... Figure 1 As shown, the method includes:

[0066] S1: During vehicle operation, acquire two-dimensional images captured by cameras on the vehicle and corresponding three-dimensional laser point clouds captured by lidar on the vehicle.

[0067] Two-dimensional images and their corresponding three-dimensional laser point clouds should theoretically be acquired simultaneously. However, due to hardware limitations, it is generally difficult to guarantee that they are acquired strictly simultaneously. Therefore, the time difference between the two can be required to be less than a set value, generally less than 100ms.

[0068] S2: Visual targets on a two-dimensional image are detected using a visual detection algorithm to obtain the two-dimensional target region corresponding to each visual target.

[0069] This invention does not limit the visual detection algorithm; it can be any method that can be used in the prior art, as long as it can detect the visual target. The detection result of the visual target is generally represented by a two-dimensional rectangular box, the interior of which is the two-dimensional target area.

[0070] S3: Using the calibrated point cloud-image single mapping matrix, project each three-dimensional laser point in the three-dimensional laser point cloud onto the two-dimensional image to obtain two-dimensional projection points.

[0071] Visual inspection is performed only on the image plane, and the detected image target only has two-dimensional pixel information. However, the points in the laser point cloud have three-dimensional information. The general method to better combine the laser points with the image target with the three-dimensional information is to project the three-dimensional laser point cloud onto the two-dimensional image.

[0072] This projection method requires pre-calibrating the point cloud-image single-mapping matrix. Generally, a camera-laser joint calibration method can be used to obtain the point cloud-image single-mapping matrix that maps the laser point cloud to the camera plane. Calibration is usually performed under relatively stable and ideal calibration conditions. By calculating the correspondence between the feature corner points of the 3D laser point cloud and the corresponding image corner points of the 2D image (usually finding at least four pairs of matching laser point-camera image points), the one-to-one mapping from the 3D laser point cloud to the camera image plane of the 2D image is obtained, thus calibrating the point cloud-image single-mapping matrix.

[0073] After calibration, all 3D laser point clouds can be mapped to the image plane through the obtained single mapping relationship during subsequent use.

[0074] S4: For each two-dimensional target region, determine the set of corresponding three-dimensional laser points based on the two-dimensional projection points within the two-dimensional target region to obtain the three-dimensional laser point set corresponding to the two-dimensional target region.

[0075] After projection, the laser points corresponding to the target area of ​​the image are obtained based on all the points on the image plane that fall on the target area of ​​the image.

[0076] S5: Calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set.

[0077] That is, by comprehensively calculating and optimizing the results of the target area in the image and its corresponding laser points, the final three-dimensional position of the visual target is obtained.

[0078] This invention combines two-dimensional images with three-dimensional laser point clouds to comprehensively calculate the position of visual targets, thereby improving the accuracy of visual target position detection.

[0079] As an example, S4 includes:

[0080] S41: Retain the two-dimensional projection points that fall within the two-dimensional target area to obtain the projection point set.

[0081] S42: Obtain the 3D laser point corresponding to each 2D projection point in the projection point set, and perform point cloud clustering to obtain the 3D laser point set.

[0082] In actual calibration operations, joint calibration errors are inevitable and difficult to overcome. This is because the laser equipment used in practice generally has low resolution, while the image resolution is very high. Therefore, for corner points in the image, it is usually difficult to find an absolutely corresponding laser point in the laser point cloud; typically, the laser point that most closely approximates the image point is selected. Furthermore, the accuracy of laser equipment varies with distance, so the single mapping matrix obtained from actual calibration inevitably contains errors and may not be suitable for the current scenario.

[0083] When estimating the location of targets in images detected while the autonomous vehicle is in operation, the impact of vehicle vibration is introduced, increasing the error in the established point cloud-image mapping relationship. As a result, the point cloud actually representing the target area cannot be perfectly projected onto the target area in the image; only a portion can be successfully mapped, while other interfering points that do not belong to the target are introduced. For example, when the target is a pedestrian, the pedestrian's point cloud will never completely overlap with the pedestrian image, causing the calculated target position to deviate from reality and affecting the accuracy of subsequent planning decisions.

[0084] The effect of projecting lasers onto camera images also has the following drawbacks: Due to the location of the laser installation and the relative angle between the laser and the target, the laser points of all targets, after being projected onto the image target, cannot cover the entire image target position. Furthermore, due to the influence of the camera's intrinsic and extrinsic parameter calibration accuracy, the method of relying solely on point cloud projection onto the image target bounding box will cause a large number of points that actually belong to the target to be missed, while points that do not belong to the target may be retained, causing trouble for subsequent target position calculations. This can easily lead to a deviation between the target position calculated based on the mapped laser points and the actual position.

[0085] To address the aforementioned shortcomings, this invention improves upon the use of combined laser point cloud computing to determine the actual location of the target in the image. Specifically, step S4 includes:

[0086] S41': Retain the two-dimensional projection points that fall within the two-dimensional target area to obtain the projection point set, and obtain the mapping point pair composed of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point.

[0087] S42': Obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair.

[0088] S43': Based on the initial mapping point pair, calculate the average scaling ratio for each visual target according to the mapping point pair corresponding to the visual target.

[0089] S44': Remove mapping point pairs that deviate from the average scaling ratio setting and recalculate the average scaling ratio;

[0090] S45': Based on the recalculated average scaling factor d0, and the length and height of the two-dimensional target region, obtain the length and height of the target region corresponding to the visual target;

[0091] S46': Obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set;

[0092] S47': Based on the length, height, and width of the target area, obtain the 3D bounding box of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the bounding box, and perform point cloud clustering to obtain the 3D laser point set.

[0093] This invention first identifies the laser points falling within the target bounding box region of a two-dimensional image. It then tracks the mapping point pairs from each laser point to an image pixel, forming a set of point cloud to pixel mapping point pairs (each laser point corresponds one-to-one with a pixel in the image). Next, the mapping point pair closest to the image center is used as the initial mapping point pair. For each other mapping point pair, the difference between its world position (the world coordinate system of the lidar) and its pixel position (the pixel coordinate system of the two-dimensional image) is calculated with respect to the initial mapping point pair. The average scaling ratio (in m / pixel) in the u and v directions of each region's image is then calculated. Finally, the average of all scaling ratios is calculated to obtain the final average scaling ratio d0. Based on the detected rectangular bounding box in the two-dimensional image, a primary target 3D bounding box is calculated using d0. The intersection of this primary target 3D bounding box with the original laser point cloud is then used to select more accurate and effective laser points. Finally, through clustering, an improved target 3D bounding box, i.e., a three-dimensional laser point set, is obtained.

[0094] Cameras can capture 30-100 images per second, providing a high-speed image stream for real-time monitoring of target position and movement. In contrast, lasers can only acquire 10 frames per second, resulting in a relatively low frame rate. Due to the different hardware characteristics of laser and camera devices, their trigger times may differ, leading to a time difference when the two devices detect the same target. While hardware trigger synchronization mechanisms can be used to ensure data capture at the same time, achieving this synchronization between cameras and lasers presents significant technical challenges. Furthermore, the inconsistency in camera and laser frequencies ensures a persistent time difference between the detection of the same target in the image and on the laser. This time difference causes a slight deviation between the laser point projection area and the target area in the image when the laser point is projected onto the target area according to the calibrated mapping. Since the maximum time difference between the laser target and the camera target is no more than 100ms, for larger targets, the projection deviation still results in a large overlapping area, allowing the approximate position of the target in the image to be estimated from the projected point cloud. However, for small targets such as pedestrians, the point cloud itself is very small. In addition, due to the deviation, the laser point cloud of the pedestrian is basically not superimposed on the image and has little or no overlap with the pedestrian target area in the image, which leads to the failure of the target position estimation in the image.

[0095] To solve the above problems, the method of the present invention further includes, before S4:

[0096] S31: Obtain images from several frames preceding the two-dimensional image as preceding images, and detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target.

[0097] S32: Calculate the motion speed of the visual target on the two-dimensional image plane based on the previous target area on the previous image, the two-dimensional target area on the two-dimensional image, and the time difference between the previous image and the two-dimensional image.

[0098] S33: Speculate the speculated position of the visual target on the two-dimensional image at the second moment according to the first moment corresponding to the two-dimensional image, the second moment corresponding to the three-dimensional laser point cloud, and the motion speed, and set the visual target on the two-dimensional image to the speculated position.

[0099] When jointly calculating the target detected by the laser and the target detected by the camera, there will inevitably be a time difference between the two devices. The present invention proposes a method for estimating the target pixel speed based on the pixel level. The idea is to consider that the pixel movement speed of the target on the image is uniform within a very short time difference. Calculate the speed of the target on u and v through the pixel positions and time of the image target in the previous few frames. Multiply the pixel speed by the time difference between the camera target and the laser target to reverse the image target to the pixel position at the laser target moment, and then perform joint calculation with the projection of the laser point cloud. Assume that the center pixel position of the image target box at time t1 is (u1, v1), and the center pixel position of the image target box at time t2 is (u2, v2). Then, at time t, the center pixel position (u, v) of the image target box: u = u2 + (t - t2) × (u2 - u1) / (t2 - t1), v = v2 + (t - t2) × (v2 - v1) / (t2 - t1); where 0 < t - t2 < 0.1s, 0 < t2 - t1 < 0.1s; The target rectangle box with the center pixel position of u, v and unchanged length and width is involved in the calculation of the laser point cloud projection.

[0100] The present invention traces back the image target to the pixel position at the laser target moment and then performs the next joint estimation, avoiding the situation that small image targets such as pedestrians have no laser points available for position estimation, and also making the image target estimation more accurate.

[0101] When the laser point cloud is projected onto the image plane and the point cloud falling within the image target area is obtained, although a large number of background point interferences are removed through the classification of laser foreground and background points, there will still be interference points, which affect the accuracy of calculating the actual position of the image target. In order to better obtain the laser points belonging to the image target, the method of point cloud clustering is generally used to classify the laser points. As is well known, laser points are three-dimensional points. For larger targets, such as trucks, etc., the more laser points are mapped to the target area. Then, clustering will inevitably consume a large amount of time, failing to meet the real-time requirement of obstacle detection for autonomous driving.

[0102] To solve this problem, the present invention performs point cloud clustering in the following manner:

[0103] 1. Project each 3D laser point to be clustered onto the xy plane, and then rasterize the xy plane.

[0104] 2. Count the number of laser points projected into each grid. If the number is greater than the set threshold, assign the grid a value of 255; otherwise, assign the grid a value of 0.

[0105] 3. Perform connected component search on the assigned raster image to achieve point cloud clustering.

[0106] To address the issue of excessive time consumption in laser point clustering, this invention proposes a clustering method that projects point clouds onto a two-dimensional grid. Each laser point is projected onto a grid in the x and y directions, with each grid corresponding to a pixel position in the projected image. If a grid contains more than a certain number of points, the corresponding pixel value is set to 255; otherwise, it is set to 0. This process is repeated to obtain a mask image (a black and white image with only 0 and 255 pixel values). Performing a connected component search on this image, which takes milliseconds, transforms the three-dimensional clustering algorithm into a two-dimensional computation. Clustering laser points in two dimensions not only simplifies the computational algorithm but also significantly improves the point cloud clustering effect, meeting the real-time requirements for target detection location estimation.

[0107] Example 2:

[0108] This invention provides a visual target position detection device, such as... Figure 2 As shown, the device includes:

[0109] Data acquisition module 1 is used to acquire two-dimensional images captured by cameras on the vehicle and corresponding three-dimensional laser point clouds captured by lidar on the vehicle during vehicle operation.

[0110] The visual detection module 2 is used to detect visual targets on a two-dimensional image using a visual detection algorithm, and obtain the two-dimensional target region corresponding to each visual target.

[0111] Point cloud mapping module 3 is used to project each three-dimensional laser point of the three-dimensional laser point cloud onto a two-dimensional image through a calibrated point cloud-image single mapping matrix to obtain two-dimensional projection points.

[0112] Specifically, by pre-calculating the correspondence between the feature corner points of the 3D laser point cloud in the set calibration scene and the corresponding image corner points of the 2D image, a one-to-one mapping from the 3D laser point cloud to the camera image plane where the 2D image is located is calculated, and a point cloud-image single mapping matrix is ​​obtained.

[0113] The point cloud location determination module 4 is used to determine the set of corresponding three-dimensional laser points for each two-dimensional target area based on the two-dimensional projection points within the two-dimensional target area, thereby obtaining the three-dimensional laser point set corresponding to the two-dimensional target area.

[0114] The position detection module 5 is used to calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set.

[0115] As an improvement, the point cloud location determination module includes:

[0116] The projection point set determination unit is used to retain the two-dimensional projection points that fall within the two-dimensional target area to obtain the projection point set.

[0117] The three-dimensional laser point set determination unit is used to obtain the three-dimensional laser point corresponding to each two-dimensional projection point of the projection point set, and to perform point cloud clustering to obtain the three-dimensional laser point set.

[0118] Alternatively, the point cloud location determination module may include:

[0119] The mapping point pair determination unit is used to retain the two-dimensional projection points that fall within the two-dimensional target area, obtain the projection point set, and acquire the mapping point pair composed of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point.

[0120] The mapping point pair reference determination unit is used to obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair.

[0121] The first calculation unit is used to calculate the average scaling ratio for each visual target based on the initial mapping point pair.

[0122] The second calculation unit is used to remove mapping point pairs that deviate from the average scaling ratio setting and to recalculate the average scaling ratio.

[0123] The target region length and height acquisition unit is used to obtain the length and height of the target region corresponding to the visual target based on the recalculated average scaling ratio and the length and height of the two-dimensional target region.

[0124] The target area width acquisition unit is used to obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set.

[0125] The 3D bounding box determination unit is used to obtain the 3D bounding box of the target area based on the length, height and width of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the 3D bounding box, and perform point cloud clustering to obtain the 3D laser point set.

[0126] As another improvement, the device of the present invention further includes:

[0127] The preceding target region acquisition module is used to acquire images from several frames preceding the two-dimensional image as preceding images, and to detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target.

[0128] The motion speed calculation module is used to calculate the motion speed of the visual target on the two-dimensional image plane based on the preceding target region on the preceding image, the two-dimensional target region on the two-dimensional image, and the time difference between the preceding image and the two-dimensional image.

[0129] The position prediction module is used to infer the predicted position of the visual target on the two-dimensional image at the second moment based on the first moment corresponding to the two-dimensional image, the second moment corresponding to the three-dimensional laser point cloud, and the motion speed, and set the visual target on the two-dimensional image to the predicted position.

[0130] In one example, point cloud clustering can be performed as follows:

[0131] Each 3D laser point to be clustered is projected onto the xy plane, and the xy plane is rasterized.

[0132] Count the number of laser points projected into each grid. If the number is greater than the set threshold, assign the grid a value of 255; otherwise, assign the grid a value of 0.

[0133] Connectivity search is performed on the assigned raster image to achieve point cloud clustering.

[0134] The apparatus provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the apparatus embodiment can be referred to the corresponding content in the aforementioned method embodiment 1. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the apparatus and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0135] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention.

Claims

1. A method for visual target position detection, characterized in that, The method includes: S1: During vehicle operation, acquire two-dimensional images captured by cameras on the vehicle and corresponding three-dimensional laser point clouds captured by lidar on the vehicle. S2: Visual targets on the two-dimensional image are detected using a visual detection algorithm to obtain the two-dimensional target region corresponding to each visual target; S3: Using the calibrated point cloud-image single mapping matrix, project each three-dimensional laser point in the three-dimensional laser point cloud onto the two-dimensional image to obtain two-dimensional projection points; In this process, the point cloud-image single mapping matrix is ​​obtained by pre-calculating the correspondence between the feature corner points of the 3D laser point cloud in the set calibration scene and the corresponding image corner points of the 2D image, calculating the one-to-one mapping from the 3D laser point cloud to the camera image plane where the 2D image is located. S4: For each two-dimensional target region, determine the set of corresponding three-dimensional laser points based on the two-dimensional projection points within the two-dimensional target region to obtain the set of three-dimensional laser points corresponding to the two-dimensional target region; S5: Calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set; S4 includes: S41': Retain the two-dimensional projection points that fall within the two-dimensional target area to obtain a projection point set, and obtain a mapping point pair consisting of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point; S42': Obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair; S43': Using the initial mapping point pair as a reference, for each visual target, calculate the average scaling ratio based on the mapping point pair corresponding to the visual target; S44': Remove mapping point pairs that deviate from the set scaling average and recalculate the scaling average. S45': Based on the recalculated average scaling ratio, and the length and height of the two-dimensional target region, obtain the length and height of the target region corresponding to the visual target; S46': Obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set; S47': Based on the length, height, and width of the target area, obtain the 3D bounding box of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the 3D bounding box, and perform point cloud clustering to obtain the 3D laser point set; Before S4, it also includes: S31: Obtain images from several frames preceding the two-dimensional image as preceding images, and detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target; S32: Calculate the motion speed of the visual target on the two-dimensional image plane based on the preceding target region on the preceding image, the two-dimensional target region on the two-dimensional image, and the time difference between the preceding image and the two-dimensional image; S33: Based on the first moment corresponding to the two-dimensional image, the second moment corresponding to the three-dimensional laser point cloud, and the motion speed, infer the predicted position of the visual target on the two-dimensional image at the second moment, and set the visual target on the two-dimensional image to the predicted position; Point cloud clustering is performed in the following manner: Each 3D laser point to be clustered is projected onto the xy plane, and the xy plane is rasterized. The number of laser points projected into each grid is counted. If the number is greater than a set threshold, the grid is assigned a value of 255; otherwise, the grid is assigned a value of 0. Connectivity search is performed on the assigned raster image to achieve point cloud clustering.

2. A visual target position detection device, characterized in that, The device includes: The data acquisition module is used to acquire two-dimensional images captured by the camera on the vehicle and corresponding three-dimensional laser point clouds captured by the lidar on the vehicle during the vehicle's operation. The visual detection module is used to detect visual targets on the two-dimensional image using a visual detection algorithm, and obtain the two-dimensional target region corresponding to each visual target; The point cloud mapping module is used to project each three-dimensional laser point of the three-dimensional laser point cloud onto a two-dimensional image using a calibrated point cloud-image single mapping matrix, so as to obtain two-dimensional projection points. In this process, the point cloud-image single mapping matrix is ​​obtained by pre-calculating the correspondence between the feature corner points of the 3D laser point cloud in the set calibration scene and the corresponding image corner points of the 2D image, calculating the one-to-one mapping from the 3D laser point cloud to the camera image plane where the 2D image is located. The point cloud location determination module is used to determine the set of corresponding three-dimensional laser points for each two-dimensional target region based on the two-dimensional projection points within the two-dimensional target region, thereby obtaining the three-dimensional laser point set corresponding to the two-dimensional target region. The position detection module is used to calculate the position of the visual target based on the two-dimensional target area of ​​the visual target and its corresponding three-dimensional laser point set. The point cloud location determination module includes: The mapping point pair determination unit is used to retain the two-dimensional projection points that fall within the two-dimensional target area, obtain the projection point set, and obtain the mapping point pair composed of each two-dimensional projection point in the projection point set and its corresponding three-dimensional laser point; The mapping point pair reference determination unit is used to obtain the two-dimensional projection point closest to the center of the two-dimensional image and its corresponding three-dimensional laser point to obtain the initial mapping point pair; The first calculation unit is used to calculate the average scaling ratio for each visual target based on the initial mapping point pair as a reference, according to the mapping point pair corresponding to the visual target. The second calculation unit is used to remove mapping point pairs that deviate from the set scaling ratio average and recalculate the scaling ratio average. The target region length and height acquisition unit is used to obtain the length and height of the target region corresponding to the visual target based on the recalculated average scaling ratio and the length and height of the two-dimensional target region. The target area width acquisition unit is used to obtain the width of the target area based on the depth of the three-dimensional laser points corresponding to the projection point set; The 3D bounding box determination unit is used to obtain the 3D bounding box of the target area based on the length, height and width of the target area, determine the 3D laser points in the 3D laser point cloud that fall within the 3D bounding box, and perform point cloud clustering to obtain the 3D laser point set; The device further includes: The preceding target region acquisition module is used to acquire images of several frames preceding the two-dimensional image as preceding images, and to detect visual targets on the preceding images to obtain the preceding target region corresponding to each visual target. The motion speed calculation module is used to calculate the motion speed of the visual target on the two-dimensional image plane based on the preceding target region on the preceding image, the two-dimensional target region on the two-dimensional image, and the time difference between the preceding image and the two-dimensional image. The position prediction module is used to predict the position of the visual target on the two-dimensional image at the second time based on the first time corresponding to the two-dimensional image, the second time corresponding to the three-dimensional laser point cloud, and the motion speed, and set the visual target on the two-dimensional image to the predicted position. Point cloud clustering is performed in the following manner: Each 3D laser point to be clustered is projected onto the xy plane, and the xy plane is rasterized. The number of laser points projected into each grid is counted. If the number is greater than a set threshold, the grid is assigned a value of 255; otherwise, the grid is assigned a value of 0. Connectivity search is performed on the assigned raster image to achieve point cloud clustering.