Laser point cloud shielded vehicle completion method and device based on Leiyu fusion deep learning framework
Through the radar-vision fusion deep learning framework, combined with drone aerial images and laser point cloud data, high-precision completion of obscured vehicles is achieved, solving the problem of missed detection of lidar in the case of vehicle obstruction, improving the integrity and robustness of environmental perception, and reducing system costs.
Patent Information
- Application Number
- CN202510664417.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In existing technologies, lidar cannot effectively scan vehicles behind it when the vehicle is blocked, resulting in missed detections. Existing completion methods are costly or have limited applicable scenarios, and lack a high-precision occluded area completion mechanism, which affects the quality of environmental modeling and perception integrity.
The system adopts a deep learning framework based on radar-vision fusion, synchronously collects and processes UAV aerial images and laser point cloud data, uses a dual-branch deep learning network to predict the position and category of obscured vehicles, and combines it with the vehicle point cloud template library for completion to achieve high-precision vehicle point cloud reconstruction.
It significantly improves the ability to identify and locate obscured vehicles, enhances the integrity and robustness of environmental perception, reduces system deployment and operating costs, and utilizes the wide viewing angle and high maneuverability of drones to improve perception effects in complex traffic scenarios.
Smart Images

Figure CN120655819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for completing a vehicle obscured by a laser point cloud based on a radar-vision fusion deep learning framework, and belongs to the technical field of autonomous driving environment perception. Background Art
[0002] With the rapid growth in the number of motor vehicles, especially cars, road traffic pressure is increasing, with traffic accidents becoming more frequent, diverse, and complex. Although the total number of accidents has stabilized in recent years, traffic safety management remains a significant challenge due to the large base of motor vehicles and drivers. Road traffic accidents not only cause significant property damage and casualties, but also pose a serious threat to urban operational efficiency and public safety. Therefore, improving road traffic perception capabilities, particularly in accident monitoring, early warning, and cause analysis, has become a core focus of the development of intelligent transportation systems (ITS) and traffic management systems.
[0003] In current environmental perception systems, LiDAR (Light Detection and Ranging) has become the primary sensor for autonomous driving and roadside perception due to its high-precision 3D reconstruction capabilities. However, in practical deployments, LiDAR suffers from an inherent occlusion problem: when a leading vehicle completely blocks the laser path of a vehicle behind, the vehicle behind cannot be scanned, resulting in missed detection. This problem is particularly prominent in typical traffic scenarios such as dense traffic or stationary traffic jams. Existing solutions include increasing the number of LiDARs, deploying multi-view perception networks, or using motion trajectory prediction algorithms for target estimation. However, these approaches suffer from high costs, susceptibility to occlusion accumulation errors, and limited applicability. In contrast, drones offer the advantages of a wide overhead field of view, flexible deployment, and high image resolution. They can capture information in blind spots of traditional roadside perception, making them a novel sensing method to assist in blind spot filling. However, the current technology landscape still lacks a method for completing occluded vehicle point cloud data using radar-visual fusion, and a high-precision completion mechanism for occluded areas has yet to be established, hindering the overall environmental modeling quality and system perception integrity. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a method and device for completing vehicles obscured by laser point cloud based on a radar-vision fusion deep learning framework.
[0005] In order to solve the above technical problems, the present invention is implemented by adopting the following technical solutions.
[0006] In a first aspect, the present invention discloses a method for completing vehicles obscured by laser point clouds based on a deep learning framework for laser-visual fusion, comprising: Obtaining raw laser point cloud data and aerial images of road scenes that have been time-aligned and collected synchronously by lidar and drones; Performing recognition processing on the raw laser point cloud data to identify the vehicle ahead, and constructing a sector-shaped shielding area based on the geometric profile of the vehicle and the positional relationship between the laser radar; Performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image; Input the two-dimensional bounding boxes, orientations, and category labels of the fan-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the position, orientation, and category information of the occluded vehicle; Selecting a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the blocked vehicle; Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is placed at the corresponding position in the original laser point cloud data, and a scene point cloud after vehicle point cloud completion is generated as the final result of vehicle completion in the obscured area.
[0007] Furthermore, the identifying and processing the original laser point cloud data to identify the vehicle ahead includes: The original laser point cloud data is preprocessed, including denoising, ground segmentation and Euclidean clustering, to obtain a clustering result, and the front vehicle is identified based on the clustering result.
[0008] Furthermore, the performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image includes: The aerial image is input into a pre-trained YOLOv8 deep learning model, which outputs the target's two-dimensional bounding box, orientation, and category label.
[0009] Furthermore, the dual-branch deep learning network includes: a dual-branch deep learning network of an image feature extraction module, a point cloud feature extraction module, a feature fusion module and a mapping regression module; The image feature extraction module is used to extract the two-dimensional bounding box, orientation and category label of the vehicle target using ConvNeXt to obtain the semantic features of the detected target in the aerial image. ; The point cloud feature extraction module is used to extract the three-dimensional spatial structure features of the fan-shaped occlusion area using PointTransformer ; The feature fusion module is used to transform the semantic features of the detected target in the aerial image based on the pre-built projection mapping matrix 𝑃 The point cloud space projected to the fan-shaped occlusion area Ω and the three-dimensional spatial structure features of the fan-shaped occlusion area Alignment, assisted cross-attention mechanism to generate joint features of image and point cloud ; The projection mapping matrix P The representation is: ; ; in, f x 、 f y is the horizontal and vertical focal length of the image; c x 、 c y is the principal point of the image; K is the camera intrinsic parameter matrix; R The rotation matrix is used to align the point cloud coordinate system with the camera coordinate system; T is the translation vector, which is used to represent the relative position between the lidar and the drone. R and T Perform coordinate conversion between point cloud data and image data; The mapping regression module is used to combine the features of the image and the point cloud Predict whether there is an obscured vehicle and its location, orientation, and category.
[0010] Furthermore, the training process of the dual-branch deep learning network includes: Obtain a historical data set collected synchronously by the lidar and the drone, and use the historical data set to construct a training set; The feature fusion module uses the attention mechanism to achieve cross-modal information alignment, which is expressed as: ; in, A Represents the attention matrix between image and point cloud; W 1. W 2 represents the feature map weight matrix; The damage function used by the dual-branch deep learning network for: ; in, Represents the point cloud coordinates predicted from image features; Represents the center point coordinates of the vehicle in the real point cloud; Represents the weight coefficient of the loss term; The dual-branch deep learning network is trained according to the training set, the attention mechanism and the damage function to obtain a trained dual-branch deep learning network.
[0011] Furthermore, the construction of the vehicle point cloud template library includes: Point cloud data of typical vehicles are extracted from historical lidar data. Euclidean clustering is used to separate the vehicle point cloud from the background based on the spatial distribution of the point cloud data of typical vehicles. The separated vehicle point cloud is labeled and stored in the vehicle point cloud template library.
[0012] Furthermore, the step of placing the vehicle point cloud template at a corresponding position in the original laser point cloud data based on the position and orientation of the obscured vehicle to generate a scene point cloud after the vehicle point cloud is completed includes: The vehicle point cloud template is spatially registered with the original laser point cloud data through an iterative closest point algorithm. After the spatial registration, a probabilistic fusion algorithm is used to eliminate noise in the overlapping area to form a scene point cloud after the vehicle point cloud is completed.
[0013] Furthermore, the spatial registration of the vehicle point cloud template with the original laser point cloud data by using an iterative closest point algorithm includes: (1) Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is preliminarily aligned to the corresponding area of the original laser point cloud data as the initial input of the iterative closest point algorithm; (2) For each point in the vehicle point cloud template, search for the point with the closest Euclidean distance in the original laser point cloud data and establish a point pair correspondence relationship set; (3) Calculate the optimal rotation matrix by minimizing the objective function R 1 and the translation vector T 1. The expression of minimizing the objective function is: ; in, q i Represents the point in the vehicle point cloud template; p i Indicates the original laser point cloud data q i nearest point; represents the square of the Euclidean distance, min It means minimization; (4) The solution R 1 and T 1 Act on the vehicle point cloud template and update its spatial position; (5) Repeat steps (2) to (4) until one of the following convergence conditions is met: The difference between the transformation matrices of two adjacent iterations is less than the threshold , the transformation matrix is obtained by solving R 1 and T 1 constitutes the matrix; The change in the value of the objective function is lower than the preset tolerance; The maximum number of iterations has been reached.
[0014] Furthermore, the probabilistic fusion algorithm is used to eliminate noise in overlapping areas to form a scene point cloud after the vehicle point cloud is completed, including: Probabilistic fusion of the overlapping areas of the registered vehicle point cloud template and the original laser point cloud data, including: According to the point cloud density and registration error, calculate the vehicle point cloud template and the original laser point cloud data. i Confidence of point assignment ω i ; ,in, D i It is i Point cloud density factor of points, E i It is i The registration error factor of each point is α 、 β is the adjustment coefficient; Points with confidence levels lower than the threshold in the overlapping area are considered as outliers, and points with confidence levels not lower than the threshold in the overlapping area are considered as high-confidence vehicle point cloud templates and original laser point cloud data; The outliers are removed, and the high-confidence vehicle point cloud template and the original laser point cloud data are retained to generate a scene point cloud after vehicle completion.
[0015] In a second aspect, the present invention further discloses a laser point cloud occluded vehicle completion device based on a laser-visual fusion deep learning framework, comprising: An acquisition module is used to acquire the original laser point cloud data and aerial images of the road scene that have been time-aligned and collected synchronously by the laser radar and the drone; A first processing module is used to identify and process the raw laser point cloud data, identify the vehicle in front, and construct a sector-shaped blocking area based on the geometric profile of the vehicle and the positional relationship between the laser radar; A second processing module is used to perform recognition processing on the aerial image to identify the two-dimensional bounding box, orientation and category label of all vehicle targets in the image; A prediction module is configured to input the two-dimensional bounding boxes, orientations, and category labels of the sector-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the location, orientation, and category information of the occluded vehicle; A generation module is used to select a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the obscured vehicle; place the vehicle point cloud template at a corresponding position in the original laser point cloud data based on the position and orientation of the obscured vehicle, and generate a scene point cloud after the vehicle point cloud is completed as the final result of vehicle completion in the obscured area.
[0016] The beneficial effects achieved by the present invention are: The laser point cloud occluded vehicle supplementation method based on the radar-vision fusion deep learning framework described in the present invention, by introducing an image-point cloud space mapping mechanism based on the radar-vision fusion deep learning framework, achieves high-precision target correspondence between drone aerial images and ground laser point clouds, significantly improving the automatic recognition and positioning capabilities of obscured vehicles. At the same time, a vehicle point cloud template library is constructed, and a matching vehicle point cloud template can be selected from the vehicle point cloud template library based on the model detection results to achieve complete reconstruction of vehicles in obscured areas, thereby enhancing the environmental perception integrity and robustness of the intelligent transportation system in complex traffic scenarios. In addition, compared with traditional solutions that rely on multiple high-line-count laser radars, the present invention fully utilizes the advantages of drones' wide viewing angle, high maneuverability, and controllable costs, effectively reducing the overall deployment and operating costs of the system while ensuring the perception coverage effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a schematic diagram of the point cloud of the occluded vehicle. DETAILED DESCRIPTION
[0018] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0019] The terms "first," "second," etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance or to implicitly indicate the quantity of the technical features indicated. Thus, a feature specified as "first," "second," etc. may explicitly or implicitly include one or more of the features.
[0020] Example 1: This example introduces a method for completing vehicles obscured by laser point clouds based on a radar-visual fusion deep learning framework, including: Obtaining raw laser point cloud data and aerial images of road scenes that have been time-aligned and collected synchronously by lidar and drones; Performing recognition processing on the raw laser point cloud data to identify the vehicle ahead, and constructing a sector-shaped shielding area based on the geometric profile of the vehicle and the positional relationship between the laser radar; Performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image; Input the two-dimensional bounding boxes, orientations, and category labels of the fan-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the position, orientation, and category information of the occluded vehicle; Selecting a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the blocked vehicle; Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is placed at the corresponding position in the original laser point cloud data, and a scene point cloud after vehicle point cloud completion is generated as the final result of vehicle completion in the obscured area.
[0021] The identifying and processing the original laser point cloud data to identify the vehicle ahead includes: The original laser point cloud data is preprocessed, including denoising, ground segmentation and Euclidean clustering, to obtain a clustering result, and the front vehicle is identified based on the clustering result.
[0022] The step of performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image includes: The aerial image is input into a pre-trained YOLOv8 deep learning model, which outputs the target's two-dimensional bounding box, orientation, and category label.
[0023] The dual-branch deep learning network includes: an image feature extraction module, a point cloud feature extraction module, a feature fusion module and a mapping regression module; The image feature extraction module is used to extract the two-dimensional bounding box, orientation and category label of the vehicle target using ConvNeXt to obtain the semantic features of the detected target in the aerial image. ; The point cloud feature extraction module is used to extract the three-dimensional spatial structure features of the fan-shaped occlusion area using PointTransformer ; The feature fusion module is used to transform the semantic features of the detected target in the aerial image based on the pre-built projection mapping matrix 𝑃 The point cloud space projected to the fan-shaped occlusion area Ω and the three-dimensional spatial structure features of the fan-shaped occlusion area Alignment, assisted cross-attention mechanism to generate joint features of image and point cloud ; The projection mapping matrix P The representation is: ; ; in, f x 、 f y is the horizontal and vertical focal length of the image; c x 、 c y is the principal point of the image; K is the camera intrinsic parameter matrix; R The rotation matrix is used to align the point cloud coordinate system with the camera coordinate system; T is the translation vector, which is used to represent the relative position between the lidar and the drone. R and T Perform coordinate conversion between point cloud data and image data; The mapping regression module is used to combine the features of the image and the point cloud Predict whether there is an obscured vehicle and its location, orientation, and category.
[0024] The training process of the dual-branch deep learning network includes: Obtain a historical data set collected synchronously by the lidar and the drone, and use the historical data set to construct a training set. The training set can also use a KITTI data set, etc.
[0025] The feature fusion module uses the attention mechanism to achieve cross-modal information alignment, which is expressed as: ; in, Represents the attention matrix between image and point cloud; represents the feature map weight matrix; The damage function used by the dual-branch deep learning network for: ; in, Represents the point cloud coordinates predicted from image features; Represents the center point coordinates of the vehicle in the real point cloud; Represents the weight coefficient of the loss term; The dual-branch deep learning network is trained according to the training set, the attention mechanism and the damage function to obtain a trained dual-branch deep learning network.
[0026] The construction of the vehicle point cloud template library includes: Point cloud data of typical vehicles are extracted from historical lidar data. Euclidean clustering is used to separate the vehicle point cloud from the background based on the spatial distribution of the point cloud data of typical vehicles. The separated vehicle point cloud is labeled and stored in the vehicle point cloud template library.
[0027] Placing the vehicle point cloud template at a corresponding position in the original laser point cloud data based on the position and orientation of the obscured vehicle to generate a scene point cloud after the vehicle point cloud is completed includes: The vehicle point cloud template is spatially registered with the original laser point cloud data through an iterative closest point algorithm. After the spatial registration, a probabilistic fusion algorithm is used to eliminate noise in the overlapping area to form a scene point cloud after the vehicle point cloud is completed.
[0028] The spatial registration of the vehicle point cloud template with the original laser point cloud data by using an iterative closest point algorithm includes: (1) Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is preliminarily aligned to the corresponding area of the original laser point cloud data as the initial input of the iterative closest point algorithm; (2) For each point in the vehicle point cloud template, search for the point with the closest Euclidean distance in the original laser point cloud data and establish a point pair correspondence relationship set; (3) Calculate the optimal rotation matrix by minimizing the objective function R 1 and the translation vector T 1. The expression of minimizing the objective function is: ; in, q i Represents the point in the vehicle point cloud template; p i Indicates the original laser point cloud data q i nearest point; represents the square of the Euclidean distance, min It means minimization; (4) The solution R 1 and T 1 Act on the vehicle point cloud template and update its spatial position; (5) Repeat steps (2) to (4) until one of the following convergence conditions is met: The difference between the transformation matrices of two adjacent iterations is less than the threshold value, and the transformation matrix is obtained by solving R 1 and T 1 constitutes the matrix; The change in the value of the objective function is lower than the preset tolerance; The maximum number of iterations has been reached.
[0029] The probabilistic fusion algorithm is used to eliminate noise in overlapping areas to form a scene point cloud after the vehicle point cloud is completed, including: Probabilistic fusion of the overlapping areas of the registered vehicle point cloud template and the original laser point cloud data, including: According to the point cloud density and registration error, calculate the vehicle point cloud template and the original laser point cloud data. i Confidence of point assignment ω i ; ,in, D i and E i It is i The point cloud density factor and registration error factor of each point (the ICP algorithm can automatically calculate the corresponding value based on the point cloud data), α 、 β is the adjustment coefficient; Points with confidence levels lower than the threshold in the overlapping area are considered as outliers, and points with confidence levels not lower than the threshold in the overlapping area are considered as high-confidence vehicle point cloud templates and original laser point cloud data; The outliers are removed, and the high-confidence vehicle point cloud template and the original laser point cloud data are retained to generate a scene point cloud after vehicle completion.
[0030] Example 2: This example introduces a method for completing vehicles obscured by laser point cloud based on a deep learning framework for laser-visual fusion. Figure 1 As shown, the following steps are included: Step 1: Synchronous acquisition of multi-source heterogeneous data: Laser point cloud data and aerial images of road scenes are collected synchronously by laser radar and drone, and point cloud and image are time-aligned using RTK-GPS. In order to achieve the subsequent fusion processing of image and point cloud data, the internal and external parameters K, R, T of the aerial drone and laser radar are used to construct the point cloud coordinate system. To image coordinate system The projection mapping matrix P can be used to map spatial points to the image plane to achieve spatial correspondence of multimodal data; ; ; in, is the horizontal and vertical focal length of the image (in pixels); is the principal point of the image (usually the center of the image); K is the camera intrinsic parameter matrix, which comes from camera calibration or manufacturer parameters; The rotation matrix aligns the point cloud coordinate system with the camera coordinate system. is the translation vector, which represents the relative position between the two sensors.
[0031] Step 2: LiDAR occlusion area detection: Use existing advanced technologies to pre-process the laser point cloud data, including denoising, ground segmentation and Euclidean clustering, and identify the vehicle ahead based on the clustering results. Once the vehicle ahead (such as a large truck or bus) is identified through clustering in the point cloud, a fan-shaped occlusion area can be constructed based on the relationship between the vehicle's geometric outline and the LiDAR position to limit the possible occlusion area. There is a high probability that there are occluded vehicle targets in this area, but they are missing in the point cloud and cannot be clustered or identified. In the subsequent steps, combined with the drone aerial image, whether there are missing vehicles in the area will be detected. If there is a target in the image but missing in the point cloud, the "Occluded Vehicle Point Cloud Template Generation" module will be triggered to complete the occluded vehicle point cloud. The occluded area is modeled in polar coordinates and is defined as: ; in, It represents the radial distance from the laser radar to a point in space, in meters; Indicates the scanning angle of the point in the horizontal direction, in degrees or radians; Indicates the minimum and maximum radial extents of the occluded area; Shows the minimum and maximum scanning angles of the blocked area.
[0032] Step 3: Vehicle Detection in Drone Aerial Imagery: The aerial imagery is fed into a pre-trained YOLOv8 (You Only Look Once version 8) deep learning model to rapidly identify and locate all vehicle targets in the image. The model then outputs a 2D bounding box, orientation, and class label for the target. In this case, the class label primarily includes vehicle types in typical road traffic scenarios, including but not limited to cars, buses, and trucks.
[0033] Step 4: Cross-modal target matching and missing judgment in occluded areas: Construct a dual-branch deep learning network, which consists of four modules: image feature extraction module, point cloud feature extraction module, feature fusion module, and mapping regression module. The image branch uses ConvNeXt for feature extraction, the point cloud branch uses PointTransformer to extract 3D spatial structure features, and the fusion module fuses the semantic features of the detected targets in the aerial image. Spatial characteristics of the fan-shaped occlusion area Ω of laser point cloud data , learning the mapping relationship between the two, providing an "alignment" basis for feature fusion, and inputting the fusion results into the mapping regression module to predict the presence of an occluded vehicle and the location, orientation, and category of the occluded vehicle, providing a spatial basis for subsequent vehicle point cloud template generation. When a vehicle is detected in the aerial image but there is no target in the corresponding point cloud, it is determined to be an occluded vehicle. Feature fusion uses the attention mechanism to achieve cross-modal information alignment: ; in, Represents the joint features of the fused image and point cloud; Represents the attention matrix between image and point cloud; Represents the feature map weight matrix.
[0034] Through the end-to-end training of the two-branch network, supervised learning is performed using the following loss function: ; in, Represents the point cloud coordinates predicted from image features; Represents the center point coordinates of the vehicle in the real point cloud; Represents the weight coefficient of the loss term (generally set empirically as 1.0 and 0.5).
[0035] Step 5: Construct and select occluded vehicle point cloud templates: Extract point cloud data of typical vehicles from the LiDAR data, construct a vehicle point cloud template library, use Euclidean clustering to separate the vehicle point clouds from the background based on the spatial distribution of the point clouds, and label the separated vehicle point clouds with categories, including but not limited to small cars (cars), buses (buses), and trucks (trucks). Based on the occluded vehicle information predicted in Step 4, select a matching vehicle point cloud template from the vehicle point cloud template library.
[0036] Step 6: Fusion of the vehicle point cloud template and the original point cloud data: Using the occluded vehicle information determined by image and point cloud mapping in step 4, the vehicle point cloud template selected in step 5 is placed at the corresponding position. The vehicle point cloud template and the original point cloud are spatially registered using the iterative closest point (ICP) algorithm. A probabilistic fusion algorithm is used to eliminate noise in overlapping areas to form a scene point cloud after the vehicle point cloud is completed. The (ICP) algorithm specifically includes the following steps: (1) Initial alignment of the vehicle point cloud template and the original point cloud: Based on the position, orientation, and category information of the occluded vehicle determined in step 4, the vehicle point cloud template selected in step 5 is preliminarily aligned to the corresponding area in the original point cloud as the initial input of the ICP algorithm.
[0037] Closest point matching: For each point in the vehicle point cloud template , search for the point with the closest Euclidean distance in the original point cloud , establish a point pair correspondence set.
[0038] (3) Rigid body transformation solution: Calculate the optimal rotation matrix by minimizing the objective function and translation vectors ; ; in, Indicates the first points; Indicates the original point cloud nearest point; is the rotation matrix; is the translation vector.
[0039] (4) Transformation application and iterative update: and Act on the vehicle point cloud template, update its spatial position, and repeat steps (2) to (4) until one of the following convergence conditions is met: The difference between the transformation matrices of two adjacent iterations is less than the threshold ; The change in the objective function value is lower than the preset tolerance; The maximum number of iterations has been reached.
[0040] (5) Probabilistic fusion denoising: Probabilistic fusion is performed on the overlapping areas of the registered vehicle point cloud template and the original point cloud. Confidence weights are assigned to points in the vehicle point cloud template and the original point cloud; outliers with confidence below the threshold in the overlapping area are removed; high-confidence vehicle point cloud templates and original point cloud data are retained to generate the scene point cloud after vehicle completion.
[0041] (6) Output completion result: Output the fused scene point cloud as the final result of vehicle completion in the occluded area, such as Figure 2 shown.
[0042] Example 3, based on the same inventive concept as Example 1, introduces a laser point cloud occluded vehicle completion device based on a radar-visual fusion deep learning framework, comprising: An acquisition module is used to acquire the original laser point cloud data and aerial images of the road scene that have been time-aligned and collected synchronously by the laser radar and the drone; A first processing module is used to identify and process the raw laser point cloud data, identify the vehicle in front, and construct a sector-shaped blocking area based on the geometric profile of the vehicle and the positional relationship between the laser radar; A second processing module is used to perform recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image; A prediction module is configured to input the two-dimensional bounding boxes, orientations, and category labels of the sector-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the location, orientation, and category information of the occluded vehicle; A generation module is used to select a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the obscured vehicle; place the vehicle point cloud template at a corresponding position in the original laser point cloud data based on the position and orientation of the obscured vehicle, and generate a scene point cloud after the vehicle point cloud is completed as the final result of vehicle completion in the obscured area.
[0043] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0044] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0045] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0046] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0047] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A laser point cloud occluded vehicle completion method based on a radar-visual fusion deep learning framework, characterized in that: include: Obtaining raw laser point cloud data and aerial images of road scenes that have been time-aligned and collected synchronously by lidar and drones; Performing recognition processing on the raw laser point cloud data to identify the vehicle ahead, and constructing a sector-shaped shielding area based on the geometric profile of the vehicle and the positional relationship between the laser radar; Performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image; Input the two-dimensional bounding boxes, orientations, and category labels of the fan-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the position, orientation, and category information of the occluded vehicle; Selecting a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the blocked vehicle; Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is placed at the corresponding position in the original laser point cloud data, and a scene point cloud after vehicle point cloud completion is generated as the final result of vehicle completion in the obscured area.
2. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 1 is characterized in that: The identifying and processing the original laser point cloud data to identify the vehicle ahead includes: The original laser point cloud data is preprocessed, including denoising, ground segmentation and Euclidean clustering, to obtain a clustering result, and the front vehicle is identified based on the clustering result.
3. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 1 is characterized in that: The step of performing recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image includes: The aerial image is input into a pre-trained YOLOv8 deep learning model, which outputs the target's two-dimensional bounding box, orientation, and category label.
4. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 1 is characterized in that: The dual-branch deep learning network includes: an image feature extraction module, a point cloud feature extraction module, a feature fusion module and a mapping regression module; The image feature extraction module is used to extract the two-dimensional bounding box, orientation and category label of the vehicle target using ConvNeXt to obtain the semantic features of the detected target in the aerial image. ; The point cloud feature extraction module is used to extract the three-dimensional spatial structure features of the fan-shaped occlusion area using PointTransformer ; The feature fusion module is used to transform the semantic features of the detected target in the aerial image based on the pre-built projection mapping matrix 𝑃 The point cloud space projected to the fan-shaped occlusion area Ω and the three-dimensional spatial structure features of the fan-shaped occlusion area Alignment, assisted cross-attention mechanism to generate joint features of image and point cloud ; The projection mapping matrix P The representation is: ; ; in, f x 、 f y is the horizontal and vertical focal length of the image; c x 、 c y is the principal point of the image; K is the camera intrinsic parameter matrix; R The rotation matrix is used to align the point cloud coordinate system with the camera coordinate system; T is the translation vector, which is used to represent the relative position between the lidar and the drone. R and T Perform coordinate conversion between point cloud data and image data; The mapping regression module is used to combine the features of the image and the point cloud Predict whether there is an obscured vehicle and its location, orientation, and category.
5. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 4 is characterized in that: The training process of the dual-branch deep learning network includes: Obtain a historical data set collected synchronously by the lidar and the drone, and use the historical data set to construct a training set; The feature fusion module uses the attention mechanism to achieve cross-modal information alignment, which is expressed as: ; in, A Represents the attention matrix between image and point cloud; W 1. W 2 represents the feature map weight matrix; The damage function used by the dual-branch deep learning network for: ; in, Represents the point cloud coordinates predicted from image features; Represents the center point coordinates of the vehicle in the real point cloud; Represents the weight coefficient of the loss term; The dual-branch deep learning network is trained according to the training set, the attention mechanism and the damage function to obtain a trained dual-branch deep learning network.
6. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 1 is characterized in that: The construction of the vehicle point cloud template library includes: Point cloud data of typical vehicles are extracted from historical lidar data. Euclidean clustering is used to separate the vehicle point cloud from the background based on the spatial distribution of the point cloud data of typical vehicles. The separated vehicle point cloud is labeled and stored in the vehicle point cloud template library.
7. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 1 is characterized in that: Placing the vehicle point cloud template at a corresponding position in the original laser point cloud data based on the position and orientation of the obscured vehicle to generate a scene point cloud after the vehicle point cloud is completed includes: The vehicle point cloud template is spatially registered with the original laser point cloud data through an iterative closest point algorithm. After the spatial registration, a probabilistic fusion algorithm is used to eliminate noise in the overlapping area to form a scene point cloud after the vehicle point cloud is completed.
8. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 7 is characterized in that: The spatial registration of the vehicle point cloud template with the original laser point cloud data by using an iterative closest point algorithm includes: (1) Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is preliminarily aligned to the corresponding area of the original laser point cloud data as the initial input of the iterative closest point algorithm; (2) For each point in the vehicle point cloud template, search for the point with the closest Euclidean distance in the original laser point cloud data and establish a point pair correspondence relationship set; (3) Calculate the optimal rotation matrix by minimizing the objective function R 1 and the translation vector T 1. The expression of minimizing the objective function is: ; in, q i Represents the point in the vehicle point cloud template; p i Indicates the original laser point cloud data q i nearest point; represents the square of the Euclidean distance, min It means minimization; (4) The solution R 1 and T 1 Act on the vehicle point cloud template and update its spatial position; (5) Repeat steps (2) to (4) until one of the following convergence conditions is met: The difference between the transformation matrices of two adjacent iterations is less than the threshold , the transformation matrix is obtained by solving R 1 and T 1 constitutes the matrix; The change in the value of the objective function is lower than the preset tolerance; The maximum number of iterations has been reached.
9. The laser point cloud occluded vehicle completion method based on the laser-visual fusion deep learning framework according to claim 7 is characterized in that: The probabilistic fusion algorithm is used to eliminate noise in overlapping areas to form a scene point cloud after the vehicle point cloud is completed, including: Probabilistic fusion of the overlapping areas of the registered vehicle point cloud template and the original laser point cloud data, including: According to the point cloud density and registration error, calculate the vehicle point cloud template and the original laser point cloud data. i Confidence of point assignment ω i ; ,in, D i It is i Point cloud density factor of points, E i It is i The registration error factor of each point is α 、 β is the adjustment coefficient; Points with confidence levels lower than the threshold in the overlapping area are considered as outliers, and points with confidence levels not lower than the threshold in the overlapping area are considered as high-confidence vehicle point cloud templates and original laser point cloud data; The outliers are removed, and the high-confidence vehicle point cloud template and the original laser point cloud data are retained to generate a scene point cloud after vehicle completion.
10. A laser point cloud occluded vehicle completion device based on a radar-visual fusion deep learning framework, characterized in that: include: An acquisition module is used to acquire the original laser point cloud data and aerial images of the road scene that have been time-aligned and collected synchronously by the laser radar and the drone; A first processing module is used to identify and process the raw laser point cloud data, identify the vehicle in front, and construct a sector-shaped blocking area based on the geometric profile of the vehicle and the positional relationship between the laser radar; A second processing module is used to perform recognition processing on the aerial image to identify the two-dimensional bounding boxes, orientations, and category labels of all vehicle targets in the image; A prediction module is configured to input the two-dimensional bounding boxes, orientations, and category labels of the sector-shaped occluded area and all vehicle targets into a pre-trained two-branch deep learning network to predict whether there is an occluded vehicle and the location, orientation, and category information of the occluded vehicle; A generating module, configured to select a matching vehicle point cloud template from a pre-built vehicle point cloud template library based on the category information of the obscured vehicle; Based on the position and orientation of the obscured vehicle, the vehicle point cloud template is placed at the corresponding position in the original laser point cloud data, and a scene point cloud after vehicle point cloud completion is generated as the final result of vehicle completion in the obscured area.
Citation Information
Patent Citations
Three-dimensional laser radar point cloud data amplification method based on metamorphic algorithm
CN114265074A
Railway operation environment abnormity identification method based on image and laser data fusion
CN114266891A
Object detection method based on image segmentation and laser radar point cloud completion
CN115512330A
Target detection method, system and equipment based on laser radar and camera fusion
CN116205989A
Target identification method based on fusion of image information and laser radar point cloud information
CN116229408A
Cited By
Moving target rapid positioning method and device based on Leiyu fusion, and electronic equipment
CN121234109A
Mobile target rapid positioning method and device based on radar and visual fusion and electronic equipment
CN121234109B
Radar scanning and deep learning-based quantitative loading auxiliary method and system
CN121883570A
Point cloud recognition training data enhancement method for semi-closed scene
CN121982455A
Point cloud completion methods, devices, media and equipment for underground parking lot scenes
CN122415639A