A method for detecting three-dimensional targets by a laser radar-optical camera of a UAV based on canopy density perception
Patent Information
- Application Number
- CN202611043660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明为了解决现有多模态三维目标检测方法在无人机俯视场景中缺乏遮挡感知能力的问题
[0039]本发明首个提出面向单无人机平台的激光雷达-相机多模态融合三维目标检测框架,实现了两种模态互补信息在无人机场景中的有效利用。
Smart Images

Figure CN122821535A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal 3D target detection technology, specifically relating to a 3D target detection method based on the fusion of lidar and camera data for UAV platforms. Background Technology
[0002] Multimodal 3D target detection based on LiDAR and cameras is a key task in 3D scene perception, with wide applications in autonomous driving, intelligent surveillance, and military reconnaissance. LiDAR provides high-precision 3D structural information, while cameras provide rich color and texture information; their complementary fusion can overcome the limitations of single-sensor perception. Existing multimodal 3D target detection methods are all designed for vehicle-mounted platforms and are difficult to directly transfer to UAV platforms with drastically different observation geometry and degradation modes. Meanwhile, existing 3D target detection methods based on single UAV platforms are mainly single-modal, and LiDAR-camera multimodal fusion methods have not yet been explored.
[0003] In UAV-view scenarios, ground targets are frequently obscured by vegetation canopies. This obscuration leads to information degradation characterized by spatial non-uniformity and modal differences: while lidar pulses can partially penetrate canopy gaps to obtain sparse but still usable point clouds, color and texture cues relied upon by the camera are directly obscured by the canopy, resulting in more severe degradation. Experiments show that directly applying existing multimodal fusion methods under these conditions not only fails to improve performance but also leads to lower detection accuracy than pure lidar methods due to the introduction of degraded image features. Although the attenuation of signals by the canopy follows the well-known Beer-Lambert law, this physical mechanism has not yet been translated into an occlusion prior that the detection network can directly utilize, and there is a lack of adaptive mechanisms for occlusion perception at each stage of the detection process. Summary of the Invention
[0004] This invention aims to address the problem that existing multimodal 3D target detection methods lack occlusion perception capabilities in UAV overhead scenes.
[0005] A multimodal 3D target detection method based on canopy closure perception using a lidar-optical camera includes:
[0006] The processing steps of the bimodal canopy closure prediction network include:
[0007] Step 2.1: Input the image corresponding to the optical camera into the image encoder to obtain the feature pyramid, and then obtain the multi-scale aggregated features. The current frame point cloud is converted into voxels and then input into a LiDAR sparse encoder to obtain sparse BEV features.
[0008] Step 2.2, for Parallel convolutions with different kernel sizes are performed separately and concatenated along the channel dimension, then fused together by convolution to form a vegetation saliency feature map. Project the current frame point cloud onto the image coordinate system, and count the normalized number of projected points, normalized average depth, and normalized depth variance at each pixel location to form a three-channel projected feature map, which is then encoded as point cloud auxiliary features. ;Will , and After being concatenated along the channel dimension and then fused by convolution, a single-channel occlusion intensity prediction map is obtained after passing through an activation function. ;
[0009] Step 2.3, UAV location A virtual perspective model is established using the terrain-predicted ground grid as the projection plane, with the projection center as the model; for the first... Individual elements, based on their horizontal position, are predicted by querying a terrain map. Obtain predicted ground elevation To obtain the effective focal length ,based on A virtual perspective model is defined and terrain-adaptive projection is performed, projecting the shallow voxel features output from the first layer of all LiDAR sparse encoders onto BEV mesh cells. middle; The corresponding set of voxels is , for The number of voxels on the surface Projecting shallow voxel features onto the corresponding mesh For projections onto the same grid cell ,based on Calculate the weighted average characteristic The geometrical and statistical features related to the canopy occlusion state are then concatenated along the channel dimension and input into the prediction head to obtain the LiDAR domain occlusion intensity prediction map. ;
[0010] The occlusion-prior-guided fusion strategy network processing steps include:
[0011] Step 3.3, will Adjust to After resolution, and The weight graph is obtained by concatenating the weights along the channel dimension and then inputting them into a dynamic weight generation network. Used for Weighted image features are obtained by weighting. For each sparse BEV cell, extract its world coordinates and predict them based on the terrain map. Obtain predicted ground elevation plus height offset Near-ground coordinates were then obtained. ;Will Projected to The corresponding image plane, in Sampling is performed on the data, and a corresponding feature vector is obtained for each query point; the set of sampling results for all query points constitutes the image modality weighted fusion feature. ;Will By concatenating the sparse BEV features along the channel dimension, joint features are obtained. These features are then fused using sparse convolution of submanifolds to obtain sparse BEV fused features. ;
[0012] Step 3.4: Based on sparse BEV fusion features The detection is completed using a detection head. The heatmap classification branch in the detection head is used to determine the category of the detection box, and the 3D box parameter regression branch in the detection head is used to determine the position, size and orientation of the detection box.
[0013] Furthermore, the process of completing the detection using the detection head includes:
[0014] Sparse BEV fusion features The position, size, and orientation of the detection box are determined by inputting the 3D bounding box parameter regression branch;
[0015] Inject occlusion prior modulation into the input of the heatmap classification branch, i.e. ,in for Middle position The value, To implement the occlusion prior after gating Mapping to AND Same-dimensional feature residual network, for Middle position The value; To detect the gating threshold of the response modulation; is the modulation intensity coefficient, where Learnable parameters initialized to zero. To modulate the upper bound; based on sparse BEV fusion features Received The category of the detection box is determined by sending it to the heatmap classification branch.
[0016] Furthermore, the geometrical statistical features related to canopy shading status include:
[0017] Normalized point density Normalized relative ground height Standard deviation of height Ground confidence level ;
[0018] in, The maximum number of voxels in a single grid cell within the scene; For the first The relative ground height of individual units, The average relative height. Here is the height coordinate value of the m-th voxel; For unit Predicted ground elevation; This is an indicator function that takes the value 1 if the condition within the parentheses is true, and 0 otherwise. The height standard deviation threshold in ground confidence assessment; This is the height threshold used in ground confidence assessment.
[0019] Furthermore, the bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network are pre-trained, and the training process includes:
[0020] Ground truth occlusion intensity values in the image and LiDAR domains are constructed offline using dense point clouds with multiple stripes. Then, the occlusion intensity values for building-covered areas are corrected. The corrected ground truth occlusion intensity values in the image domain are then used as the basis for the final determination. True values of LiDAR domain occlusion intensity ;
[0021] Image and point cloud data are fed into a bimodal canopy closure prediction network for processing; the ground truth occlusion intensity in the image domain is then calculated. After being aligned to the feature map resolution by bilinear interpolation, the classification loss is obtained as supervision. ;
[0022] Perform occlusion intensity gating contrast feature learning, including:
[0023] The eight corner points of the target's 3D bounding box are projected onto the image plane using a calibration matrix, and the two-dimensional bounding rectangle of the projected corner points is taken as the RoI region of the target in the corresponding viewpoint; for each valid RoI sample Query the true value of the image domain occlusion intensity at the center of its two-dimensional projection frame. Only retain the true value of occlusion intensity below the threshold. The samples were used for contrastive learning; subsequently, RoIAlign was used to analyze the feature maps. Local image features of a fixed size are extracted and mapped into normalized embedding vectors by a projection head. In the sample set participating in contrastive learning, RoI samples of the same class are treated as positive samples, and RoI samples of different classes are treated as negative samples. The weighted InfoNCE loss function is used as the contrastive learning loss. ;
[0024] Based on the output of the bimodal canopy closure prediction network, an occlusion prior-guided fusion strategy network makes predictions to obtain the category, location, size, and orientation of the detection boxes; based on a learning loss... and classification loss The total loss completes the training of the bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network.
[0025] Furthermore, the process of offline construction of ground truth values for occlusion intensity in the image and LiDAR domains using multi-band dense point clouds includes:
[0026] Step 1.1: Construct a canopy closure model ,in The effective attenuation coefficient is a composite parameter of the absorption extinction coefficient and the unit conversion factor. The reference point density is used for density normalization across different scenarios. This represents the average point density of the current scene. and To construct the view frustum The number of non-ground points and cross-sectional area inside;
[0027] Step 1.2: Construct the view frustum query framework: The four ray directions corresponding to the four corner points determine the pixel view frustum. angular range in spherical coordinates , Let be the maximum and minimum values of the azimuth angle in the coordinate system. , The maximum and minimum zenith angles in the coordinate system; based on the view frustum query framework, from multi-band dense point clouds Query the set of points that satisfy the angle constraint Average ground distance The area of the spherical patch at that point is used as the cross-sectional area of the frustum. Using DEM reference elevation Separate ground points from non-ground points to obtain the number of non-ground points. ;Will and Substitute the canopy closure model, traverse all valid pixels, and generate an image domain canopy closure map. , that is, the true value of the occlusion intensity in the image domain;
[0028] Step 1.3: Construct a narrow cylindrical query framework: Construct a narrow cylindrical query region along the unit vector of the observation direction. The narrow cylinder is defined as the region with the observation ray from the UAV to the center of the ground grid as its central axis and a lateral query radius smaller than a set threshold; non-ground points are retrieved from multiple dense point clouds and their numbers are counted. The area of the ground grid used for LiDAR domain canopy closure modeling is taken as the cross-sectional area of a narrow cylinder. Substitute the data into the canopy closure model, traverse all cells, and generate a LiDAR domain canopy closure map. That is, the true value of the occlusion intensity in the LiDAR domain; Fine-grid resolution used for offline modeling of a dual-modal canopy closure prediction network.
[0029] Furthermore, the canopy closure model The construction process includes:
[0030] Based on reference point density Canopy thickness along the observation direction And the number of non-ground points inside the view frustum constructed along the observation direction. and cross-sectional area Determine the cumulative leaf area index The complement of transmittance is defined as canopy closure, and a canopy closure model is then constructed. .
[0031] Furthermore, the average ground distance The area of the spherical piece at that location is ,in ;
[0032] The fine grid resolution used in offline modeling based on the bimodal canopy closure prediction network corresponds to the grid area of the square. , Fine-grid resolution used for offline modeling of a dual-modal canopy closure prediction network.
[0033] Furthermore, in the process of correcting the occlusion intensity value of the building coverage area, based on the occlusion intensity in the image domain and LiDAR domain, the occlusion intensity value of the building coverage area is corrected to 1, resulting in the corrected occlusion intensity value.
[0034] Furthermore, the classification loss function It mainly consists of the weighted sum of mean square error loss, edge perception loss, high occlusion region consistency loss, and smoothing term.
[0035] Furthermore, the training process of the bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network also includes a step of generating samples to construct a sample library. This sample generation is performed using a bimodal occlusion matching method, including:
[0036] During the sample library construction phase, images, point clouds, and their corresponding C++ samples are included. img and C lidAs a single sample, for each target, the mean, standard deviation, and Gaussian weighted mean at the target location are extracted from the ground truth maps of occlusion intensity in both the image and LiDAR domains, resulting in six statistics. These statistics are then weighted by modal weighting coefficients and concatenated to form a six-dimensional bimodal occlusion state descriptor. KD trees are constructed according to the target detection category: each tree uses the occlusion state descriptors of all samples in that category as nodes, and the Euclidean distance between nodes directly measures the similarity of occlusion states, thereby supporting sample retrieval based on occlusion similarity.
[0037] In the sample generation stage, a location in the scene is randomly selected from the effective area of the DEM as the sample placement point. Its pixel coordinates are converted into a planar position in the LiDAR coordinate system, and the elevation of this point is used as the ground height. Combined with the height of the target itself, the vertical coordinates of the target box center are determined. The bimodal occlusion state descriptor of this placement location is extracted. In the KD tree of this target detection category, several candidate samples with the nearest neighbor of the occlusion state are retrieved based on Euclidean distance. These are the samples that are closest to the occlusion conditions of the placement location. Three-dimensional spatial collision detection and two-dimensional image overlap verification are performed on each candidate sample. The first sample that passes the verification is pasted to achieve sample generation. After pasting, the ground value of the occlusion intensity in the image domain is updated by maximum value fusion in the corresponding region.
[0038] Beneficial effects:
[0039] This invention is the first to propose a multimodal fusion 3D target detection framework based on lidar and camera for a single UAV platform, which realizes the effective utilization of complementary information from two modes in UAV scenarios.
[0040] This invention constructs a directly usable bimodal occlusion prior for the detection network, enabling the network to perceive the degree of information degradation of each modality at each spatial location, thus providing a physical basis for adaptive occlusion detection.
[0041] The occlusion prior-guided fusion strategy designed in this invention covers all four stages of the detection process, enabling the detector to adaptively adjust according to occlusion conditions in data augmentation, feature encoding, multimodal fusion, and detection head stages, especially showing more significant advantages in severely occluded and difficult samples. Attached Figure Description
[0042] Figure 1 This is a diagram of the overall logical architecture of the present invention;
[0043] Figure 2 Here is the architecture diagram of the DPM module;
[0044] Figure 3 This is a diagram of the DCPNet architecture.
[0045] Figure 4 OPF strategy architecture diagram;
[0046] Figure 5 A visualization comparison of the detection results of BEVFusion, TG-ADet, and the present invention on the SI3D-DI dataset. Detailed Implementation
[0047] Specific implementation method one: Combining Figure 1 This implementation method is described below. Figure 1 The overall detection architecture diagram shown includes three core components: DPM offline modeling, DCPNet online prediction, and OPF fusion strategy, as well as their data flow relationships. The green arrows indicate the guidance path used only in the training phase, while the orange arrows indicate the guidance path used in both the training and inference phases.
[0048] The lidar-optical camera multimodal 3D target detection method based on canopy closure perception described in this embodiment is a 3D target detection method for unmanned aerial vehicle (UAV) platforms, specifically including:
[0049] Step 1: Using the dual-modal canopy closure physical modeling module (DPM), based on physical modeling and building mask correction, offline construction of occlusion intensity ground truth values in the image domain and LiDAR domain is carried out using multi-band dense point cloud.
[0050] Figure 2 The diagram shows the DPM module architecture, including the Beer-Lambert canopy closure modeling principle (left), image domain pixel frustum modeling (middle), LiDAR domain BEV ground mesh modeling (right), and building mask correction. Figure 2 As shown, the specific processing steps of step 1 include:
[0051] Step 1.1, Beer-Lambert canopy closure modeling:
[0052] The vegetation canopy, composed of randomly distributed branches and leaves, can be approximated as a random scattering medium to which the Beer-Lambert law applies. For traversing a thickness of... The laser pulse in the canopy, at position Leaf area density at that location is At that time, the transmittance is:
[0053]
[0054] in Extinction coefficient, This represents the cumulative leaf area index.
[0055] Canopy density is defined as the complement of transmittance, i.e. .
[0056] Because airborne lidar cannot directly measure , Statistical approximation is performed using dense point clouds with multiple bands. These dense point clouds are derived from repeated scans taken by UAVs flying in multiple directions, which reduces canopy blind spots. For a given location on the ground, an appearance cone is constructed along the observation direction. Count the number of non-ground points within it. and cross-sectional area Assuming the canopy is approximately uniformly distributed along the observation path, the surface density at non-ground points can be used as... Statistical proxy volume:
[0057]
[0058] in, The thickness of the canopy along the observation direction;
[0059] Substituting the above equation into the Beer-Lambert form, we introduce the effective attenuation coefficient. To absorb unknown extinction coefficients and unit conversion factors, a density normalization factor is also introduced. To compensate for the differences in sampling density between different flight strips, the canopy closure was obtained:
[0060]
[0061] in, The effective attenuation coefficient is a composite parameter of the absorption extinction coefficient and the unit conversion factor, with a default value of 0.001. This is the reference point density at the dataset level, used for density normalization across different scenarios; This represents the average point density of the current scene.
[0062] The aforementioned model is a physically inspired statistical approximation method used to capture the relative spatial variation trend of canopy attenuation, rather than precisely inverting canopy optical parameters. The separation between ground points and non-ground points is achieved by referencing elevations using a digital elevation model (DEM). Achieve: Lower the elevation within the visual cone than Points with a height of m (where 1 m is the ground separation height threshold) are classified as ground points, and the rest are classified as non-ground points.
[0063] Step 1.2, Image Domain Canopy Closure Modeling:
[0064] Image domain canopy occlusion modeling uses each pixel of the image as a spatial unit to quantify the degree of canopy occlusion in the direction of the optical camera's observation.
[0065] For the coordinates in the image, The pixels are taken, and their four corner points (top, bottom, left, and right) are used to correct the intrinsic parameter matrix of the optical camera. Rotation matrix from world coordinate system to optical camera coordinate system Backproject the corner points into ray directions in the world coordinate system: ,in Number the corner points. For the first Image coordinates of the corner points This is the corresponding ray direction vector.
[0066] View frustum lookup framework: The four ray directions corresponding to the four corner points determine the pixel view frustum. angular range in spherical coordinates From dense point clouds with multiple bands Query the set of points that satisfy the angle constraint:
[0067]
[0068] in, Let q be the position of the optical center of the optical camera in the world coordinate system, and let q be the set of dense point clouds along multiple bands. The position of a point in the world coordinate system; , Let q be the azimuth and zenith angles in the spherical coordinate system centered at the camera's optical center. , Let be the maximum and minimum values of the azimuth angle in the coordinate system. , These are the maximum and minimum values of the zenith angle in the coordinate system.
[0069] The cross-sectional area of the apparent cone is approximately equal to the average distance from the ground. Area of the spherical piece at:
[0070]
[0071] in, .
[0072] Using DEM reference elevation Separate ground points from non-ground points to obtain the number of non-ground points. .Will and Substituting the canopy closure surrogate value formula from step 1.1, we iterate through all valid pixels to generate an image domain canopy closure map. H and W are the height and width of the image.
[0073] Step 1.3, LiDAR domain canopy closure modeling:
[0074] LiDAR domain canopy occlusion modeling uses bird's-eye view (BEV) ground grid cells as spatial units to quantify the degree of canopy occlusion in the direction of LiDAR observation.
[0075] According to the detection coordinate system, the ground is divided into Each cell has a spatial resolution of [number] cells. (That is, the grid size). It should be noted that... The fine mesh resolution used for offline modeling of DPM, compared to the resolution of the BEV feature map. These are two independent parameters—the former is used to accurately model the true value of occlusion intensity, and the latter determines the spatial granularity of the detected features. The ground elevation of each grid cell is provided by the DEM. For each grid cell... Define from the drone's location to ground position unit vector of observation direction .
[0076] Narrow cylindrical query frame: along Construct a narrow cylindrical query region The narrow cylinder is centered on the observation ray from the UAV to the center of the ground grid cell, and has a small lateral query radius (default 0.1m) to filter point cloud points close to this observation ray, preventing a large number of points from adjacent grid cells from being introduced. Non-ground points are retrieved from dense point clouds along multiple bands, and their numbers are counted. .make Substitute the canopy closure surrogate value formula from step 1.1, traverse all cells, and generate the LiDAR domain canopy closure map. Its spatial layout is aligned with the BEV detection point cloud space.
[0077] Step 1.4, Building Mask Correction
[0078] The Beer-Lambert model is suitable for randomly scattering media (such as vegetation canopies) but not for rigid obstructions such as buildings. Therefore, it is necessary to modify the coverage area of buildings.
[0079] Building point clouds are identified using semantic labels of point clouds. The view frustum query framework in step 1.2 and the narrow cylinder query framework in step 1.3 are used to determine whether each pixel or grid cell is covered by a building, thereby obtaining the corresponding initial building mask.
[0080] To improve the spatial continuity of the initial building mask, post-processing is performed: First, hole filling is applied to the connected regions of the buildings. This involves expanding the connectivity of the background region from the outer boundary of the initial building mask, identifying background meshes that are not connected to the outer boundary but are surrounded by building regions as internal holes and filling them with building regions. Then, morphological closing operations are performed on each connected region of the buildings. This involves first expanding the building regions to connect local narrow gaps, fractures, and small gaps, and then eroding them to restore their main boundaries. This process corrects local discontinuities while avoiding large-scale adhesion between different buildings, resulting in a more spatially continuous and complete building mask. It should be noted that the masking is performed separately for the two modalities. The image modality uses the same query frame as the image view frustum to obtain the building mask; the Lidar modality uses the same query frame as the narrow cylinder to obtain the building mask.
[0081] Then, the shading intensity value of the building's coverage area is corrected:
[0082]
[0083] Based on the above modified formula, for image domain canopy closure maps LiDAR domain canopy closure map The revised notation is as follows: and This is the final ground truth value for bimodal occlusion intensity, characterizing both canopy attenuation and rigid occlusion. This ground truth value has two uses in subsequent processes: as training supervision for DCPNet (step 2), and as a guiding signal for submodules used only during the training phase in OPF (steps 3.1 and 3.2). This ground truth value is not involved in the inference phase.
[0084] Step 2: Using the dual-modal canopy closure prediction network (DCPNet), with the ground truth of occlusion intensity generated in Step 1 as supervision, the intermediate features of the encoder are reused to predict the occlusion intensity distribution in the image domain and LiDAR domain respectively under single-frame inference conditions.
[0085] The specific process of step 2 includes:
[0086] The ground truth occlusion intensity generated in step 1 relies on computationally intensive frustum point cloud queries and requires dense point clouds across multiple bands, making it unsuitable for real-time online execution. DCPNet aims to predict the occlusion intensity distribution across modalities under single-frame inference conditions. To avoid introducing additional heavyweight perceptual branches, DCPNet reuses intermediate features from the detection network. For example... Figure 3 As shown, Figure 3 (a) in the diagram represents the image domain subnetwork (which includes multi-scale vegetation feature extraction and point cloud projection-assisted coding). Figure 3(b) in the diagram represents the LiDAR domain subnetwork (including terrain adaptive perspective projection and geometric statistical feature extraction).
[0087] The bimodal canopy closure prediction network comprises an image domain subnetwork and a LiDAR domain subnetwork. These subnetworks are structurally independent and integrated into the image branch and voxel branch, respectively. Each subnetwork uses the ground truth occlusion intensity of the corresponding modality from step 1 as training supervision. DCPNet runs during both the training and inference phases. The processing of the image domain subnetwork and the LiDAR domain subnetwork includes:
[0088] Step 2.1: Process using an image encoder and a LiDAR sparse encoder respectively:
[0089] Initial processing of the image domain sub-network: The image corresponding to the optical camera is fed into the image encoder to obtain the feature pyramid, and then multi-scale aggregated features are obtained. ;in and These are the height and width of the image feature map, respectively. Here, represents the number of channels. This implementation uses the Swing Transformer, but the present invention does not limit the specific structure of the image encoder; any image encoder capable of outputting multi-scale feature maps is applicable.
[0090] Initial processing of the LiDAR domain sub-network: The point cloud obtained by the LiDAR is fed into the LiDAR sparse encoder (voxel encoder) to obtain BEV features. The point cloud input to the LiDAR sparse encoder is denoted as the current frame point cloud. It is first voxelized into voxels and then input into the LiDAR sparse encoder. Optionally, the sparse encoder adopts a multi-stage sparse convolution and sparse residual block structure based on VoxelNext, extracting multi-scale 3D sparse features through progressive downsampling. In specific implementation, it can be implemented using the paper "Voxelnext: Fully sparse voxelnet for 3dobject detection and tracking" by Chen Y, Liu J, Zhang X et al., where sparse voxel features are converted into BEV features through elevation compression. It should be noted that Voxelnext contains features for obtaining terrain prediction maps. The invention includes, but is not limited to, Voxelnext. When using other encoders, it is sufficient to ensure that a corresponding terrain prediction branch is set in the sparse encoder. The terrain prediction branch can adopt existing technologies, such as "TG-ADet: Terrain-Guided Network for 3D Object Detection in ALS Point Clouds" by Jiang Y, Li X, Gu Y, et al. This invention will not provide a detailed description of the terrain prediction branch.
[0091] Step 2.2, Image Domain Canopy Closure Prediction Subnetwork Processing:
[0092] The purpose of the image domain subnetwork is to predict the occlusion intensity pixel by pixel.
[0093] Input: Multi-scale aggregated features output by the image encoder , and the point cloud of the current frame.
[0094] The processing procedure is divided into three parts:
[0095] The first part is multi-scale vegetation feature extraction. Apply convolution kernels with sizes of 1000 and 1000 respectively , and The three sets of parallel convolutions capture canopy texture details, medium-scale canopy morphology, and large-scale vegetation distribution information, respectively. The outputs of the three sets of convolutions... ( After being connected in series along the channel dimension, via Convolutional fusion to create vegetation saliency feature map .
[0096] The second part is point cloud projection-assisted feature encoding. The point cloud of the current frame is projected onto the image coordinate system, and the normalized number of projected points, normalized average depth, and normalized depth variance are counted at each pixel location to form a three-channel projection feature map.
[0097] More specifically, the initial 3D point cloud obtained by LiDAR is projected onto the image plane and divided into corresponding pixels or grid cells according to the multi-scale aggregated feature resolution. Then, the number of projection points within each pixel or grid cell is counted to obtain the number of projection points, which is then normalized using logarithmic compression. Simultaneously, the average depth value (depth of the point in the direction of the camera's optical axis) of all projection points within the cell is calculated to obtain the average depth, which is then normalized according to the maximum average depth within the effective region (grid cells in the feature map that contain at least one projection point). Furthermore, the dispersion of the depth value of each projection point within the cell relative to the average depth is calculated to obtain the depth variance, which is then normalized. Thus, a three-channel projection feature map is constructed using the normalized number of projection points, the normalized average depth, and the normalized depth variance.
[0098] The three-channel projection feature map is encoded into point cloud auxiliary features through two layers of convolution. Its spatial dimensions and Consistent. Point cloud projection auxiliary features help the subnetwork perceive occlusion boundaries from a geometric perspective, because the distribution of projection points in occluded areas differs statistically from that in open areas.
[0099] The third part is feature fusion and occlusion intensity prediction. , and After being concatenated along the channel dimension, the data is fused through two convolutional layers and finally activated by a Sigmoid activation function to output a single-channel occlusion intensity prediction map.
[0100] Output: .
[0101] During training, the ground truth value of occlusion intensity in the image domain After being aligned to the feature map resolution using bilinear interpolation, it serves as supervision. Loss function It consists of a weighted sum of mean squared error loss, edge perception loss, high occlusion region consistency loss, and smoothing term.
[0102] Step 2.3, LiDAR domain canopy closure prediction subnetwork:
[0103] The purpose of the LiDAR domain subnetwork is to predict occlusion intensity in the BEV space. Unlike the DPM offline modeling in step 1.3, this subnetwork does not query the point cloud along the view frustum. Instead, it starts from existing sparse voxels, projects them onto the terrain-defined ground grid, and extracts geometric and statistical features that indicate the local occlusion state.
[0104] Input: Shallow voxel features from the first layer output of the LiDAR sparse encoder and topographic prediction maps .in As a world coordinate system, These are the eigenvectors.
[0105] The processing procedure is divided into two parts:
[0106] The first part is a terrain-adaptive perspective projection. This projection layer is based on the drone's position. Using the ground grid defined by terrain prediction as the projection plane and the projection center as the projection center, a virtual perspective model is established. For the first... Individual elements, based on their horizontal position, are predicted by querying a terrain map. Obtain predicted ground elevation Thus, the effective focal length is determined. , The grid size is used to adaptively correct the impact of terrain undulations on the projection.
[0107] The projection coordinates of the virtual perspective model are calculated as follows:
[0108] ,
[0109] in, The coordinates of the UAV's position in the world coordinate system; Let be the coordinates of the UAV's ground projection point in the BEV grid, and denominator be . This determines the perspective scaling relationship along the drone's line of sight. The formula means that the closer a voxel is to the ground, the smaller the offset when projected onto the BEV mesh, and the farther away from the ground (e.g., high-altitude canopy voxels), the larger the offset, thus reflecting the spatial coverage relationship of occlusions in the BEV space.
[0110] Terrain adaptive projection was performed using a virtual perspective model, projecting all shallow voxel features onto the BEV mesh cells. For the mesh cells... Its corresponding voxel set is defined as , For projection onto BEV mesh cells The number of voxels on the surface Projecting shallow voxel features onto the corresponding mesh ;
[0111] The second part is geometric statistical feature extraction and occlusion intensity prediction. For projections onto the same grid cell... voxel set Feature vectors obtained based on projection The projection layer calculates the weighted average feature. And extract four geometrical statistical features related to canopy shading status. Let For unit Predicted ground elevation, For the first The relative ground height of individual units, The average relative height. Let be the height coordinate value of the m-th voxel.
[0112] The four geometric statistical characteristics are as follows:
[0113] Normalized point density: ;in, The maximum number of voxels in a single grid cell within the scene;
[0114] Normalized relative ground height: ;
[0115] Height standard deviation: ;
[0116] Ground confidence level: ;in This is an indicator function that takes the value 1 if the condition within the parentheses is true, and 0 otherwise. This is the height standard deviation threshold for ground confidence assessment, with a default value of 0.5. This is the height threshold for ground confidence assessment, with a default value of 0.5.
[0117] The above four features describe the local occlusion state from different perspectives: point density reflects the degree of canopy obstruction, relative height and standard deviation distinguish echoes near the ground from canopy echoes, and ground confidence determines whether the location is open ground.
[0118] Four geometrical statistical features and weighted average voxel features The data is concatenated along the channel dimension and then input into the prediction head. The prediction head mainly consists of three layers of sub-manifold sparse convolutions, which fuse projection features while maintaining sparsity, and finally compress it into a single channel and crop it to... The interval, i.e., the LiDAR domain occlusion intensity prediction map.
[0119] Output: , which is a prediction map of LiDAR domain occlusion intensity in sparse BEV format.
[0120] During training, For supervision. Loss function The mean squared error loss and L1 loss are used to assign higher weights to highly occluded areas to alleviate the problem of unbalanced distribution of occlusion levels.
[0121] Step 3: Using the Occlusion Prior Guided Fusion (OPF) strategy, the offline ground truth from Step 1 (used only in the training phase) and the online prediction from Step 2 (used in both the training and inference phases) are embedded into four stages: data augmentation, feature encoding, multimodal fusion, and detection head, to achieve adaptive multimodal detection under occlusion conditions.
[0122] The OPF strategy embeds the ground truth occlusion intensity values of DPM (including image domain and LiDAR domain ground truth values) and the occlusion intensity predictions of DCPNet into the four stages of the detection process through four dedicated sub-modules. The internal structure of each sub-module and its correspondence with the detection stages are as follows: Figure 4 As shown, Figure 4 Of the four sub-modules of the OPF strategy shown, (a) is for bimodal occlusion matching sample generation, (b) is for occlusion intensity-gated contrastive feature learning, (c) is for terrain-aware occlusion prior weighted fusion, and (d) is for occlusion-aware detection response modulation. Among them, (a) and (b) are used only in the training phase, while (c) and (d) are used in both the training and inference phases.
[0123] In this design, steps 3.1 (data augmentation) and 3.2 (feature encoding) use DPM ground truth only during the training phase, while steps 3.3 (multimodal fusion) and 3.4 (detection head) use DCPNet predictions during both the training and inference phases. This design allows offline ground truth to be used to improve training quality, while online predictions are used to guide actual detection.
[0124] Step 3.1, Generation of dual-modal occlusion matching samples (training phase only):
[0125] Sample pasting is a well-known data augmentation method in 3D object detection, enriching training data by pasting targets from a sample library to new locations in the training scene. However, in UAV-viewed scenarios, canopy occlusion causes significant differences in target point cloud density and image visibility at different locations. Standard sample pasting randomly selects and places samples without considering whether the occlusion state of the pasted samples matches the occlusion conditions of the placement location, potentially introducing a distribution mismatch between augmented and real samples. For example, pasting a complete target collected in an open area to a highly occluded area will produce a high-density point cloud and clear image features that do not match the actual observations of that area.
[0126] This invention proposes a sample generation method for bimodal occlusion matching, the specific process of which is as follows:
[0127] Sample pasting is a well-known data augmentation method in 3D object detection, enriching training data by pasting targets from a sample library to new locations in the training scene. However, in UAV-viewed scenarios, canopy occlusion causes significant differences in target point cloud density and image visibility at different locations. Standard sample pasting randomly selects and places samples without considering whether the occlusion state of the pasted samples matches the occlusion conditions of the placement location, potentially introducing a distribution mismatch between augmented and real samples. For example, pasting a complete target collected in an open area to a highly occluded area will produce a high-density point cloud and clear image features that do not match the actual observations of that area.
[0128] This invention proposes a sample generation method for bimodal occlusion matching, the specific process of which is as follows:
[0129] During the sample library construction phase, images, point clouds, and their corresponding C++ samples are included. img and C lid As a single sample, for each target, the mean, standard deviation, and Gaussian weighted mean at the target location are extracted from the ground truth maps of occlusion intensity in both the image and LiDAR domains, resulting in six statistics. These statistics are then weighted by modal weighting coefficients and concatenated to form a six-dimensional bimodal occlusion state descriptor. The modal weighting coefficients are manually set hyperparameters (0.75 for the LiDAR domain and 0.25 for the image domain in this embodiment), reflecting the relative reliability of the two modalities under canopy occlusion conditions, with the LiDAR domain contributing dominantly to the occlusion state description. KD trees are constructed according to the target detection category (man-made targets, camouflaged targets, canopy, vehicles): each tree uses the occlusion state descriptors of all samples in that category as nodes, and the Euclidean distance between nodes directly measures the similarity of occlusion states, thus supporting sample retrieval based on occlusion similarity.
[0130] During the sample generation phase, a location is randomly selected from the effective area of the DEM as the sample placement point. Its pixel coordinates are converted to a planar position in the LiDAR coordinate system, and the elevation of this point is used as the ground height. Combined with the target's own height, the vertical coordinates of the target bounding box center are determined. A bimodal occlusion state descriptor for this placement location is extracted. In the KD tree of this target detection category, several nearest-neighbor candidate samples with occlusion states are retrieved based on Euclidean distance—that is, samples whose occlusion conditions are closest to those of the placement location. Three-dimensional spatial collision detection and two-dimensional image overlap verification are performed on each candidate sample. The first sample that passes the verification is pasted. After pasting, the ground truth occlusion intensity in the image domain is updated by maximum value fusion in the corresponding region to maintain the consistency of the supervision signal.
[0131] The generated images and point clouds are also sent to the network processing in step 2.
[0132] Step 3.2, Occlusion Intensity Gated Contrast Feature Learning (Training Phase Only):
[0133] In drone scenarios, targets with similar geometric shapes but different colors and textures frequently exist (e.g., man-made targets and camouflaged targets have similar shapes but different colors). Image modalities are better than LiDAR at distinguishing such fine-grained differences. Contrastive learning can enhance the class discrimination ability of image encoders, causing image features of similar targets to cluster in the embedding space and distancing dissimilar targets. However, image features in severely occluded areas degrade significantly, and the visual appearance of the target no longer has class discrimination. Including these areas in contrastive learning will adversely affect feature space optimization. Therefore, this invention uses occlusion intensity as a sample quality gating signal to ensure that the contrastive loss is calculated only on samples with reliable image information.
[0134] The specific process is as follows: The eight corner points of the target's 3D bounding box are projected onto the image plane using a calibration matrix, and the two-dimensional bounding rectangle of the projected corner points is taken as the RoI region of the target in the corresponding viewpoint. The RoI region is the two-dimensional bounding rectangle obtained after projecting the 3D bounding box onto the image plane; RoI samples are training samples composed of local image features extracted from this region and their category labels. A valid RoI sample is one where the projected region has a valid intersection with the image range and has a non-zero width and height after cropping.
[0135] For each valid RoI sample Query the true value of the image domain occlusion intensity at the center of its two-dimensional projection frame. Only retain occlusion intensity below the threshold. The samples were used for contrastive learning. Subsequently, RoIAlign was used to extract features from the image feature maps. Local image features of a fixed size are extracted and mapped into normalized embedding vectors by a projection head. After gating (less than) In the sample set, RoI samples of the same class are treated as positive samples, and RoI samples of different classes are treated as negative samples. A weighted InfoNCE loss is used:
[0136]
[0137] in This is the gated index set. The image projection coordinates of the target center. RoI samples The set of positive samples ( (for 3D bounding box category labels) Category-specific weights are assigned to the main categories that need to be distinguished, while the auxiliary categories are assigned lower weights to participate in training. This is used to expand the contrast relationship between categories and stabilize the learning of the embedding space. This is the temperature coefficient for InfoNCE loss, with a default value of 0.07.
[0138] Step 3.3, Terrain-Aware Occlusion Prior Weighted Fusion (used in both training and inference phases):
[0139] Existing fusion methods based on a unified feature space (such as BEVFusion) transform image features to the BEV space through depth distribution prediction. In UAV-view scenarios, this paradigm faces three problems: large uncertainty in depth estimation, high computational cost of large-size BEV feature maps, and image feature degradation under severe occlusion leading to the introduction of erroneous information into the fusion result. This invention designs a terrain-aware occlusion prior weighted fusion method, which uses sparse LiDAR BEV features as a starting point to query image features in reverse, avoiding depth estimation. The computational cost is proportional to the sparsity of the BEV features.
[0140] Input: Sparse BEV features output from the LiDAR sparse encoder, terrain prediction map Multi-scale aggregation features Image domain occlusion intensity prediction .
[0141] The processing is divided into three stages:
[0142] Occlusion intensity modulation stage: Interpolation to image features After resolution, and The network is generated by concatenating the data along the channel dimension and then inputting it into a dynamic weight generation network.
[0143]
[0144] in For interpolation processing and two-layer convolutional encoders (used to adjust resolution). For the Sigmoid function, Indicates serial connection along the channel dimension. This is the lower bound of the weight (default value is 0.05).
[0145] The resulting weighted graph and With consistent resolution, the reliability of image information at each location is reflected simultaneously in both spatial and channel dimensions: the weights of severely occluded areas converge. The weights for lightly occluded and open areas approach 1. Compared to simply using... As scalar weights, dynamic weight networks allow different feature channels to learn different occlusion sensitivities. Applying weights to the graph The weighted image features are obtained as follows:
[0146]
[0147] Projection sampling stage: The size of the sparse BEV feature is ,in For non-empty BEV cell elements, The feature dimension is defined as follows. For each sparse BEV cell, its world coordinates are extracted. Query terrain prediction map Obtain predicted ground elevation plus height offset The near-ground coordinates were then obtained. Terrain prediction provides near-ground 3D coordinates at each BEV cell location, allowing sampled image features to correspond to the most likely spatial location of the target. However, assuming a fixed ground height can lead to image sampling position shifts in drone scenarios with significant terrain undulations. (The query coordinates are then used to further refine the prediction.) Projected onto the intrinsic and extrinsic parameter matrices of the optical camera The corresponding image plane is weighted by bilinear interpolation in the image features. Sampling is performed on the above, and the size of each query point is obtained as follows: eigenvectors. All The set of sampling results from each query point constitutes the image modality weighted fusion feature. The size is .
[0148] Feature fusion stage: Concatenated with sparse BEV features along the channel dimension, resulting in a size of The joint features are fused through sparse convolution of sub-manifolds, which enables cross-modal feature interaction while maintaining the sparse structure.
[0149] Output: Sparse BEV fusion features .
[0150] Step 3.4, Occlusion Perception Detection Response Modulation (used in both training and inference phases):
[0151] After the fused features are input into the detection head, they enter the heatmap classification branch and the 3D bounding box parameter regression branch, respectively. It should be noted that the input of the heatmap classification branch is based on sparse BEV fused features. Received The heatmap classification branch is used to determine the category of the detection box, while the 3D box parameter regression branch is used to determine the location, size, and orientation of the detection box.
[0152] Targets in severely occluded areas exhibit degraded BEV features due to sparse point clouds, resulting in low heatmap responses from detection networks and a high risk of missed detections. This invention proposes an occlusion-aware detection response modulation strategy that utilizes LiDAR domain occlusion intensity prediction from DCPNet as a conditional prior to compensate for and enhance the heatmap classification branch.
[0153] For fusion features Each valid BEV position in Sampling the predicted occlusion intensity at the corresponding location , for Middle position The value of is used. Severely occluded regions are filtered out using a soft gating function, and auxiliary features are generated by a lightweight mapping network and superimposed onto the input of the heatmap classification branch.
[0154]
[0155] in, for Middle position The value, For multilayer perceptron mapping networks, the gated occlusion prior is used. (The predicted value is actually the predicted occlusion prior) mapped to... Same-dimensional feature residuals; The default value for the gate threshold used to detect response modulation is 0.6; is the modulation intensity coefficient, where Learnable parameters initialized to zero. This is the upper bound for modulation (default value is 0.1). Zero-initialization design ensures that the modulation branch is equivalent to the identity mapping in the early stages of training. hour This maintains the stability of the pre-trained feature distribution and gradually learns the appropriate modulation amplitude as training progresses.
[0156] shading intensity is lower than The location does not pass through The mapping retains its heatmap features; the occlusion intensity is higher than The location receives feature compensation proportional to the degree of occlusion, helping the network generate higher target responses in these regions. This modulation is applied only to the heatmap classification branch, while the regression branch (including center offset, center height, 3D size, and orientation) remains unaffected. This limits the impact of occlusion priors to the target discrimination stage, avoiding interference with the accuracy of geometric regression.
[0157] During training, the total loss function includes not only the contrastive loss but also the classification loss. Ground truth of image domain occlusion intensity. After being aligned to the feature map resolution by bilinear interpolation, it is used as supervision for image domain occlusion intensity prediction. Classification loss function The total loss function is primarily composed of a weighted sum of mean squared error loss, edge perception loss, high occlusion region consistency loss, and a smoothing term. In addition, the total loss function also includes occlusion intensity loss from the LiDAR domain, terrain prediction loss, and classification and regression losses from the detection head. It should also be noted that the total loss function can be improved and adjusted based on actual training results; for example, it may also include other losses.
[0158] Step 4: The inference stage takes only a single frame of LiDAR point cloud and a single frame of image as input. The input passes through the LiDAR sparse encoder and the image encoder, respectively, and then through the occlusion prior weighted fusion module and the occlusion perception detection head to complete the end-to-end detection.
[0159] Step 4 (complete reasoning process) is implemented through the following procedure:
[0160] The inference phase does not require multi-band dense point clouds and ground truth occlusion intensity values; it only uses a single-frame image and a single-frame LiDAR point cloud as input. Steps 3.1 (Binocular Occlusion Matching Sample Generation) and 3.2 (Occlusion Intensity Gated Contrast Feature Learning) are not involved in the inference phase. The complete inference flow is as follows:
[0161] (1) After voxelization (voxel size is 0.1 m) and voxel feature encoding, the LiDAR point cloud is input into the LiDAR sparse encoder. The sparse encoder adopts a multi-stage sparse convolution and sparse residual block structure based on VoxelNeXt, and extracts multi-scale three-dimensional sparse features by progressive downsampling.
[0162] The terrain prediction branch compresses multi-scale 3D sparse features into BEV sparse features along the height dimension. Using the terrain prediction branch proposed by the TG-ADet method, it predicts the terrain distribution of the scene, i.e., the terrain prediction map. .
[0163] LiDAR sparse encoder based on Elevation attention weights are constructed: using the terrain prediction elevation plus offset as the Gaussian distribution center, voxel features at different altitudes are weighted, giving greater weight to features closer to the ground. The weighted 3D sparse features are accumulated along the height dimension and output as the final sparse BEV feature through sparse convolution.
[0164] (2) LiDAR domain canopy closure prediction subnetwork (step 2.3) Read the shallow voxel features and terrain prediction map output from the first layer of the sparse encoder. After terrain-adaptive perspective projection and geometric statistical feature extraction, the output is a LiDAR domain occlusion intensity prediction. .
[0165] (3) A single frame image is processed by an image encoder and a feature pyramid network to output multi-scale aggregated features. .
[0166] (4) Image domain canopy closure prediction subnetwork (step 2.2) Taking the current frame's point cloud projection features as input, after multi-scale vegetation feature extraction and point cloud projection-assisted encoding, the output is a predicted image domain occlusion intensity. .
[0167] (5) Terrain-aware occlusion prior weighted fusion module (step 3.3) uses sparse BEV features, , and Using query point generation, projection sampling, and occlusion intensity modulation as input, sparse BEV fusion features are generated through three stages: query point generation, projection sampling, and occlusion intensity modulation. .
[0168] (6) Occlusion sensing detection head (step 3.4) and As input, occlusion prior modulation is injected into the heatmap classification branch. The detector head outputs prediction results for six branches: category, center offset, center height, 3D dimensions, orientation, and IoU quality score.
[0169] (7) Decode the output of the detection head and suppress non-maximum values to obtain the final three-dimensional detection box.
[0170] The training strategy of this invention adopts a three-stage process to ensure the stable convergence of each component: The first stage trains the single-modal LiDAR detection network (including the terrain prediction branch) for 20 rounds, enabling the LiDAR sparse encoder and detector head to obtain good initial feature representation capabilities; the second stage freezes the detection-related modules and independently trains the bimodal canopy closure prediction sub-network for 20 rounds, enabling DCPNet to learn to fit the ground truth of occlusion intensity generated by DPM; the third stage jointly trains the complete multimodal fusion network for 10 rounds, setting differentiated learning rates for different modules—protecting the already trained LiDAR backbone and canopy closure prediction module (learning rate multiplier 0.5), and accelerating the adaptation of randomly initialized fusion layers to multimodal features (learning rate multiplier 1.5).
[0171] Compared with the prior art, the present invention has the following advantages:
[0172] First, this method invented the first multimodal fusion 3D target detection framework of lidar-optical camera for single UAV platform, realizing the effective utilization of complementary information of two modes in UAV scenarios.
[0173] Second, this invention constructs a directly usable bimodal occlusion prior for the detection network through physically inspired Beer-Lambert modeling, enabling the network to perceive the degree of information degradation of each modality at each spatial location, thus providing a physical basis for adaptive occlusion detection.
[0174] Third, the occlusion prior-guided fusion strategy designed in this invention covers the four stages of the detection process, enabling the detector to adaptively adjust according to occlusion conditions in data augmentation, feature encoding, multimodal fusion, and detection head stages, especially showing more significant advantages in severe occlusion and difficult samples.
[0175] To verify the performance of the algorithm proposed in this invention, a dataset was constructed and experiments were conducted. The experimental results verified the effectiveness of the proposed method.
[0176] 1. Experimental Dataset
[0177] This invention uses a self-built dataset, SI3D-DI, for experimental verification. This dataset was collected by a DJI Matrice 350 RTK drone equipped with a Zenmuse L2 sensor. The scene contains four target categories: vehicles, man-made targets, camouflaged targets, and celestial dome. Samples were generated by accumulating time windows of one second according to image timestamps, resulting in a dataset containing 4892 training samples and 3765 test samples.
[0178] 2. Experimental Setup
[0179] Evaluation metrics: Average accuracy of bird's-eye view (AVR) ) and three-dimensional average accuracy ( As a performance indicator, 40-point interpolation was used for calculation. The cross-union ratio (CUI) threshold for all categories was set to 0.5. The test set was evaluated along two dimensions: the first dimension classified targets into three difficulty levels—"easy," "medium," and "hard"—based on the number of points inside the target; the second dimension classified targets into "unobstructed" based on the occlusion intensity of the LiDAR domain in the target's neighborhood. "Partial occlusion" ) and "high occlusion" ( There are three levels.
[0180] Implementation details: Voxel size was set to 0.1 m. Point cloud detection range (XYZ axes) was set to [-64, -16, 0, 64, 16, 16] (m). The original image was downsampled by 2 times and input into the network at 2640×1978 pixels, and normalized using ImageNet standard parameters. All experiments were performed on two NVIDIA V100 GPUs using the AdamW optimizer.
[0181] 3. Results Analysis
[0182] 3.1 Comparative analysis with existing methods:
[0183] Table 1 presents the performance analysis results of the 3D target detection methods on the SI3D-DI test set. CAMF-Det, compared with several representative methods on the SI3D-DI test set, achieved the best results at all three difficulty levels. They reached 71.48%, 70.14%, and 63.87% respectively. They reached 58.64%, 57.09%, and 50.78%, respectively. Compared to the single-modal baseline TG-ADet, The improvement was 8.11%, 8.76%, and 9.43% for easy, medium, and difficult difficulty levels, respectively, with the largest improvement observed at the difficult level, indicating that the method is most effective in improving targets with sparse point clouds and severe occlusion.
[0184] It is noteworthy that existing multimodal methods generally perform poorly on this dataset, failing to surpass pure LiDAR baselines in multiple metrics. The fundamental reason is that these methods assume that cross-modal information space is uniform and reliable, and do not consider image feature degradation caused by canopy occlusion. CAMF-Det, by explicitly modeling occlusion priors, fully leverages its complementary advantages in regions with reliable image information and actively suppresses its contribution in degraded regions, thereby consistently outperforming baseline methods across all categories and difficulty levels.
[0185] A visualization comparison of the detection results of BEVFusion, TG-ADet, and the present invention on the SI3D-DI dataset is shown below. Figure 5 As shown.
[0186] 3.2 Ablation Experiment:
[0187] Two sets of ablation experiments were conducted to verify the contributions of the overall framework and the OPF submodule, respectively.
[0188] The first set of overall framework ablation analysis shows the multimodal baseline (TG-ADet plus image branching and base fusion). The improvements were 65.60%, 63.85%, and 56.80% respectively, showing limited gains compared to TG-ADet. The slight decrease in performance after introducing DCPNet alone indicates that DCPNet's occlusion prediction only generates positive gains through explicit OPF embedding in the detection process. The complete CAMF-Det achieved improvements of 5.88%, 6.29%, and 7.07% compared to the multimodal baseline, demonstrating that the synergy between DCPNet and OPF is the core source of the performance gains.
[0189] The second set of OPF submodules features progressive ablation display, with occlusion prior weighted fusion applied to varying difficulty levels. The occlusion prior contributes the most, improving performance by 4.69%, and is the core mechanism for achieving occlusion perception fusion. Detection response modulation, occlusion matching data enhancement, and occlusion gating contrastive learning contribute 1.50%, 1.13%, and 2.11% respectively at the hardness level. The four sub-modules work together to embed occlusion priors into different stages of the detection process, jointly achieving stable performance improvement under complex occlusion conditions.
[0190] Table 1 Performance analysis of the 3D target detection method on the SI3D-DI test set, with the cross-union threshold of the target set to 0.5.
[0191]
[0192]
[0193]
[0194] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for 3D target detection using a UAV lidar-optical camera based on canopy closure perception, characterized in that, include: The processing steps of the bimodal canopy closure prediction network include: Step 2.1: Input the image corresponding to the optical camera into the image encoder to obtain the feature pyramid, and then obtain the multi-scale aggregated features. The current frame point cloud is converted into voxels and then input into a LiDAR sparse encoder to obtain sparse BEV features. Step 2.2, for Parallel convolutions with different kernel sizes are performed separately and concatenated along the channel dimension, then fused together by convolution to form a vegetation saliency feature map. Project the current frame point cloud onto the image coordinate system, and count the normalized number of projected points, normalized average depth, and normalized depth variance at each pixel location to form a three-channel projected feature map, which is then encoded as point cloud auxiliary features. ;Will , and After being concatenated along the channel dimension and then fused by convolution, a single-channel occlusion intensity prediction map is obtained after passing through an activation function. ; Step 2.3, UAV location A virtual perspective model is established using the terrain-predicted ground grid as the projection plane, with the projection center as the model; for the first... Individual elements, based on their horizontal position, are predicted by querying a terrain map. Obtain predicted ground elevation To obtain the effective focal length ,based on A virtual perspective model is defined and terrain-adaptive projection is performed, projecting the shallow voxel features output from the first layer of all LiDAR sparse encoders onto BEV mesh cells. middle; The corresponding set of voxels is , for The number of voxels on the surface Projecting shallow voxel features onto the corresponding mesh For projections onto the same grid cell ,based on Calculate the weighted average characteristic The geometrical and statistical features related to the canopy occlusion state are then concatenated along the channel dimension and input into the prediction head to obtain the LiDAR domain occlusion intensity prediction map. ; The occlusion-prior-guided fusion strategy network processing steps include: Step 3.3, will Adjust to After resolution, with The weight graph is obtained by concatenating the weights along the channel dimension and then inputting them into a dynamic weight generation network. Used for Weighted image features are obtained by weighting. For each sparse BEV cell, extract its world coordinates and predict them based on the terrain map. Obtain predicted ground elevation plus height offset The near-ground coordinates were then obtained. ;Will Projected to The corresponding image plane, in Sampling is performed on the data, and a corresponding feature vector is obtained for each query point; the set of sampling results for all query points constitutes the image modality weighted fusion feature. ;Will By concatenating the sparse BEV features along the channel dimension, joint features are obtained. These features are then fused using sparse convolution of submanifolds to obtain sparse BEV fused features. ; Step 3.4: Based on sparse BEV fusion features The detection is completed using a detection head. The heatmap classification branch in the detection head is used to determine the category of the detection box, and the 3D box parameter regression branch in the detection head is used to determine the position, size and orientation of the detection box.
2. The UAV LiDAR-optical camera 3D target detection method based on canopy closure perception according to claim 1, characterized in that, The process of using a detection head to complete the detection includes: Sparse BEV fusion features The position, size, and orientation of the detection box are determined by inputting the 3D bounding box parameter regression branch; Inject occlusion prior modulation into the input of the heatmap classification branch, i.e. ,in for Middle position The value, To implement the occlusion prior after gating Mapping to AND Same-dimensional feature residual network, for Middle position The value; To detect the gating threshold of the response modulation; is the modulation intensity coefficient, where Learnable parameters initialized to zero. To modulate the upper bound; based on sparse BEV fusion features Received The category of the detection box is determined by sending it to the heatmap classification branch.
3. The UAV lidar-optical camera three-dimensional target detection method based on canopy closure perception according to claim 1, characterized in that, Geometric statistical features related to canopy shading status include: Normalized point density Normalized relative ground height Standard deviation of height Ground confidence level ; in, The maximum number of voxels in a single grid cell within the scene; For the first The relative ground height of individual units, The average relative height. Here is the height coordinate value of the m-th voxel; For unit Predicted ground elevation; This is an indicator function that takes the value 1 if the condition within the parentheses is true, and 0 otherwise. The height standard deviation threshold in ground confidence assessment; This is the height threshold used in ground confidence assessment.
4. A method for three-dimensional target detection using a UAV lidar-optical camera based on canopy closure perception, as described in any one of claims 1 to 3, characterized in that, The bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network are pre-trained, and the training process includes: Ground truth occlusion intensity values in the image and LiDAR domains are constructed offline using dense point clouds with multiple stripes. Then, the occlusion intensity values for building-covered areas are corrected. The corrected ground truth occlusion intensity values in the image domain are then used as the basis for the final determination. True values of LiDAR domain occlusion intensity ; Image and point cloud data are fed into a bimodal canopy closure prediction network for processing; the ground truth occlusion intensity in the image domain is then calculated. After being aligned to the feature map resolution by bilinear interpolation, the classification loss is obtained as supervision. ; Perform occlusion intensity gating contrast feature learning, including: The eight corner points of the target's 3D bounding box are projected onto the image plane using a calibration matrix, and the two-dimensional bounding rectangle of the projected corner points is taken as the RoI region of the target in the corresponding viewpoint; for each valid RoI sample Query the true value of the image domain occlusion intensity at the center of its two-dimensional projection frame. Only retain the true value of occlusion intensity below the threshold. The samples were used for contrastive learning; subsequently, RoIAlign was used to analyze the feature maps. Local image features of a fixed size are extracted and mapped into normalized embedding vectors by a projection head. In the sample set participating in contrastive learning, RoI samples of the same class are treated as positive samples, and RoI samples of different classes are treated as negative samples. The weighted InfoNCE loss function is used as the contrastive learning loss. ; Based on the output of the bimodal canopy closure prediction network, an occlusion prior-guided fusion strategy network makes predictions to obtain the category, location, size, and orientation of the detection boxes; based on a learning loss... and classification loss The total loss completes the training of the bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network.
5. The UAV lidar-optical camera three-dimensional target detection method based on canopy closure perception according to claim 4, characterized in that, The process of offline construction of ground truth occlusion intensity values in the image and LiDAR domains using multi-strip dense point clouds includes: Step 1.1: Construct a canopy closure model ,in The effective attenuation coefficient is a composite parameter of the absorption extinction coefficient and the unit conversion factor. The reference point density is used for density normalization across different scenarios. This represents the average point density of the current scene. and To construct the view frustum The number of non-ground points and cross-sectional area inside; Step 1.2: Construct the view frustum query framework: The four ray directions corresponding to the four corner points determine the pixel view frustum. angular range in spherical coordinates , Let be the maximum and minimum values of the azimuth angle in the coordinate system. , The maximum and minimum zenith angles in the coordinate system; based on the view frustum query framework, from multi-band dense point clouds Query the set of points that satisfy the angle constraint Average ground distance The area of the spherical patch at that point is used as the cross-sectional area of the frustum. Using DEM reference elevation Separate ground points from non-ground points to obtain the number of non-ground points. ;Will and Substitute the canopy closure model, traverse all valid pixels, and generate an image domain canopy closure map. , that is, the true value of the occlusion intensity in the image domain; Step 1.3: Construct a narrow cylindrical query framework: Construct a narrow cylindrical query region along the unit vector of the observation direction. The narrow cylinder is defined as the region with the observation ray from the UAV to the center of the ground grid as its central axis and a lateral query radius smaller than a set threshold; non-ground points are retrieved from multiple dense point clouds and their numbers are counted. The area of the ground grid used for LiDAR domain canopy closure modeling is taken as the cross-sectional area of a narrow cylinder. Substitute the data into the canopy closure model, traverse all cells, and generate a LiDAR domain canopy closure map. That is, the true value of the occlusion intensity in the LiDAR domain; Fine-grid resolution used for offline modeling of a dual-modal canopy closure prediction network.
6. The UAV lidar-optical camera three-dimensional target detection method based on canopy closure perception according to claim 5, characterized in that, Canopy Closure Model The construction process includes: Based on reference point density Canopy thickness along the observation direction And the number of non-ground points inside the view frustum constructed along the observation direction. and cross-sectional area Determine the cumulative leaf area index The complement of transmittance is defined as canopy closure, and a canopy closure model is then constructed. .
7. The UAV lidar-optical camera three-dimensional target detection method based on canopy closure perception according to claim 6, characterized in that, Average ground distance The area of the spherical piece at that location is ,in ; The fine grid resolution used in offline modeling based on the bimodal canopy closure prediction network corresponds to the grid area of the square. , Fine-grid resolution used for offline modeling of a dual-modal canopy closure prediction network.
8. The UAV lidar-optical camera three-dimensional target detection method based on canopy closure perception according to claim 4, characterized in that, In the process of correcting the occlusion intensity value of the building coverage area, the occlusion intensity value of the building coverage area is corrected to 1 based on the occlusion intensity in the image domain and LiDAR domain, thus obtaining the corrected occlusion intensity value.
9. A method for three-dimensional target detection of a UAV LiDAR-optical camera based on canopy closure perception according to claim 4, characterized in that, The classification loss function It mainly consists of the weighted sum of mean square error loss, edge perception loss, high occlusion region consistency loss, and smoothing term.
10. A method for three-dimensional target detection of a UAV LiDAR-optical camera based on canopy closure perception according to claim 4, characterized in that, The training process of the bimodal canopy closure prediction network and the occlusion prior-guided fusion strategy network also includes a step of generating samples to build a sample library. Sample generation is performed using a bimodal occlusion matching method, including: During the sample library construction phase, images, point clouds, and their corresponding C++ samples are included. img and C lid As a single sample, for each target, the mean, standard deviation, and Gaussian weighted mean at the target location are extracted from the ground truth maps of occlusion intensity in both the image and LiDAR domains, resulting in six statistics. These statistics are then weighted by modal weighting coefficients and concatenated to form a six-dimensional bimodal occlusion state descriptor. KD trees are constructed according to the target detection category: each tree uses the occlusion state descriptors of all samples in that category as nodes, and the Euclidean distance between nodes directly measures the similarity of occlusion states, thereby supporting sample retrieval based on occlusion similarity. In the sample generation stage, a location in the scene is randomly selected from the effective area of the DEM as the sample placement point. Its pixel coordinates are converted into a planar position in the LiDAR coordinate system, and the elevation of this point is used as the ground height. Combined with the height of the target itself, the vertical coordinates of the target box center are determined. The bimodal occlusion state descriptor of this placement location is extracted. In the KD tree of this target detection category, several candidate samples with the nearest neighbor of the occlusion state are retrieved based on Euclidean distance. These are the samples that are closest to the occlusion conditions of the placement location. Three-dimensional spatial collision detection and two-dimensional image overlap verification are performed on each candidate sample. The first sample that passes the verification is pasted to achieve sample generation. After pasting, the ground value of the occlusion intensity in the image domain is updated by maximum value fusion in the corresponding region.