Low-slow small target detection and trajectory prediction tracking method based on laser radar

By combining lidar point cloud data preprocessing with deep learning and extended Kalman filters, the problem of sparse target perception and tracking interruption in low, slow and small target detection is solved, achieving high-precision, stable tracking and all-weather monitoring of low, slow and small targets.

CN121069407AActive Publication Date: 2025-12-05CHINA UNIV OF MINING & TECH

Patent Information

Application Number
CN202511606088.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2025-12-05
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing lidar technology suffers from problems such as insufficient perception of sparse targets, frequent tracking interruptions, and low recognition accuracy in detecting low-speed, small targets, making it difficult to achieve stable and accurate monitoring in complex environments.

Method used

A method combining lidar point cloud data preprocessing, deep learning detection network and extended Kalman filter is adopted to improve the recognition ability of weakly reflective and small-volume targets and achieve stable tracking through point cloud saliency screening, pillar construction and encoding, multi-scale feature extraction and trajectory prediction.

Benefits of technology

It improves the detection accuracy and tracking stability of low, slow, and small targets, enabling efficient detection and continuous tracking of such targets, enhancing their adaptability to complex environments, and providing all-weather monitoring capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069407A_ABST
    Figure CN121069407A_ABST
Patent Text Reader

Abstract

The invention discloses a low-slow small target detection and trajectory prediction tracking method based on a laser radar, and the method comprises the steps: firstly collecting the point cloud data of the laser radar, and carrying out the preprocessing of spatial modeling and coordinate transformation of the point cloud data of the laser radar; performing significance screening; based on distance partition driving, Pilllar construction and coding are carried out; constructing a deep learning detection network; based on Anchor design and a matching strategy, carrying out size adaptation on a weak target in the air in the fused features; and performing time sequence prediction and observation updating on the target state based on an extended Kalman filter (EKF), and completing low-slow small target detection and trajectory prediction tracking. The method can maintain the high precision advantage of the laser radar, improves the recognition capability of the laser radar on weak-reflection, small-size and irregular-motion targets, has high robustness and environment adaptability, and achieves the stable and precise sensing and continuous tracking of low, slow and small flight targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of control theory and signal processing, and particularly relates to a method for detecting and predicting the trajectory of small, slow targets based on lidar. Background Technology

[0002] With the widespread application of small flight platforms such as drones and model aircraft, new types of flying targets characterized by "low altitude, slow speed, and small size" (referred to as "low, slow, and small") are frequently appearing in urban, transportation, energy, and military fields. Due to their small size, low flight altitude, slow speed, weak radar reflection characteristics, and ease of concealment, these targets are increasingly posing a potential threat to current airspace security management. Especially in urban environments, low, slow, and small targets can easily bypass the monitoring of traditional air defense systems, creating serious security risks. Therefore, how to achieve high-precision, all-weather, and all-time-domain detection and stable tracking of low, slow, and small targets has become an important research direction in the fields of intelligent sensing and low-altitude security.

[0003] Among numerous detection technologies, LiDAR (LiDAR) stands out for its high precision, high spatial resolution, and independence from lighting conditions, demonstrating immense potential in target 3D perception and spatial localization. Compared to traditional millimeter-wave radar, LiDAR can acquire high-density point cloud data, possessing stronger object contour reconstruction capabilities and making it suitable for geometric feature extraction and motion trajectory tracking of small targets in complex environments. Especially in obstacle-ridden urban environments, LiDAR can effectively acquire spatial location information of targets, assisting in accurately determining their position, attitude, and motion state.

[0004] However, lidar also faces a series of challenges when used for detecting low-altitude, slow-moving, and small targets. On the one hand, due to their small size and irregular flight trajectories, low-altitude, slow-moving, and small targets reflect weak laser signals, appearing sparse and discontinuous in point clouds. They are easily obscured by background ground points, building reflections, or tree canopies, resulting in low detection accuracy. On the other hand, limited by the lidar's scanning frequency and field of view, when the target moves rapidly, changes direction abruptly, or is briefly occluded, the differences between point cloud data frames are large, significantly hindering continuous tracking and state prediction. Furthermore, laser echoes in low-altitude environments are easily affected by multipath reflections and rain / fog obstruction, further complicating target identification.

[0005] While some studies have attempted to combine point cloud segmentation, feature clustering, and deep learning methods to improve the detection capabilities of lidar for low, slow, and small targets, they still face numerous challenges in practical applications. For example, when low, slow, and small targets have weak reflectivity, are tiny, or have complex trajectories, existing methods often rely on strong prior information or ideal environmental conditions, making it difficult to guarantee the stability of detection and tracking in general scenarios. Furthermore, many existing algorithms are primarily designed for ground targets (such as vehicles and pedestrians), which differ fundamentally from low, slow, and small aerial targets in terms of target scale, motion patterns, and spatial distribution, resulting in limited generalization ability in aerial target scenarios.

[0006] Compared to lidar, traditional visible light imaging systems, while providing rich texture and color information, suffer from a sharp decline in performance under unstable lighting conditions, backlighting, or nighttime conditions, and lack three-dimensional spatial positioning capabilities. Especially when the color contrast between the target and the background is not significant, this can easily lead to false positives and false negatives, and target loss can occur during brief periods of obstruction or high-speed flight.

[0007] Event cameras, with their high temporal resolution and sensitivity to dynamic targets, can quickly respond to sudden movement events. However, they have limited capabilities in representing the structural and textural features of targets, making them suitable only for assisting in determining target trajectories, not as the primary detection method. Infrared thermal imaging is suitable for nighttime and low-visibility environments, but its accuracy in identifying long-distance targets is limited due to issues such as weak thermal radiation and low thermal contrast of small, slow-moving targets. While radio frequency and acoustic signature detection offer advantages such as all-weather operation and low cost, their accuracy in noisy environments is low, and they cannot provide precise spatial location information, making high-precision tracking difficult.

[0008] In summary, in complex urban or low-altitude environments, the detection and tracking of low-speed, small targets requires systems to simultaneously possess high spatial resolution, wide field-of-view coverage, strong anti-interference capabilities, target geometry reconstruction capabilities, and stable tracking capabilities. LiDAR, with its rich 3D spatial information and ability to detect non-cooperative targets, is one of the important technological pathways to achieve this goal. However, existing LiDAR detection technologies, when applied to low-speed, small target scenarios, still suffer from insufficient detection capabilities for sparse targets, frequent tracking interruptions, and low recognition accuracy, making it difficult to support the urgent need for continuous and reliable monitoring of low-speed, small targets in real-world combat environments. Summary of the Invention

[0009] Purpose of the Invention: The purpose of this invention is to provide a method for detecting and predicting the trajectory of low-speed, small targets based on lidar. This method retains the high precision advantage of lidar while enhancing its ability to identify targets with weak reflection, small size, and irregular movement, and possesses strong robustness and environmental adaptability, thereby achieving stable, accurate perception and continuous tracking of low-speed, small flying targets.

[0010] Technical solution: The present invention provides a method for detecting and predicting the trajectory of small, slow targets based on lidar, comprising the following steps: Step 1: Acquire lidar point cloud data and perform spatial modeling and coordinate transformation preprocessing on the lidar point cloud data;

[0011] Step 2: Perform saliency filtering on the preprocessed data to obtain the filtered point cloud data.

[0012] Step 3: Based on distance partitioning, the filtered point cloud data is constructed and encoded using Pillar to obtain the Pillar representation vector.

[0013] Step 4: Construct a deep learning detection network, using the pillar representation vector as input and outputting the fused features.

[0014] Step 5: Based on the Anchor design and matching strategy, perform size adaptation on the weak targets in the air in the fused features.

[0015] Step 6: Based on the extended Kalman filter (EKF), perform time-series prediction and observation update of the target state to complete the detection and trajectory prediction and tracking of low, slow, and small targets.

[0016] Furthermore, step 1 specifically involves measuring the vertical height of the radar from the mounting reference plane to the bottom of the sensor. Record the elevation angle at the initial installation of the radar. Horizontal rotation deflection And the horizontal position offset of the radar relative to the origin of the carrier coordinate system. The data is input into the calibration module to construct the extrinsic parameter matrix of the radar sensor in the platform coordinate system. During point cloud acquisition, each frame of raw point cloud data is represented in local coordinate form as follows:

[0017] .

[0018] in, For each frame of point cloud raw data, , , Provide the x-coordinate, y-coordinate, and height information for each point cloud in the local coordinate system. The reflection intensity of each point cloud, where N is the number of point clouds in each frame of raw point cloud data; read the extrinsic parameters. And perform coordinate transformations on each point:

[0019] .

[0020] The transformed point is located in the platform's global space. .

[0021] Calculate the vertical height of the point relative to the ground: .

[0022] Calculate the spatial distance for each point: Thresholds set based on experience The point cloud is divided into three spatial regions: .

[0023] in, For the spatial region of point cloud, These are threshold limits for point clouds at close, medium, and long distances, set based on experience.

[0024] Furthermore, step 2 specifically involves: after each frame of point cloud acquisition is completed, scanning the reflection intensity values ​​of all points and extracting the maximum value for that frame. and minimum value The reflection intensity at each point Normalize: .

[0025] KD-Tree neighborhood search is used to count the number of neighboring points in the three-dimensional neighborhood of each point. : .

[0026] in, For the target point whose significance is currently being calculated, For point clouds Any other point in the array;

[0027] Normalize the number of neighborhood points: .

[0028] in, To normalize the local density values, It represents the maximum local density value of all points in the current frame's point cloud.

[0029] After normalization in both dimensions, a significance score S is constructed for each point. i : .

[0030] in, An adaptive weight adjustment mechanism is introduced: during runtime, when a point belongs to the Far Field region, the weight is increased. Weighting, while lowering the significance threshold : .

[0031] in, and T represents the dynamic baseline threshold coefficient.s far This represents the saliency threshold of a point cloud within the Far Field region. Ultimately, the threshold is used to determine whether to retain the point for subsequent pillar construction. .

[0032] Furthermore, step 3 specifically involves: reading the pillar parameter table, dividing the pillar using the raster indexing method, and dividing the pillar according to the selected area. The BEV space is divided into grids.

[0033] Each grid point corresponds to a pillar. The selected and retained point cloud is projected onto the BEV grid, and the pillar number is marked. If the number of points in a pillar is greater than or equal to the number of points in the corresponding area, then... If a point is found to be valid, it is considered a valid pillar and proceeds to the next stage of encoding; the remaining point clouds are discarded. For each valid pillar, the points... After calculating the offset between the point and the centroid within the pillar and the offset between the point and the central grid position of the pillar, the feature vectors are concatenated. The features of each point are encoded in higher dimensions through a shared MLP, and then max pooling is performed in the pillar dimension to form a fixed-length pillar representation vector.

[0034] Furthermore, step 4 specifically involves: introducing local convolutional responses into the shallow and mid-layer feature outputs of the backbone network for self-attention estimation, at each spatial location on the feature map. Calculate its significance weight And use it as a scale factor to reweight the original feature map: .

[0035] in, Represents the original BEV feature map. To enhance the final result, a top-down resolution restoration mechanism is employed, progressively upsampling high-level semantic features and then concatenating and fusing them with low-level detail features. The fusion strategy is expressed in the following form: .

[0036] in, This is the feature map at the current scale. The feature maps of the previous level are concatenated and then integrated through convolution to generate a fusion result with a unified number of channels.

[0037] Further, step 5 specifically involves: loading multiple sets of preset anchor box templates into a parameter file, each anchor containing three-dimensional dimensions (length, width, and height), center point height, and orientation angle information; obtaining the average value by fitting 3D bounding boxes from multiple frames of measured data based on the actual envelope size of the point cloud of weak aerial targets; employing a positive and negative sample partitioning mechanism based on 3D IoU to calculate the volume intersection-union ratio (IoU) between the predicted box and the ground truth target box (GT) for each anchor; and implementing a dynamic adjustment strategy based on the IoU lower limit across distance intervals, as follows: .

[0038] in, The spatial distance from the Anchor center to the radar. , Set the partition thresholds in the system; and adjust the center height of each Anchor for alignment.

[0039] Furthermore, step 6 specifically involves: for the spatial flight characteristics of low, slow, and small targets, a six-dimensional state vector is used to model their trajectory. The first three dimensions represent the target's three-dimensional spatial position in the current frame, and the last three dimensions represent the target's velocity components.

[0040] Assuming the target follows a uniform linear flight model within a short time, establish a state prediction model: .

[0041] Among them, X k-1 This is the posterior estimate of the target state vector at the previous time step. The state transition matrix is ​​defined as follows: .

[0042] in, The sampling period is determined based on the LiDAR scanning frequency.

[0043] Process noise It follows a zero-mean Gaussian distribution to absorb nonlinear disturbances or control errors outside the model; the observation model is constructed based on the lidar detection results, with the spatial coordinates of the detected target in the current frame serving as the observation input: .

[0044] in: , To observe noise, model the uncertainty of the detection results.

[0045] Each time a new frame of point cloud detection results is generated, the filter completes state prediction and update: State prediction: , .

[0046] in, Let be the prior state estimate at time k, and let represent the result predicted based on the state and motion model at time k-1. The posterior state estimate at time k-1 represents the optimal estimate that incorporates all observations from time k-1 and earlier. Let be the prior error covariance matrix at time k, representing the uncertainty of the predicted state. Let be the posterior error covariance matrix at time k-1, representing the uncertainty of the optimal estimate at the previous time step. F is the process noise covariance matrix, representing the degree of inaccuracy of the state transition model; k T The state transition matrix F k The transpose of .

[0047] Observation Update: .

[0048] .

[0049] .

[0050] in, Let Kalman gain be the value at time k. H is the observation matrix. k T Observation matrix The transpose of the matrix maps the six-dimensional state space to the three-dimensional observation space. The noise covariance matrix is ​​used to observe the measurement error and uncertainty of the lidar. Let be the posterior state estimate at time k, which is the optimal and final state output obtained after observation updates. Let be the posterior error covariance matrix at time k, representing the uncertainty of the final optimal estimate.

[0051] After completing the EKF state update in each frame, the latest estimated target position is... Converted to azimuth (Yaw) and pitch (Pitch) relative to the gimbal origin, used to drive real-time gimbal turning: Azimuth calculation: .

[0052] Pitch angle calculation: .

[0053] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0054] The present invention also discloses a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the method of the present invention.

[0055] The present invention also discloses a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method of the present invention.

[0056] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. In terms of lidar point cloud acquisition and preprocessing: This invention addresses the characteristics of small radar cross-sections and weak reflected signals of low-speed, small targets by designing a point cloud preprocessing strategy suitable for small aerial targets, thus improving the saliency of target point clouds in low signal-to-noise ratio backgrounds. It enhances the ability to retain weak echoes during the data acquisition and filtering stages, while suppressing background interference points such as ground clutter, trees, and buildings, thereby providing clearer and more stable point cloud input for subsequent identification models.

[0057] 2. Deep Learning Model Design: This invention features customized optimizations to the deep learning network structure, adapting mainstream 3D detection models based on ground targets (such as cars and pedestrians) to a structure suitable for low-altitude, slow-moving, and small flying targets. The model is specifically designed in terms of volume scale adaptation, feature extraction layers, and detection frame scale matching to better suit the characteristics of small target size, irregular shape, and sparse point clouds, thereby improving detection accuracy and robustness.

[0058] 3. Extended Kalman Filter Tracking and Prediction: Considering the volatile trajectory and unstable attitude of small, slow-moving targets, this invention introduces an extended Kalman filter algorithm to perform state estimation and trajectory prediction based on the detection results. This method combines multi-frame point cloud data from the lidar with a target dynamic model to update the target's spatial position, velocity, and acceleration in real time, effectively solving problems such as short-term occlusion, detection jitter, and temporary missed detections, thereby improving tracking continuity and prediction accuracy.

[0059] 4. This invention constructs a lidar perception and tracking system for low-altitude, slow-flying, and small targets, which has good target detection accuracy, tracking stability, and environmental adaptability, and can provide effective technical support for application scenarios such as low-altitude safety monitoring and UAV management. Attached Figure Description

[0060] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0062] This invention addresses the characteristics of low-flying, slow-moving, and small targets in point cloud data. Combining the data input requirements of the PointPillars model, it designs a point cloud preprocessing framework comprising multiple sub-modules, including high-frequency lidar acquisition, data spatial regularization, feature enhancement, background removal, and voxelization construction. This enables efficient extraction and structural representation of low-flying, slow-moving, and small targets at the data input layer. Figure 1 As shown, the technical solution of the present invention includes the following main steps: Step 1: LiDAR point cloud acquisition and distance perception preprocessing.

[0063] 1.1 Distance-Aware Point Cloud Region Modeling: The lidar equipment used in this invention is fixedly mounted on a two-degree-of-freedom gimbal with pitch and rotation capabilities, providing dynamic sensing capabilities for subsequent spatial modeling and target localization. To ensure the system's accurate interpretation of the three-dimensional structure of the point cloud, the measurement of the radar installation position and the configuration of spatial parameters must be completed in the initial stage of equipment deployment. The specific operation is as follows: First, using measuring tools such as a tape measure, accurately measure the vertical height of the radar from the installation reference surface (such as the platform ground) to the bottom of the sensor, and record it as the installation height. The unit is meters. Simultaneously, using instruments such as a level and goniometer, the pitch angle at the initial installation of the radar was recorded. Horizontal rotational deflection (Yaw) And the horizontal position offset relative to the origin of the carrier coordinate system. .

[0064] After completing the above measurements, input the data into the calibration module to construct the extrinsic parameter matrix of the radar sensor in the platform coordinate system. This matrix is ​​stored in the form of a four-dimensional homogeneous coordinate transformation and written to the system configuration file for real-time access during runtime. During point cloud acquisition, each frame of raw point cloud data is represented in local coordinates as follows: .

[0065] To transform it to an absolute coordinate system, the system reads extrinsic parameters during the data preprocessing stage. And perform coordinate transformations on each point: .

[0066] The transformed point is located in the platform's global space. This is used for subsequent height calculation and spatial selection.

[0067] To further identify potential point cloud features of low-flying targets, the vertical height information of the points relative to the ground is calculated: .

[0068] This altitude will serve as one of the input dimensions for subsequent saliency scoring and classification, used to determine whether the target is likely a low-altitude, slow, and small aircraft. Simultaneously, the spatial distance (i.e., its three-dimensional Euclidean distance to the radar center) will be calculated for each point: Thresholds set based on experience The point cloud is divided into three spatial regions: .

[0069] This information will be used throughout the subsequent saliency screening and pillar construction process, serving as an important basis for distance-aware adaptive processing.

[0070] 1.2 Adaptive Saliency Filtering Mechanism: In practical tasks, the point cloud data acquired by lidar contains a large number of non-target points, such as stray echoes from the ground, building edges, and background vegetation. Especially in long-distance areas, target points often exist only in a very small number of sparse points, making them easily misfiltered or overwhelmed. Therefore, reasonable filtering of the point cloud before voxel encoding is particularly crucial.

[0071] To this end, the present invention designs a saliency scoring mechanism that integrates "reflection characteristics" and "local geometry" in order to prioritize the retention of key points that may belong to the target object, especially to give extra retention opportunities to sparse points of small targets at a distance.

[0072] (1) Reflection intensity normalization: After each frame of point cloud acquisition is completed, the system first scans the reflection intensity values ​​of all points and extracts the maximum value of the frame. and minimum value Then the reflection intensity at each point was... Normalize: .

[0073] This process unifies the intensity range differences between different frames, enhancing the stability of the scoring mechanism.

[0074] (2) Local density estimation: in the three-dimensional neighborhood of each point (sphere radius) ) Count the number of neighboring points It employs a KD-Tree fast neighborhood search, which is computationally efficient and suitable for large-scale point clouds.

[0075] To facilitate comparisons between different points, the number of neighborhood points is normalized: .

[0076] This density value reflects the local geometric complexity of the point. High-density points usually belong to well-structured targets or their boundary regions.

[0077] (3) Significance scoring and adaptive enhancement mechanism: After normalizing the above two dimensions, a significance score is constructed for each point: .

[0078] in, In the system configuration, this parameter group can be specified by the user; for example, the default value is... .

[0079] To improve the retention of sparse targets at long distances, this invention introduces an "adaptive weight adjustment" mechanism: during runtime, when a point belongs to the "Far Field" region, the system will automatically increase the weight. Weighting (increasing emphasis on reflection intensity) while lowering the significance threshold. ,For example: .

[0080] in and By optimizing settings using measured data, the system dynamically adjusts parameters based on the distance distribution of the detected targets, ensuring that sparse points at greater distances are more likely to pass the screening. Finally, a threshold is used to determine whether to retain the point for subsequent pillar construction. .

[0081] 1.3 Distance-Partitioned Pillar Construction and Encoding Strategy: In LiDAR 3D target detection tasks, the PointPillars encoding structure effectively improves feature processing efficiency by mapping spatial point clouds to a 2D Bird's Eye View (BEV) plane. However, when facing low-flying, slow-moving, and small targets, especially when they appear in distant regions and the point cloud is sparse, the traditional uniform-size voxel partitioning strategy often encounters the following problems: In distant regions: the number of points is sparse and scattered, and there are insufficient points within voxels, making it difficult to form pillars.

[0082] In close-range areas: the pixel density is high but the pillar resolution is too low, resulting in the loss of boundary details.

[0083] Unified parameter strategy: cannot meet the feature representation needs of targets at different distances.

[0084] To address these issues, this invention proposes a distance-aware adaptive pillar construction mechanism. The core idea is to "dynamically set pillar structure parameters according to distance regions" to enhance the universality and effectiveness of pillar generation.

[0085] (1) Region classification and parameter loading: When the system starts, this module automatically reads the preset pillar parameter table in the configuration file, as shown in Table 1.

[0086] .

[0087] These parameter values ​​were optimized based on a large number of flight-measured point clouds and were set according to the point count characteristics and structural distribution features of the target at different distances.

[0088] During each frame processing, the system relies on previously calculated... The value (i.e., the spatial distance of the point) is used to assign a region label (near, middle, far) to each point, and different pillar division parameters are used accordingly.

[0089] (2) Pillar generation process: This module uses the raster index method to divide the pillar, according to the selected area. The BEV space is divided into grids; each grid point corresponds to a pillar. The selected and retained point cloud is projected onto the BEV grid, and the pillar number is marked. If the number of points in a pillar is greater than or equal to the number of points in the corresponding area, the BEV space is divided into grids. If a point cloud is found to be valid, it is considered a valid pillar and proceeds to the next stage of encoding. The remaining point clouds are either discarded or cached for global optimization.

[0090] (3) Feature encoding method: For each point in a valid pillar The system sequentially calculates the following types of local features: the offset of a point from the centroid within the pillar: .

[0091] Offset of the point from the center grid position of the pillar: .

[0092] in Let be the coordinates of the pillar's center in the BEV plane.

[0093] Complete point feature vector concatenation: .

[0094] Feature encoding: Each point feature is encoded in higher dimensions by sharing the MLP, and then max pooling is performed in the pillar dimension to form a fixed-length pillar representation vector.

[0095] Finally, all pillar encoding results are combined into a sparse tensor input, which is then fed into the backbone network for spatial perception and target detection.

[0096] Step 2: Detection network design and enhancement module mechanism based on multi-scale structure: Based on the aforementioned sparse BEV coding results, this invention constructs a multi-scale sensing backbone network and embeds spatial enhancement mechanism and channel attention mechanism at key levels to significantly improve the robustness and sensitivity of detection.

[0097] 2.1 Multi-scale pyramid structure backbone network: In point cloud detection scenarios, the volume scale of low-speed small targets changes significantly at different distances: the voxels are dense in the near-distance region, the target features are clear, and it is easy to capture the edges and geometry; while the number of points in the far-distance region is sparse, the target outline is blurred, the feature expression is limited, and it is easily covered by the background.

[0098] To adapt to this distribution characteristic, the BEV backbone network constructed in this invention adopts a pyramid structure to expand the feature extraction pathway, which is constructed sequentially by multiple downsampling layers. Feature layers of different scales have different receptive fields.

[0099] The first layer maintains the original BEV size, focusing on preserving the spatial details of small targets; the second layer expands the receptive field by 2x downsampling to capture mesoscale structures; the third layer further downsamples to obtain a wide range of background and contextual information.

[0100] Each layer is constructed by stacking convolutional layers, with the number of channels gradually increasing from bottom to top according to the network depth to enhance semantic expressiveness. The feature maps output by each layer will be cascaded or upsampled and converged in the subsequent fusion module to achieve information backflow and take into account multi-scale structural representation.

[0101] 2.2 Spatial Attention and Edge-Guided Enhancement Mechanism: When processing sparse point cloud features, the structural information of the target itself is weak. Therefore, a spatial attention mechanism needs to be introduced to highlight potential target regions and suppress background interference. This invention designs a spatial enhancement submodule that applies spatial weight adjustment to the BEV feature map, thereby dynamically enhancing structurally sensitive regions.

[0102] In the shallow and mid-layer feature outputs of the backbone network, local convolutional responses are introduced for self-attention estimation at each spatial location on the feature map. Calculate its significance weight And use it as a scale factor to reweight the original feature map: .

[0103] in, Represents the original BEV feature map. To enhance the results, this mechanism is particularly effective in responding to sparse features in distant regions, thus highlighting the contours of weak targets.

[0104] Furthermore, considering that small targets are mostly distributed at the edge of the background, a boundary response map is constructed in the BEV plane using local gradient information to guide features to focus on edge details, thereby improving the ability to preserve small structures.

[0105] 2.3 Feature Fusion and Resolution Reversion Strategy: To address the issue of semantic inconsistency between multi-scale features, a top-down resolution restoration mechanism is adopted, which progressively upsamples high-level semantic features and then fused them with low-level detail features.

[0106] This fusion strategy can be expressed in the following form: .

[0107] in, This is the feature map at the current scale. The previous level (low-resolution) feature maps are concatenated and then convolved to integrate the feature dimensions, generating a fusion result with a uniform number of channels. Through progressive upward fusion, the system obtains multi-scale feature representations that combine detail and context before the final detection head input.

[0108] Step 3: Anchor Design and Matching Strategy Based on Weak Aerial Targets: In point cloud target detection represented by BEV, the anchor mechanism is the foundation for establishing the association between predicted and real targets. Its matching quality directly affects the regression accuracy and positive / negative sample balance of the subsequent detection network. Traditional anchor designs are mostly based on fixed-size vehicles / pedestrian tasks that are biased towards large targets. For weak aerial targets such as low-altitude, slow-moving, and small aircraft, there are problems such as low matching rate, insufficient positive samples, and large localization errors.

[0109] This invention addresses the characteristics of micro-targets flying in the airspace by redesigning the anchor frame size system and introducing a distance partitioning adjustment strategy and an IoU dynamic lower limit adjustment mechanism into the matching mechanism, thereby comprehensively improving the anchor frame matching rate of weak targets and the quality of training samples.

[0110] 3.1 Anchor Geometric Parameter Design and Pre-configuration: To cover the shape features of flying targets of various scales (such as micro UAVs, remote-controlled gliders, etc.), this invention loads multiple sets of preset anchor frame templates through parameter files during the detection network initialization stage. Each anchor includes three-dimensional dimensions (length, width, and height), center point height, and orientation angle information. The anchor frame template styles are shown in Table 2.

[0111] .

[0112] During the anchor frame design process, the average value was obtained by fitting 3D bounding boxes of multiple frames of measured data, based on the actual envelope size of the point cloud of weak aerial targets. Simultaneously, multiple orientation angles were configured for each type of anchor to improve the matching flexibility for targets flying in multiple directions.

[0113] 3.2 Anchor Matching Mechanism and IoU Lower Limit Adjustment: To improve anchor matching efficiency, the system adopts a positive and negative sample partitioning mechanism based on 3DIoU. The volume intersection-union ratio (IoU) is calculated for each predicted anchor bounding box and the ground truth bounding box (GT). .

[0114] Traditional matching mechanisms often set fixed IoU thresholds (e.g., 0.6 for positive samples and 0.3 for negative samples). However, this performs poorly in scenarios with sparse points and distant targets. This is because the weak target point cloud envelope is blurred, and the anchor box and ground plane center are prone to shift, leading to a large number of valid anchors being misclassified as negative samples. Therefore, this invention proposes a dynamic adjustment strategy for the lower limit of IoU across different distance intervals, as follows: .

[0115] in, The spatial distance from the Anchor center to the radar. , This is the partition threshold set in the system. This mechanism makes it easier for distant sparse targets to be matched with suitable anchors, increases the number of positive samples during training, and reduces detection blind spots.

[0116] 3.3 Aerial Target Center Alignment Strategy: Low-altitude, slow-moving, small targets in flight generally have a higher center altitude than static ground targets. To avoid detection bias caused by mismatch between anchor altitude and target center, the system performs alignment of each anchor during the matching phase. Adjust center alignment: For GT center height Choose with The closest Anchor is preferred.

[0117] If the height difference between the matched Anchor and the center of the ground plane exceeds the threshold (e.g., ±0.3m), the structure will not be matched and will be considered as mismatched.

[0118] This height alignment strategy can significantly improve the matching stability of anchors in the z-axis direction, making an important contribution to the accuracy of 3D detection.

[0119] 3.4 Anchor box balancing and training sample construction: To prevent class imbalance, the present invention adopts the following strategy when constructing positive and negative samples: For each GT box, at least one anchor is guaranteed to match successfully; for anchors with an overlap rate close to the threshold, those closer to the center are given priority; the ratio of positive to negative samples in each frame is controlled within a reasonable range (e.g., 1:3) to balance the loss calculation.

[0120] Once the network is built, all positive samples will participate in the calculation of regression loss and classification loss, and the network parameters will be optimized through error backpropagation.

[0121] Step 4: Trajectory association of low-speed small targets based on lidar point clouds and extended Kalman filter tracking method: The system uses an extended Kalman filter (EKF) to perform time-series prediction and observation update of the target state. At the same time, the estimated three-dimensional spatial position of the target is used to drive the two-degree-of-freedom gimbal to turn, forming a "detection-prediction-turning" closed loop to ensure continuous tracking and stable pointing of UAV targets.

[0122] 4.1 State Modeling and Target Representation: For the space flight characteristics of low, slow, and small targets, the system uses a six-dimensional state vector to model their trajectory. .

[0123] The first three dimensions represent the target's three-dimensional spatial position in the current frame, while the latter three dimensions represent the target's velocity components. This state vector accurately expresses the target's dynamic behavior in the radar coordinate system and supports the prediction of target motion trends.

[0124] During the system initialization phase, when the detection module stably detects a potential target for multiple consecutive frames, the system will establish the EKF trajectory object corresponding to the target and initialize its state vector and covariance matrix.

[0125] 4.2 State Transition and Observation Model: This invention assumes that the target follows a uniform linear flight model within a short time, and establishes a state prediction model: .

[0126] in, The state transition matrix is ​​defined as follows: .

[0127] Process noise It follows a zero-mean Gaussian distribution to absorb nonlinear disturbances or control errors outside the model.

[0128] Meanwhile, the observation model is constructed based on the lidar detection results, with the spatial coordinates of the detected target in the current frame serving as the observation input: .

[0129] in: .

[0130] To observe noise, model the uncertainty of the detection results.

[0131] 4.3 Filter operation process and trajectory prediction update: During system operation, whenever a new frame of point cloud detection results is generated, the filter completes state prediction and update according to the following steps.

[0132] (1) State prediction: , .

[0133] (2) Observation update: .

[0134] .

[0135] .

[0136] Through the above prediction and update process, the system can continuously estimate the target's three-dimensional spatial position and velocity, alleviating the trajectory instability caused by short-term detection of missing frames, false detections, or occlusions.

[0137] 4.4 Gimbal Control and Target Tracking Driven Mechanism: After completing the EKF state update in each frame, the system will update the latest estimated target position. These are converted into azimuth and pitch angles relative to the gimbal origin, used to drive the gimbal to turn in real time.

[0138] Azimuth calculation: .

[0139] Pitch angle calculation: .

[0140] The system converts the above angle values ​​into PWM signals or motor control commands, which are then transmitted to the control board (such as an STM32 or Jetson platform). The control board drives two servos to control the horizontal and vertical rotation of the gimbal, so that the lidar faces the target position and maintains continuous tracking.

[0141] When the target moves near the edge of the field of view or is briefly obstructed, the system can still turn the gimbal according to the EKF prediction value to maintain the tracking direction and improve anti-interference and responsiveness.

[0142] Example

[0143] First, the raw data is received by rotating and scanning at a frequency of 10Hz using the LeiShen intelligent LiDAR. Each frame of data contains tens of thousands of three-dimensional point clouds, and the information of each point is a data vector. According to the calibrated extrinsic matrix Transform the local coordinates of each point to the global coordinate system to obtain the point's position in the global space. Calculate the spatial distance of each point. Scan the entire frame point cloud, extract the maximum / minimum reflection intensity of that frame, and normalize it to... The KD-Tree algorithm is used to count the number of neighborhood points, calculate the density of local points, and calculate a significance score. A threshold is used to determine whether to retain the point in subsequent pillar construction. The selected points are projected onto the BEV plane, and an adaptive mechanism is used to detect the region where the target point is located. Corresponding mesh parameters are loaded, and the offsets between the target point and the centroids of all points within the pillar, as well as the offset of the mesh center, are calculated. The original coordinates, reflection intensity, and the two offsets are concatenated into a 6-dimensional feature vector. The feature vectors of all points within the Pillar are input into a shared MLP, each point is upscaled to 128 dimensions, and Max Pooling is performed to extract the maximum value across the 128 channels. The processed Pillar features are then scattered back into a BEV grid based on their spatial location, forming an H×W×128 pseudo-image. This pseudo-image is input into a pyramid CNN, which outputs a multi-scale feature map. High-level features are upsampled and concatenated with low-level features to obtain a final feature map that integrates details and semantics. Preset anchors are used on this final feature map for classification and regression, determining whether each anchor box contains "background" or a "small, low-speed target." Finally, the first three dimensions of the detected small target information are used. The input is fed into the Kalman filter; based on the observed covariance matrix... and observation noise The Kalman gain is calculated, and the prediction is corrected through state updates to output the optimal estimate of the trajectory prediction. Finally, the predicted target position is converted into azimuth and pitch angles relative to the origin of the gimbal, and the gimbal is driven by a motor to achieve real-time tracking of low, slow, and small targets.

Claims

1. A laser radar-based low, slow and small target detection and trajectory prediction tracking method, characterized in that, Comprise the following steps: Step 1, collect laser radar point cloud data, and carry out spatial modeling and coordinate transformation preprocessing on the laser radar point cloud data; Step 2, significant screening is carried out on the preprocessed data, and the screened point cloud data is obtained; Step 3, based on distance partition driving, the screened point cloud data is subjected to Pillar construction and coding, and the pillar expression vector is obtained; Step 4, a deep learning detection network is constructed, and the pillar expression vector is taken as the input, and the fused feature is output; Step 5, based on the anchor design and matching strategy, the size of the aerial weak target in the fused feature is adapted; Step 6, based on the extended Kalman filter EKF, the target state is time series predicted and observed updated, and the low slow small target detection and trajectory prediction tracking are completed.

2. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 1 specifically: measure the vertical height from the installation datum plane to the bottom of the sensor , record the initial installation pitch angle of the radar , horizontal rotation deflection , and the horizontal position offset of the radar relative to the origin of the carrier coordinate system , input the data into the calibration module, and construct the external parameter matrix of the radar sensor in the platform coordinate system ; during the point cloud collection process, each frame of original point cloud data is expressed in the form of local coordinates: , wherein, is the original data of each frame of point cloud, , , is the horizontal coordinate, vertical coordinate and height information of each point cloud in the local coordinate system, is the reflection intensity of each point cloud, and N is the number of point clouds in the original point cloud data of each frame. Reading the external parameters and performing coordinate transformation for each point: , transformed to get the position of the point in the global space of the platform ; The vertical height information of the calculation point relative to the ground is calculated: , The spatial distance of each point is calculated: , Thresholds set empirically , the point cloud is divided into three spatial regions: , wherein, is a spatial region of the point cloud, are empirically set threshold limits for the near, mid, and far distances of the point cloud.

3. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 2 is specifically: After each frame of point cloud acquisition is completed, the reflection intensity values ​​of all points are scanned, and the maximum value of that frame is extracted. and minimum value The reflection intensity at each point Normalize: , KD-Tree neighborhood search is used to count the number of neighborhood points within a three-dimensional neighborhood of each point : , wherein, is the current point for which saliency is being computed, is the point cloud is any other point in the point cloud; The number of neighborhood points is normalized: , wherein, is the normalized local density value, is the maximum value of the local density values of all points in the current frame point cloud; After normalization in both dimensions, a saliency score S is constructed for each point i : , wherein, ; introduce adaptive weight adjustment mechanism: at runtime, when a point belongs to Far Field region, up-regulate weight, while lowering saliency threshold ; , wherein, and denotes a dynamic base threshold coefficient, T s far denotes a significance threshold that the point cloud is in the Far Field region; Finally, according to the threshold value, it is judged whether the point is retained to enter the subsequent pillar construction: 。 4. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 3 is specifically: reading the pillar parameter table, using the grid index method to divide the pillars, and dividing the selected area according to the pillar number Grid division of BEV space; Each grid point corresponds to a pillar column, and the screened and retained point cloud is projected to the BEV grid, and the corresponding pillar number is marked. If the number of points in a certain pillar is greater than or equal to the corresponding number of points in the region , it is regarded as a legal pillar, and the next stage of coding is entered, and the remaining point cloud is removed. For each point within a legal pillar , the offset of the point to the centroid of the pillar and the offset of the point to the center grid location of the pillar are calculated, and then the feature vectors are spliced, each point feature is dimensionally coded through a shared MLP, and then maximum pooling is performed in the pillar dimension to form a fixed-length pillar expression vector.

5. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 4 is specifically: in the shallow and middle layer feature output of the backbone network, local convolution response is introduced for self-attention estimation, and each spatial position on the feature map The significance weight thereof is calculated and is used as a scale factor to reweight the original feature map: , wherein, denotes the original BEV feature map, is the enhanced result; An up-down resolution recovery mechanism is adopted to gradually up-sample the high-level semantic features and splice and fuse with the low-level detail features; The fusion strategy is expressed in the following form: , wherein, is the current scale feature map, is the upper level feature map, and after splicing, the feature dimension is integrated by convolution to generate a fusion result with a unified channel number.

6. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 5 is specifically: a plurality of preset anchor box templates are loaded through a parameter file, each Anchor contains three-dimensional size, center point height and direction angle information; According to the actual envelope size of the aerial weak target, the mean value is obtained by fitting the 3D bounding box of the measured data; A 3D IoU-based positive and negative sample division mechanism is adopted to calculate the volume intersection ratio between each Anchor prediction box and the real target box GT; The IoU lower limit based on distance interval is dynamically adjusted, which is specifically as follows: , wherein, is the spatial distance from the Anchor center to the radar, , is the partition threshold set in the system; and the center height of each Anchor is adjusted in alignment.

7. The low, slow and small target detection and trajectory prediction tracking method based on laser radar according to claim 1, characterized in that, Step 6 is specifically: for the spatial flight characteristics of the low slow small target, a six-dimensional state vector is used to model its trajectory, the first three dimensions represent the three-dimensional spatial position of the target in the current frame, and the last three dimensions represent the velocity component; Assuming that the target satisfies the uniform straight flight model in a short time, a state prediction model is established: , where X k-1 is the posterior estimate of the state vector at the previous time step, is the state transition matrix, defined as: , wherein, is a sampling period determined from the laser radar scan frequency; Process noise The observation model is constructed based on the detection results of the lidar, and the spatial coordinates of the detected targets in the current frame are taken as the observation input: , Wherein: , To observe noise, the uncertainty of the detection result is modeled; Every time a new frame of point cloud detection result is generated, the filter completes state prediction and update: State prediction: , , wherein, is the prior state estimate at time k, representing the result predicted from the state at time k-1 and the motion model, is the posterior state estimate at time k-1, representing the optimal estimate incorporating all observations up to time k-1, is the prior error covariance matrix at time k, representing the uncertainty of the predicted state, is the posterior error covariance matrix at time k-1, representing the uncertainty of the optimal estimate at the previous time, is the process noise covariance matrix, representing the inaccuracy of the state transition model; F k T is the state transition matrix F k is the transpose of the state transition matrix F Observation update: , , , wherein, is the Kalman gain at time k, is the observation matrix, H k T is the observation matrix is the transpose of H, mapping the six-dimensional state space to the three-dimensional observation space, is the observation noise covariance matrix, representing the measurement error and uncertainty of the LiDAR, is the posterior state estimate at time k, which is the optimal, final state output after the observation update, is the posterior error covariance matrix at time k, representing the uncertainty of the final optimal estimate; After each frame completes the EKF state update, the latest estimated target position Convert to azimuth Yaw and pitch Pitch relative to the gimbal origin, used to drive the gimbal real-time steering: Azimuth angle calculation: , Pitch angle calculation: 。 8. A computer apparatus comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, causes the processor to perform the method of any one of claims 1 to 7. The processor executes the computer program to realize the steps of the method of claim 1.

9. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to realize the steps of the method of claim 1.

10. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to realize the steps of the method of claim 1.

Citation Information

Patent Citations

  • Low and slow small target tracking method based on polar coordinate system

    CN109100714A

  • Low-speed small target tracking device and method adapting to urban complex background

    CN112381856A

  • Radar target processing method and device, storage medium and vehicle

    CN118131230A

  • Multi-vehicle target tracking method based on roadside laser radar point cloud data

    CN118864531A

  • Unmanned aerial vehicle low-altitude target positioning and identification system

    CN120411824A

Cited By

  • Horizontal panoramic unmanned aerial vehicle detection method and system based on rotation event camera

    CN121353960A

  • A method and system for detecting horizontal panoramic UAVs based on a rotating event camera

    CN121353960B

  • Track measurement method and device based on line structured light and storage medium

    CN121363916A

  • Air target automatic detection and following lock positioning method

    CN121721653A