A low, slow and small target detection and recognition method based on radar and variable focus camera

By combining the calibration of lidar and a zoom camera with multimodal fusion technology, the problem of identifying small, slow, and low-speed targets in complex environments using traditional detection methods has been solved, achieving high-precision and rapid target identification and strike capabilities.

CN121069415BActive Publication Date: 2026-02-10CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511606094.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-10
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Traditional detection methods struggle to efficiently and reliably identify small, slow-moving targets in complex environments, especially in complex terrain and adverse weather conditions, where their detection capabilities are limited and their dynamic response is insufficient.

Method used

By employing joint calibration and coordinated control of LiDAR and a variable-focus camera, combined with multimodal fusion technology, and through deformable BEV projection and cross-modal attention fusion, real-time detection and recognition of low, slow, and small targets can be achieved.

Benefits of technology

It significantly improves the identification accuracy and response speed of low, slow and small targets in complex environments, ensuring that the system can work stably under all-weather conditions and has efficient target threat assessment and strike capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069415B_ABST
    Figure CN121069415B_ABST
Patent Text Reader

Abstract

The application discloses a low, slow and small target detection and identification method based on a radar and a variable-focus camera, and comprises the following steps: step 1, joint calibration and parameter prestorage of a laser radar and a variable-focus camera; step 2, laser radar target detection and variable-focus camera dynamic focusing control; step 3, identification of a low, slow and small target based on multi-modal fusion; and step 4, target striking control and execution. The application breaks through the physical limitation of a single sensor, realizes the perception ability leap of 1+1>2 through intelligent fusion, establishes an adaptive processing framework, enables the system to dynamically cope with various complex scenes, and provides reliable real-time performance, and meets the strict requirements of key security applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting and recognizing small, slow targets, and more particularly to a method for detecting and recognizing small, slow targets based on radar and a zoom camera. Background Technology

[0002] With the rapid development and widespread application of unmanned aerial vehicle (UAV) technology, the demand for the detection and identification of "low, slow, and small" targets has increased significantly. In complex battlefield environments, traditional detection methods can no longer meet the requirements for efficient and reliable detection of moving targets. The widespread use of "low, slow, and small" weapons such as UAVs and loitering munitions has brought new challenges to the modern battlefield. Due to their small size and small radar cross-section, "low, slow, and small" weapons are difficult to detect effectively by traditional radar systems. Their stealth and mobility make these targets a significant hidden danger, thus requiring rapid and accurate identification and interception of these targets.

[0003] There are four common methods for detecting "low, slow and small" flying objects: radar detection, radio detection, acoustic detection and photoelectric detection.

[0004] Radar detection utilizes the reflection of radio waves to detect and locate targets. However, radar is susceptible to interference from ground reflections and clutter, especially in complex terrain; it is particularly difficult to detect low-flying, slow-moving, and small objects, especially those with low radar cross-sections (RCS); it is also costly, requires complex equipment, and the false alarm rate of radar systems increases by more than 40% in complex terrain.

[0005] Radio detection utilizes the interaction between a target and radio signals, analyzing changes in these signals to detect the target. However, it has a limited detection range, is suitable for close-range targets, is susceptible to interference from other radio equipment leading to signal misinterpretation, and performs poorly in detecting high-speed targets.

[0006] Acoustic detection utilizes the sound signals emitted by the target or detects targets by actively emitting sound waves and receiving reflected waves. However, it has a limited detection range, making it suitable for short-range monitoring; it is greatly affected by environmental noise and easily interfered with by natural factors such as wind and rain; and it is less effective at detecting low-noise or silent flying objects.

[0007] Photoelectric detection utilizes optical sensors to capture images or infrared radiation of targets, and then uses image processing techniques and algorithms to analyze and identify them. However, photoelectric detection is greatly affected by weather and lighting conditions, such as fog, rain, snow, and strong light, which can affect detection results; its detection range is relatively limited, requiring line-of-sight detection; and processing image data requires high computing resources, with optical sensor performance decreasing by up to 70% in foggy or hazy weather. Summary of the Invention

[0008] Purpose of the invention: To address the shortcomings of existing technologies, this invention provides a method for detecting and identifying low-speed, small targets based on radar and a variable-focus camera. This method solves the problems of poor environmental adaptability, limited detection capability, and insufficient dynamic response in detecting low-speed, small flying targets, improves the real-time detection and identification performance of low-speed, small targets in complex environments, and meets the urgent needs of modern urban air defense.

[0009] Summary of the invention: A method for detecting and recognizing small, slow targets based on radar and a zoom camera, comprising the following steps: Step 1: Joint calibration and parameter pre-storage of LiDAR and zoom camera; The LiDAR and zoom camera are mounted on the same two-degree-of-freedom gimbal, ensuring no relative motion between them, and the LiDAR and zoom camera are jointly calibrated to obtain their intrinsic and extrinsic parameter matrices.

[0010] Step 2: LiDAR target detection and dynamic focusing control of the zoom camera; extract potential low-speed small targets from the original radar point cloud data, calculate the target's azimuth and distance, adjust the gimbal pointing according to the target's azimuth to ensure that the target is always in the center of the zoom camera's field of view, and automatically adjust the camera's focal length according to the target's distance to ensure that the target's resolution in the image meets the requirements.

[0011] Step 3: Identify low-speed, small targets based on multimodal fusion; in the image feature extraction module, embed conditional parameterized convolutions into each residual block; in the radar point cloud feature extraction module, extract the height, density, and reflection intensity features of the target and concatenate them with the learned features; perform cross-modal attention fusion calculation through deformable BEV projection to optimize small targets in real time.

[0012] Step 4: Target strike control and execution; Based on the target identification module, the category confidence level and distance to the gimbal are obtained, the target threat is assessed, and then laser strike is carried out. The strike effect is evaluated. If the target is not destroyed on the first strike, the system automatically triggers parameter adjustment, and the gimbal aiming point is dynamically compensated according to the offset until the target is successfully shot down and the attack stops.

[0013] Furthermore, in step 1, a coordinate system is established, with the world coordinate system as W, the gimbal coordinate system as G, the lidar coordinate system as L, and the zoom camera coordinate system as C. The joint calibration includes the intrinsic parameter calibration of the zoom camera, the relative extrinsic parameter calibration of the lidar and the zoom camera, and the gimbal motion compensation.

[0014] Ideally, the intrinsic parameters of a zoom camera should be calibrated to reduce focal length divergence within the focal length range. , Select N calibration points within [the specified area]. , ... Fixed focal length to Multiple sets of chessboard images with different poses were captured, and the intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: Where α is the weighting coefficient, which calculates the contribution ratio of the target focal length in the known focal length range.

[0015] The relative extrinsic parameter calibration between the LiDAR and the zoom camera is performed by placing the AprilTag calibration board in a static scene, ensuring it appears in the field of view of both the LiDAR and the zoom camera, and measuring the coordinates of the corner points of the calibration board in the world coordinate system W. That is, to calculate the transformation from the world coordinate system to the calibration plate. ; Detect the AprilTag to obtain the pose of the calibration board in the zoom camera coordinate system. Extract the point cloud from the calibration board, fit the plane equation, and calculate the pose in the lidar coordinate system. The pose of the lidar to the zoom camera was calculated using the world coordinates of the calibration board. .

[0016] Gimbal motion compensation is used to calibrate the installation positions of the radar and camera in the gimbal coordinate system G. and The pitch angle is obtained through the gimbal encoder. and azimuth Calculate the rotation transformation of the gimbal coordinate system G relative to the world coordinate system W: Real-time positions of the lidar and zoom camera in the world coordinate system: ; .

[0017] radar point coordinates Transform to world coordinate system Then to the camera coordinate system Finally, the image pixel coordinates are projected onto the camera intrinsic parameters to obtain... : ; ; .

[0018] Ideally, the calibration parameters should be pre-stored for each focal length. ,storage When detecting a target, based on the current focal length f, intrinsic parameters are called or interpolated, and calculations are performed in real time as the gimbal rotates. and Ensure coordinate alignment during movement.

[0019] Furthermore, in step 2, the detection and control of low-speed, small targets includes the following steps: Step 21: LiDAR point cloud acquisition and preprocessing; raw LiDAR point cloud Each point contains coordinates (x, y, z) and reflection intensity. The point cloud is filtered, and statistical outlier filtering is used to remove noise. If a certain point... Average distance to neighboring points Greater than a threshold calculated based on the distribution of neighboring points , for If the standard deviation of the k nearest neighbors is given, then this point is identified as a noise point and removed. The noise-removed point cloud is subjected to target clustering. Euclidean distance is used to merge neighboring points into candidate target clusters. A threshold is set to fit the size of small, slow, and large targets. Clusters with sufficient volume are retained, while objects that are too large or too small are excluded.

[0020] Step 22: Gimbal-based collaborative control; Based on the target's 3D point cloud position detected by the lidar, control the azimuth and pitch angles of the two-degree-of-freedom gimbal in real time to ensure that the target is always located in the center of the camera's field of view, providing stable target imaging for subsequent image recognition.

[0021] Step 23: Dynamic focusing of the zoom camera; target distance based on radar point cloud computing. In addition, the ideal focal length is calculated to meet imaging requirements, ensuring the target's resolution in the image. Dynamic corrections are then performed, and relevant focal length calculations are conducted. Given that the drone's altitude is approximately 0.5m, the required focal length is calculated. Where h is the sensor height and H is the actual target height, which in this case is the drone's height. The radar measures the target distance; the parameters corresponding to the current focal length f are retrieved from the pre-stored calibration lookup table: the intrinsic parameter matrix K(f); if there is no pre-stored point for f, linear interpolation is used for calculation.

[0022] Step 24: Real-time optimization; Point cloud downsampling, using voxel filtering to reduce point cloud density and reduce computation; Traditional serial processes have high cumulative latency, so gimbal control and camera focusing are designed to be executed in parallel, target position information, distance, and timestamp are transmitted through shared memory, and a circular buffer is used to avoid read and write conflicts; Through hardware acceleration, the point cloud algorithm is deployed on FPGA, making the latency <10ms.

[0023] In step 3, the recognition of low, slow, and small targets includes the following steps: Step 31: In the image feature extraction module, a ResNet-50 network is used, and conditional parameterized convolutions are embedded in each residual block; an FPN network is used, in which focal length-aware feature options are introduced, and the fusion weights of features at each scale are input. When using telephoto lenses, the P5 layer, i.e., high semantic features, is strengthened, and when using wide-angle lenses, the P2 layer, i.e., detail features, is emphasized.

[0024] Step 32: In the radar point cloud feature extraction module, the point cloud is first voxelized and converted into a voxel mesh. VoxelNet is used to generate a Bev feature map to extract the target's height, density, and reflection intensity features, which are then concatenated with the learned features.

[0025] Step 33: Perform deformable BEV projection.

[0026] Step 34: Let the radar BEV characteristic be... The BEV features of the image are Cross-modal attention fusion computation is performed using a learnable weight matrix.

[0027] Step 35: Optimize small targets in real time by improving the detection head.

[0028] Furthermore, in step 31, the parameterized convolution is represented as: ;in, The focal length-related weights are generated by a lightweight MLP: ; The pre-defined expert convolutional kernels are used; the fusion weight input FPN network employs a dynamic feature fusion algorithm. After the FPN features are constructed, dynamic weights are injected during multi-scale feature fusion, and weights are applied using element-wise multiplication. When the system is in telephoto mode (f>100mm), the P5 layer is strengthened in the following way: weight allocation, MLP output... The weight is increased by 0.6-0.8; when the system is in a wide-angle state (f<50mm), the P2 layer is strengthened in the following way: MLP output The weight increases by 0.5-0.7.

[0029] Furthermore, in step 33, the deformable BEV projection includes: for each BEV grid point (x,y), predicting its sampling offset in the image based on the scale factor S. :

[0030] s; where z is the depth provided by the radar, and f is the current focal length; dynamic extrinsic parameter compensation is performed, and calibration parameters are integrated in the view transformation. Corrected projected coordinates: ;in To correct the initial depth value of the foreground point in the camera coordinate system.

[0031] Furthermore, in step 34, the cross-modal attention fusion calculation includes: using a learnable weight matrix Generate Q, K, and V matrices, where Q = K= V= ,F L For lidar point cloud features, F I For image features, d k The dimension of the key vector K: B(f) is generated by a lightweight MLP based on the current focal length f and is used to adjust the mode weights. Learnable scaling parameters: At long distances (telephoto), increase the weight of image details; at close distances (wide-angle), increase the weight of point cloud features.

[0032] Ideally, in step 35, the dynamic dimensions of the anchor are designed as follows: for a focal length range of 20–50 mm, the anchor size is 1.0 m × 1.0 m; for a focal length range of 50–100 mm, the anchor size is 0.5 m × 0.5 m; and for a focal length range of 100–200 mm, the anchor size is 0.3 m × 0.3 m.

[0033] Point cloud voxelization and BEV generation are deployed on FPGA using a parallel architecture, processing point cloud data once per clock cycle, resulting in a latency of <5ms. The expert convolutional kernels of CondConv are preloaded into the NPU cache, requiring only updates to the weight coefficients π(f) when the focal length changes, with a communication volume of <1KB. The model is lightweighted, channels are pruned, and L1 regularization is added to the convolutional layers during training. During the evaluation phase, channels with weights below the threshold are removed for real-time optimization.

[0034] Ideally, in step 4, achieving a successful target strike includes the following steps: S41: Establish a strike decision module; based on the category confidence level and distance to the gimbal obtained from the target identification module, perform a target threat assessment: ;in, For category confidence, 1 represents the weighting coefficient for the class confidence score. 1 represents the weighting coefficient of the distance penalty term. This represents the weighting coefficient for the speed threat term.

[0035] = D is the distance parameter.

[0036] = Speed ​​threat standardization.

[0037] A threshold judgment is performed. If the threat index S is greater than 0.7, the following strike steps are taken.

[0038] S42: Laser emission control; precise gimbal aiming, adjusting pitch / azimuth angles based on predicted position, activating the laser emitter, gimbal control algorithm, motion prediction results: .

[0039] = Communication delay + gimbal response time; a is the acceleration obtained by the Kalman filter, x, y, z are the three-dimensional spatial coordinates at the current moment, and v is the velocity at the current moment.

[0040] If the laser fails to destroy the target on its first strike, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated based on the offset.

[0041] S43: Strike effect assessment; UAV-side detection, including laser receiver and visual verification device; Laser receiver: photodiode array is set to 16×160; Trigger threshold: >500W / The duration is 100ms; Visual verification device: A smoke alarm device is installed on the drone. If the hit is successful, the smoke system is triggered.

[0042] S44: Real-time performance assurance; FPGA-accelerated decision-making, parallel scoring calculation, SIMD architecture, threat scoring completed in a single cycle, reducing computational latency, real-time control loop, gimbal calibration using dual-loop control, communication optimization, low-latency data transmission, protocol stack optimization, data compression, and differential encoding for strike commands.

[0043] Beneficial Effects: Compared with existing technologies, the beneficial effects of this invention are: 1. Multimodal dynamic fusion significantly improves recognition accuracy. Traditional target recognition methods often use a single sensor, making it difficult to balance long-range detection accuracy and short-range recognition capability. This invention innovatively combines the precise three-dimensional ranging of LiDAR with the multi-scale visual features of a zoom camera, achieving optimal fusion of cross-modal features through a focal length-adaptive dynamic convolutional network and deformable BEV projection technology.

[0044] 2. Intelligent zoom-coordinated processing ensures rapid response. Addressing the dynamic characteristics of low-altitude, slow-moving, and small targets, this invention designs a distance-based adaptive zoom control algorithm and a hardware-accelerated processing pipeline. Point cloud voxelization (<5ms) and NPU-accelerated dynamic convolution calculation (0.5ms) are implemented using FPGA, enabling the entire process latency from target detection to strike decision to be controlled within 30ms.

[0045] 3. Robust in complex environments, stable operation in all weather conditions. The system integrates multispectral sensor data (visible light / infrared / LiDAR) and, through a deep learning-assisted clutter suppression algorithm, maintains a high detection rate even in harsh environments such as fog (visibility <100m), strong light (100,000 lux), rain, and snow. A unique focal length compensation mechanism ensures continuous and stable tracking during zooming, solving the performance degradation problem caused by environmental changes in traditional systems.

[0046] 4. Modular design for flexible and convenient deployment. Utilizing standardized hardware interfaces and the ROS communication framework, it supports plug-and-play compatibility with various LiDAR models and zoom cameras. Innovative dynamic calibration technology allows for pre-calibration parameters to cover the entire focal length range, shortening deployment time. The system can be flexibly integrated into various platforms such as vehicle-mounted and fixed stations to meet diverse operational needs.

[0047] 5. It has a high cost-effectiveness ratio and the potential for large-scale application. Through lightweight model design (INT8 quantization + channel pruning) and heterogeneous computing architecture optimization, it can achieve 4K@30fps real-time processing on edge devices such as Jetson Orin, reducing hardware costs.

[0048] In summary, this invention effectively solves the core problems of "unclear visibility, inaccurate tracking, and inability to hit" in the detection of low, slow, and small targets through an innovative "perception-fusion-decision-strike" end-to-end technical solution. Attached Figure Description

[0049] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are illustrative and should not be construed as limiting the invention in any way.

[0050] Figure 1 A schematic diagram for improving the BEV network architecture.

[0051] Figure 2 The flowchart for the closed-loop logic control. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] The present invention will be further illustrated below with reference to specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Modifications to the present invention in various equivalent forms all fall within the scope defined by the appended claims.

[0054] This invention provides a method for detecting and recognizing small, slow targets based on radar and a zoom camera, including the following steps: Step 1: Joint calibration and parameter pre-storage of lidar and zoom camera.

[0055] The lidar and the zoom camera are mounted on the same two-degree-of-freedom gimbal, with no relative motion between them. Joint calibration of the two sensors yields the intrinsic and extrinsic parameter matrices for each sensor.

[0056] Coordinate system conventions: W is the world coordinate system; G is the gimbal coordinate system; L is the lidar coordinate system; C is the camera coordinate system.

[0057] 1.1 Camera intrinsic parameter calibration.

[0058] Focal distance divergence, within the focal length range [ , Select N calibration points within a range of 20mm to 200mm (e.g., ...). =20mm, =30mm,……, =200mm). Fixed focal length to Multiple sets of chessboard images with different poses were captured. The intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( ), where α is the weighting coefficient, which calculates the contribution ratio of the target focal length in the known focal length range.

[0059] Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: .

[0060] 1.2 Radar and camera relative extrinsic parameter calibration.

[0061] Place the AprilTag calibration board in a static scene, ensuring it appears in the field of view of both the radar and the camera. Accurately measure the coordinates of the corner points of the calibration board in the world coordinate system W. The pose of the calibration board in the camera coordinate system is obtained by detecting the AprilTag. Extract the point cloud from the calibration board, fit the plane equation, and calculate the pose in the radar coordinate system. World coordinates via calibration plate Calculate the camera pose in the world coordinate system: .

[0062] 1.3 Gimbal Motion Compensation.

[0063] Calibrate the installation positions of the radar and camera in the gimbal coordinate system G. and (Measurement requires the gimbal to be in the zero position). The pitch angle is obtained via the gimbal encoder. and azimuth Calculate the rotation transformation of the gimbal coordinate system G relative to the world coordinate system W: Real-time positions of radar and camera in the world coordinate system: ; .

[0064] radar point Projected onto image pixels : ; ; .

[0065] 1.4 Pre-storage and recall of multi-focal length parameters.

[0066] Pre-stored calibration parameters for each focal length ,storage When detecting a target, intrinsic parameters are called or interpolated based on the current focal length f. Real-time calculations are performed as the gimbal rotates. and Ensure coordinate alignment during movement.

[0067] Step 2: Radar target detection and camera dynamic focus control.

[0068] Potential low-altitude, slow-moving, small targets (drones, hot air balloons) are extracted from raw radar point cloud data, and their precise azimuth and distance are calculated. The gimbal is adjusted based on the target's azimuth to ensure the target remains centered in the sensor's field of view. The camera's focal length is automatically adjusted based on the target's distance to guarantee optimal target resolution in the image.

[0069] 2.1 Radar point cloud acquisition and preprocessing.

[0070] Original point cloud of lidar Each point contains coordinates (x, y, z) and reflection intensity I. The point cloud is filtered, and statistical outlier filtering is used to remove noise. If a certain point... Average distance to neighboring points Greater than a threshold calculated based on the distribution of neighboring points , for If the standard deviation of the k nearest neighbors is given, then this point is identified as a noise point and removed. .

[0071] The noise-removed point cloud is subjected to target clustering. Using Euclidean distance, neighboring points are merged into candidate target clusters. A threshold is set to fit the size of small, slow, and large targets. Clusters with sufficient volume are retained, while objects that are too large or too small are excluded.

[0072] 2.2 Gimbal-based collaborative control.

[0073] Based on the target's 3D point cloud position detected by the lidar, the azimuth and pitch angles of the two-degree-of-freedom gimbal are controlled in real time to ensure that the target is always located in the center of the camera's field of view, providing stable target imaging for subsequent image recognition.

[0074] 2.3 Camera dynamic focus adjustment.

[0075] Target distance based on radar point cloud computing To calculate the ideal focal length based on imaging requirements and ensure the target maintains optimal resolution in the image, dynamic corrections are performed, and relevant focal length calculations are conducted: Given that the drone's altitude is approximately 0.5m, the required focal length is calculated as follows: Where h is the sensor height and H is the actual target height, which is 0.5m for the drone in this case. To determine the target distance using radar.

[0076] Finally, the parameter corresponding to the current focal length f is retrieved from the pre-stored calibration lookup table: the intrinsic parameter matrix K(f). If no pre-stored point exists for f, linear interpolation is used for calculation. .

[0077] 2.4 Real-time performance optimization.

[0078] Point cloud downsampling is achieved by using voxel filtering to reduce point cloud density and computational load. Traditional serial processes have high cumulative latency, so gimbal control and camera focusing are designed to be executed in parallel. Target position information, distance, and timestamp are transmitted through shared memory, and a circular buffer is used to avoid read / write conflicts. Hardware acceleration is achieved by deploying the point cloud algorithm on an FPGA, resulting in latency of <10ms.

[0079] Step 3: Low-speed, small target recognition method based on multimodal fusion.

[0080] This invention proposes an improved BEV fusion framework, which solves the challenge of identifying small, slow, and low-speed targets through the following core innovations. For example... Figure 1 For the overall network architecture.

[0081] 3.1 Multi-scale feature fusion network.

[0082] In the image feature extraction module, traditional CNNs have limited ability to extract features from zoom camera images, resulting in the loss of details of small targets at a distance. This invention embeds conditionally parameterized convolutions (CondConv) into each residual block of the ResNet-50: ; The focal length-related weights are generated by a lightweight MLP: ; The default expert convolution kernel.

[0083] In FPN, a focal length-aware feature option is introduced, with fusion weights for features at various scales. For telephoto lenses, the P5 layer (high semantic features) is strengthened, while for wide-angle lenses, the P2 layer (detail features) is emphasized. When the system is in telephoto mode (f>100mm), the P5 layer is strengthened using the following methods: weight allocation, and MLP output... The weights are increased by 0.6-0.8. When the system is in a wide-angle state (f<50mm), the P2 layer is enhanced in the following way: MLP output The weight increases by 0.5-0.7.

[0084] In the point cloud feature extraction module, the point cloud is first voxelized and converted into a voxel mesh. VoxelNet is then used to generate a Bev feature map to extract features such as the height, density, and reflection intensity of the target, which are then concatenated with the learned features.

[0085] 3.2 Deformable BEV projection.

[0086] For each BEV grid point (x, y), predict its sampling offset in the image based on the scale factor S. : s; where z is the depth provided by the radar, and f is the current focal length.

[0087] Perform dynamic extrinsic parameter compensation and integrate calibration parameters during view transformation. The corrected projected coordinates, where To correct the initial depth value of the foreground point in the camera coordinate system: .

[0088] 3.3 Cross-modal attention fusion.

[0089] Radar BEV characteristics are The BEV features of the image are Through learnable weight matrices Generate Q / K / V. Q= K= V= Perform attention calculations: B(f) is generated by a lightweight MLP based on the current focal length f and is used to adjust the mode weights. At long distances (telephoto), image details are more reliable, so their weight is increased; at close distances (wide-angle), radar spatial accuracy is higher, so the weight of point cloud features is increased.

[0090] 3.4 Optimize the detection head for small targets.

[0091] Improved CenterPoint header, input For small targets, optimization is performed using high-resolution BEV feature maps. Anchor dynamic sizes are designed: 1.0m x 1.0m for focal lengths of 20-50mm; 0.5m x 0.5m for focal lengths of 50-100mm; and 0.3m x 0.3m for focal lengths of 100-200mm.

[0092] 3.5 Real-time performance optimization.

[0093] Point cloud voxelization and BEV generation are deployed on an FPGA using a parallel architecture, processing point cloud data once per clock cycle, resulting in latency <5ms. CondConv's expert convolutional kernels are preloaded into the NPU cache, requiring only updates to the weight coefficients π(f) when the focal length changes, with communication overhead <1KB. The model is lightweighted, and channel pruning is implemented. During training, L1 regularization is added to the convolutional layers; during evaluation, channels with weights less than a threshold (threshold = 0.01 × max_weight) are removed.

[0094] Step 4: Target strike control and execution.

[0095] After target identification, it is necessary to engage low-lying, slow-moving, and small targets. Figure 2 This refers to the overall strike process.

[0096] 4.1 Strike Decision Module.

[0097] Based on the category confidence score derived from the target recognition module and the distance to the gimbal, a target threat assessment is performed. ; Category confidence level. Drones = 1.0, Birds = 0.1.

[0098] = D is the distance parameter. = Speed ​​threat standardization; threshold judgment is performed, and if the threat index S is greater than 0.7, the following strike steps are taken.

[0099] 4.2 Laser emission control.

[0100] The gimbal precisely aims and adjusts the pitch / azimuth angles based on the predicted position. The laser emitter is activated. The gimbal control algorithm predicts motion. , = Communication delay + gimbal response time; 'a' is estimated by a Kalman filter; x, y, z are the current three-dimensional spatial coordinates, and v is the current velocity.

[0101] If the laser fails to destroy the target on its first strike, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated based on the offset.

[0102] 4.3 Evaluation of the effectiveness of the strike.

[0103] Unmanned aerial vehicle (UAV) inspection: laser receiving device, visual verification device.

[0104] Laser receiver: photodiode array (16*160); trigger threshold: >500W / , lasting 100ms.

[0105] Visual verification: A smoke alarm device is installed on the drone. If the drone is hit, the smoke system is triggered.

[0106] 4.4 Real-time performance guarantee.

[0107] Decision-making is accelerated by FPGA, employing parallel scoring computation and a SIMD architecture to complete threat scoring in a single cycle, reducing computational latency. Real-time control loops and dual-loop control are used for gimbal calibration. Communication optimizations include low-latency data transmission, protocol stack optimization, data compression, and differential encoding for strike commands.

[0108] The technical objectives of this invention are reflected in the following multiple dimensions: First, at the sensor collaboration level, this invention aims to establish an optimal cooperation mechanism between LiDAR and a zoom camera. In traditional solutions, these two sensors often operate independently, leading to the accumulation of spatiotemporal registration errors, especially during camera zooming, where the continuity of target tracking is difficult to guarantee. This invention develops a real-time calibration system based on zoom camera feedback to achieve dynamic parameter compensation when the focal length changes, ensuring that the registration error between the radar point cloud and the optical image is always controlled within a certain pixel range within the 20-200mm zoom range. Encoder feedback is used to update the sensor pose in real time, solving the coordinate offset caused by motion. This innovation enables the system to adapt to changes in target distance, maintaining high resolution during long-range detection and a large field of view coverage during close-range observation.

[0109] Secondly, at the target recognition level, single sensors such as radar require high point cloud density for target identification. When detecting small aerial targets at long distances, the point cloud often fails to meet the requirements of the target recognition algorithm. Cameras, on the other hand, rely heavily on image semantic information and cannot track targets within range in real time. This invention innovatively constructs an intelligent recognition architecture based on multimodal data collaborative perception. By deeply integrating the precise spatial perception capabilities of lidar with the rich semantic information of a zoom camera, it achieves high-precision recognition of low-flying, slow-moving, and small targets.

[0110] Finally, at the feature extraction level, this invention designs a dedicated multi-scale fusion network specifically for the characteristics of small, slow-moving targets. Traditional BEV fusion methods often result in the loss of small target information due to feature projection with fixed parameters when processing zoom camera data. This invention innovatively introduces a focal length conditional convolutional layer, enabling the network weights to be dynamically adjusted according to the real-time focal length value. This strengthens local texture feature extraction in telephoto mode and emphasizes global context awareness in wide-angle mode.

[0111] In summary, this invention integrates multidisciplinary technologies to construct a complete low-speed, small target detection solution. Its core value lies in: breaking through the physical limitations of a single sensor and achieving a leap in perception capability (1+1>2) through intelligent fusion; establishing an adaptive processing framework that enables the system to dynamically respond to various complex scenarios; and providing reliable real-time performance to meet the stringent requirements of critical security applications.

Claims

1. A method for detecting and recognizing small, slow-moving targets based on radar and a zoom camera, characterized in that... Includes the following steps: Step 1: Joint calibration and parameter pre-storage of LiDAR and zoom camera; The lidar and zoom camera are mounted on the same two-degree-of-freedom gimbal, and the relative motion between them is ensured. The lidar and zoom camera are jointly calibrated to obtain their intrinsic and extrinsic parameter matrices. Step 2: LiDAR target detection and dynamic focusing control of the zoom camera; Potential low-speed, small targets are extracted from the raw radar point cloud data, and their azimuth and distance are calculated. The gimbal is adjusted according to the target's azimuth to ensure that the target is always located in the center of the variable-focus camera's field of view. The variable-focus camera automatically adjusts its focal length according to the target's distance to ensure that the target's resolution in the image meets the requirements. In step 2, the detection and control of low-speed, small targets includes the following steps: Step 21: LiDAR point cloud acquisition and preprocessing; Original point cloud of lidar Each point contains coordinates (x, y, z) and reflection intensity. The point cloud is filtered, and statistical outlier filtering is used to remove noise. If a certain point... Average distance to neighboring points Greater than a threshold calculated based on the distribution of neighboring points , for If the standard deviation of the k nearest neighbors is given, then this point is identified as a noise point and removed. ; The point cloud after noise removal is subjected to target clustering. Euclidean distance is used to merge neighboring points into candidate target clusters. A threshold is set to fit the size of small, slow, and large targets. Clusters with sufficient volume are retained, while objects that are too large or too small are excluded. Step 22: Gimbal-based collaborative control; Based on the target's 3D point cloud position detected by the lidar, the azimuth and pitch angles of the two-degree-of-freedom gimbal are controlled in real time to ensure that the target is always located in the center of the camera's field of view, providing stable target imaging for subsequent image recognition. Step 23: Dynamic focusing of the zoom camera; Target distance based on radar point cloud computing In addition, the ideal focal length is calculated to meet imaging requirements, ensuring the target's resolution in the image, and dynamic corrections are performed, along with relevant focal length calculations. Given that the drone's altitude is approximately 0.5m, the required focal length can be calculated using the focal length calculation method. ; Where h is the sensor height, and H is the actual target height, which in this case is the drone height. To determine the target distance using radar; The parameter corresponding to the current focal length f is retrieved from the pre-stored calibration lookup table: the intrinsic parameter matrix K(f); if there is no pre-stored point for f, linear interpolation is used for calculation. Step 24: Real-time optimization; Point cloud downsampling is achieved by using voxel filtering to reduce point cloud density and computational load. Traditional serial processes have high cumulative latency, so the gimbal control and camera focusing are designed to be executed in parallel. Target position information, distance, and timestamp are transmitted through shared memory, and a circular buffer is used to avoid read / write conflicts. Hardware acceleration is achieved by deploying the point cloud algorithm on an FPGA, resulting in latency of <10ms. Step 3: Identify low-speed, small targets based on multimodal fusion; In the image feature extraction module, conditional parameterized convolutions are embedded in each residual block; in the radar point cloud feature extraction module, the height, density, and reflection intensity features of the target are extracted and concatenated with the learned features; cross-modal attention fusion calculation is performed through deformable BEV projection to optimize small targets in real time; in step 3, the identification of low, slow, and small targets includes the following steps: Step 31: In the image feature extraction module, a ResNet-50 network is used, embedding conditional parameterized convolutions in each residual block; an FPN network is used, introducing focal length-aware feature options, inputting the fusion weights of features at each scale. For telephoto lenses, the P5 layer (high semantic features) is strengthened, while for wide-angle lenses, the P2 layer (detail features) is emphasized. In Step 31, the focal length-aware feature option is a multi-scale feature fusion mechanism based on dynamic weight allocation. Its network structure is a conditional parameterized convolution, which is represented as: ; in, The focal length-related weights are generated by a lightweight MLP: ; The pre-defined expert convolution kernel; The fusion weight input FPN network adopts a dynamic feature fusion algorithm. After the FPN features are constructed, dynamic weights are injected when multi-scale feature fusion is performed. The weights are applied by multiplying element by element. When the system is in telephoto mode (f>100mm), the P5 layer is enhanced in the following ways: weight allocation, MLP output... The weight is increased by 0.6-0.8; when the system is in a wide-angle state (f<50mm), the P2 layer is strengthened in the following way: MLP output The weighting increases by 0.5-0.7; Step 32: In the radar point cloud feature extraction module, the point cloud is first voxelized and converted into a voxel mesh. VoxelNet is used to generate a Bev feature map to extract the target's height, density, and reflection intensity features, which are then concatenated with the learned features. Step 33: Perform deformable BEV projection; in step 33, deformable BEV projection includes: For each BEV grid point (x, y), predict its sampling offset in the image based on the scale factor S. : s; Where z is the depth provided by the radar, and f is the current focal length; Perform dynamic extrinsic parameter compensation and integrate calibration parameters during view transformation. Corrected projected coordinates: ; in To correct the initial depth value of the foreground point in the camera coordinate system; In step 34, the cross-modal attention fusion computation includes: Through learnable weight matrix Generate Q, K, and V matrices, where Q = K= V= ,F L For lidar point cloud features, F I For image features, d k The dimension of the key vector K: ; B(f) is generated by a lightweight MLP based on the current focal length f and is used to adjust the mode weights. Learnable scaling parameters: ; At long distances (telephoto), increase the weight of image details; at close distances (wide-angle), increase the weight of point cloud features. Step 34: Let the radar BEV characteristic be... The BEV features of the image are Cross-modal attention fusion computation is performed using a learnable weight matrix; Step 35: Real-time optimization of small targets by improving the detection head; Step 4: Target strike control and execution; Based on the target recognition module, the system determines the category confidence level and the distance to the gimbal, performs a target threat assessment, and then conducts a laser strike. The system also evaluates the strike effect. If the first strike fails to destroy the target, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated based on the offset until the target is successfully shot down and the attack stops.

2. The method for detecting and recognizing small, slow-moving targets based on radar and a variable-focus camera as described in claim 1, characterized in that, In step S1, a coordinate system is established, with the world coordinate system as W, the gimbal coordinate system as G, the lidar coordinate system as L, and the zoom camera coordinate system as C. The joint calibration includes the intrinsic parameter calibration of the zoom camera, the relative extrinsic parameter calibration of the lidar and the zoom camera, and the gimbal motion compensation.

3. The method for detecting and recognizing small, slow-moving targets based on radar and a variable-focus camera as described in claim 2, characterized in that, The intrinsic parameter calibration of a zoom camera is to perform focal length divergence on the zoom camera within the focal length range [ , Select N calibration points within [the specified area]. , ... Fixed focal length to Multiple sets of chessboard images with different poses were captured, and the intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: ; Where α is the weighting coefficient, which calculates the contribution ratio of the target focal length in the known focal range; The relative extrinsic parameter calibration between the LiDAR and the zoom camera is performed by placing the AprilTag calibration board in a static scene, ensuring it appears in the field of view of both the LiDAR and the zoom camera, and measuring the coordinates of the corner points of the calibration board in the world coordinate system W. That is, to calculate the transformation from the world coordinate system to the calibration plate. ; Detecting the AprilTag reveals the pose of the calibration board in the zoom camera coordinate system. ; Extract the point cloud from the calibration board, fit the plane equation, and calculate the pose in the lidar coordinate system. The pose of the lidar to the zoom camera was calculated using the world coordinates of the calibration board. ; Gimbal motion compensation is used to calibrate the installation positions of the radar and camera in the gimbal coordinate system G. and The pitch angle is obtained through the gimbal encoder. and azimuth Calculate the rotational transformation of the gimbal coordinate system G relative to the world coordinate system W: ; Real-time positions of the lidar and zoom camera in the world coordinate system: ; ; radar point coordinates Transform to world coordinate system Then to the camera coordinate system Finally, the image pixel coordinates are projected onto the camera intrinsic parameters to obtain... : ; ; 。 4. The method for detecting and recognizing small, slow-moving targets based on radar and a zoom camera as described in claim 3, characterized in that, Pre-store the calibration parameters for each focal length. ,storage When detecting a target, based on the current focal length f, intrinsic parameters are called or interpolated, and calculations are performed in real time as the gimbal rotates. and Ensure coordinate alignment during movement.

5. The method for detecting and recognizing small, slow-moving targets based on radar and a zoom camera as described in claim 1, characterized in that, In step 35, the dynamic dimensions of the anchor are designed: for a focal length range of 20–50 mm, the anchor size is 1.0 m × 1.0 m; for a focal length range of 50–100 mm, the anchor size is 0.5 m × 0.5 m; and for a focal length range of 100–200 mm, the anchor size is 0.3 m × 0.3 m. Point cloud voxelization and BEV generation are deployed on FPGA using a parallel architecture, processing point cloud data once per clock cycle, resulting in a latency of <5ms. The expert convolutional kernels of CondConv are preloaded into the NPU cache, requiring only updates to the weight coefficients π(f) when the focal length changes, with a communication volume of <1KB. The model is lightweighted, channels are pruned, and L1 regularization is added to the convolutional layers during training. During the evaluation phase, channels with weights below the threshold are removed for real-time optimization.

6. The method for detecting and recognizing small, slow-moving targets based on radar and a variable-focus camera as described in claim 1, characterized in that, In step 4, achieving a successful strike on the target includes the following steps: S41: Establish a strike decision-making module; Based on the target recognition module, the category confidence level and the distance to the gimbal are used to perform a target threat assessment: ; in, For category confidence, 1 represents the weighting coefficient for the class confidence score. 1 represents the weighting coefficient of the distance penalty term. The weighting coefficient for the speed threat term; = D is the distance parameter; = Speed ​​threat standardization; A threshold judgment is performed. If the threat index S is greater than 0.7, the following strike steps are taken. S42: Laser emission control; The gimbal precisely aims, adjusts the pitch / azimuth angle based on the predicted position, activates the laser emitter, and the gimbal control algorithm predicts the motion. = Communication delay + PTZ response time; a is the acceleration obtained by the Kalman filter, x, y, z are the three-dimensional spatial coordinates at the current moment, and v is the velocity at the current moment; If the laser fails to destroy the target on the first strike, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated according to the offset. S43: Assessment of strike effectiveness; Unmanned aerial vehicle (UAV) end-to-end inspection, including laser receiving devices and visual verification devices; Laser receiver: Photodiode array is 16×160; trigger threshold: >500W / Lasts for 100ms; Visual verification device: A smoke alarm device is installed on the drone. If the drone is hit successfully, the smoke system is triggered. S44: Real-time performance guarantee; The decision-making is accelerated by FPGA, adopts parallel scoring calculation, SIMD architecture, completes threat scoring in a single cycle, reduces computing latency, real-time control loop, gimbal calibration adopts calibration dual-loop control, communication optimization, low-latency data transmission, protocol stack optimization, data compression, and attack commands adopt differential encoding.

Citation Information

Patent Citations

  • Target detection and identification device and method based on multi-fusion sensor

    CN110428008A

  • Low, small and slow target identification and positioning method and system based on mixed vision

    CN115035470A