Low-slow small target detection and identification method based on radar and zoom camera

By combining the calibration of lidar and a zoom camera with multimodal fusion technology, the problem of identifying small, slow, and low-speed targets in complex environments using traditional detection methods has been solved, achieving high-precision, fast-response, and all-weather stable target identification and strike capabilities.

CN121069415AActive Publication Date: 2025-12-05CHINA UNIV OF MINING & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511606094.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2025-12-05
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Traditional detection methods struggle to efficiently and reliably identify low, slow, and small targets in complex environments, especially under conditions of small radar cross-section, complex terrain, and adverse weather, where their detection capabilities are limited and their dynamic response is insufficient.

Method used

By employing joint calibration and parameter pre-storage of LiDAR and variable-focus camera, combined with multimodal fusion technology, and through dynamic focusing of variable-focus camera and cross-modal attention fusion, real-time detection and recognition of low, slow, and small targets can be achieved.

Benefits of technology

It significantly improves the identification accuracy of low, slow, and small targets and the system's adaptability to complex environments, ensuring stable operation around the clock, shortening deployment time, reducing hardware costs, and enabling rapid response and efficient strikes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069415A_ABST
    Figure CN121069415A_ABST
Patent Text Reader

Abstract

The invention discloses a low-slow small target detection and identification method based on a radar and a zoom camera. The method comprises the following steps: step 1, joint calibration and parameter pre-storage of the laser radar and the zoom camera; 2, laser radar target detection and variable-focus camera dynamic focusing control are carried out; step 3, identifying a low-slow small target based on multi-modal fusion; and step 4, target strike control and execution. The physical limitation of a single sensor is broken through, and 1 + 1gt is realized through intelligent fusion; 2, the sensing capability is improved; an adaptive processing framework is established, so that the system can dynamically deal with various complex scenes; reliable real-time performance is provided, and strict requirements of key security application are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting and recognizing small, slow targets, and more particularly to a method for detecting and recognizing small, slow targets based on radar and a zoom camera. Background Technology

[0002] With the rapid development and widespread application of unmanned aerial vehicle (UAV) technology, the demand for the detection and identification of "low, slow, and small" targets has increased significantly. In complex battlefield environments, traditional detection methods can no longer meet the requirements for efficient and reliable detection of moving targets. The widespread use of "low, slow, and small" weapons such as UAVs and loitering munitions has brought new challenges to the modern battlefield. Due to their small size and small radar cross-section, "low, slow, and small" weapons are difficult to detect effectively by traditional radar systems. Their stealth and mobility make these targets a significant hidden danger, thus requiring rapid and accurate identification and interception of these targets.

[0003] There are four common methods for detecting "low, slow and small" flying objects: radar detection, radio detection, acoustic detection and photoelectric detection.

[0004] Radar detection utilizes the reflection of radio waves to detect and locate targets. However, radar is susceptible to interference from ground reflections and clutter, especially in complex terrain; it is particularly difficult to detect low-flying, slow-moving, and small objects, especially those with low radar cross-sections (RCS); it is also costly, requires complex equipment, and the false alarm rate of radar systems increases by more than 40% in complex terrain.

[0005] Radio detection utilizes the interaction between a target and radio signals, analyzing changes in these signals to detect the target. However, it has a limited detection range, is suitable for close-range targets, is susceptible to interference from other radio equipment leading to signal misinterpretation, and performs poorly in detecting high-speed targets.

[0006] Acoustic detection utilizes the sound signals emitted by the target or detects targets by actively emitting sound waves and receiving reflected waves. However, it has a limited detection range, making it suitable for short-range monitoring; it is greatly affected by environmental noise and easily interfered with by natural factors such as wind and rain; and it is less effective at detecting low-noise or silent flying objects.

[0007] Photoelectric detection utilizes optical sensors to capture images or infrared radiation of targets, and then uses image processing techniques and algorithms to analyze and identify them. However, photoelectric detection is greatly affected by weather and lighting conditions, such as fog, rain, snow, and strong light, which can affect detection results; its detection range is relatively limited, requiring line-of-sight detection; and processing image data requires high computing resources, with optical sensor performance decreasing by up to 70% in foggy or hazy weather. Summary of the Invention

[0008] Purpose of the invention: To address the shortcomings of existing technologies, this invention provides a method for detecting and identifying low-speed, small targets based on radar and a variable-focus camera. This method solves the problems of poor environmental adaptability, limited detection capability, and insufficient dynamic response in detecting low-speed, small flying targets, improves the real-time detection and identification performance of low-speed, small targets in complex environments, and meets the urgent needs of modern urban air defense.

[0009] Summary of the invention: A method for detecting and recognizing small, slow targets based on radar and a zoom camera, comprising the following steps: Step 1: Joint calibration and parameter pre-storage of LiDAR and zoom camera; The LiDAR and zoom camera are mounted on the same two-degree-of-freedom gimbal, ensuring no relative motion between them, and the LiDAR and zoom camera are jointly calibrated to obtain their intrinsic and extrinsic parameter matrices.

[0010] Step 2: LiDAR target detection and dynamic focusing control of the zoom camera; extract potential low-speed small targets from the original radar point cloud data, calculate the target's azimuth and distance, adjust the gimbal pointing according to the target's azimuth to ensure that the target is always in the center of the zoom camera's field of view, and automatically adjust the camera's focal length according to the target's distance to ensure that the target's resolution in the image meets the requirements.

[0011] Step 3: Identify low-speed, small targets based on multimodal fusion; in the image feature extraction module, embed conditional parameterized convolutions into each residual block; in the radar point cloud feature extraction module, extract the height, density, and reflection intensity features of the target and concatenate them with the learned features; perform cross-modal attention fusion calculation through deformable BEV projection to optimize small targets in real time.

[0012] Step 4: Target strike control and execution; Based on the target identification module, the category confidence level and distance to the gimbal are obtained, the target threat is assessed, and then laser strike is carried out. The strike effect is evaluated. If the target is not destroyed on the first strike, the system automatically triggers parameter adjustment, and the gimbal aiming point is dynamically compensated according to the offset until the target is successfully shot down and the attack stops.

[0013] Furthermore, in step 1, a coordinate system is established, with the world coordinate system as W, the gimbal coordinate system as G, the lidar coordinate system as L, and the zoom camera coordinate system as C. The joint calibration includes the intrinsic parameter calibration of the zoom camera, the relative extrinsic parameter calibration of the lidar and the zoom camera, and the gimbal motion compensation.

[0014] Ideally, the intrinsic parameters of a zoom camera should be calibrated to reduce focal length divergence within the focal length range. , Select N calibration points within [the specified area]. , ... Fixed focal length to Multiple sets of chessboard images with different poses were captured, and the intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: Where α is the weighting coefficient, which calculates the contribution ratio of the target focal length in the known focal length range.

[0015] The relative extrinsic parameter calibration between the LiDAR and the zoom camera is performed by placing the AprilTag calibration board in a static scene, ensuring it appears in the field of view of both the LiDAR and the zoom camera, and measuring the coordinates of the corner points of the calibration board in the world coordinate system W. That is, to calculate the transformation from the world coordinate system to the calibration plate. ; Detect the AprilTag to obtain the pose of the calibration board in the zoom camera coordinate system. Extract the point cloud from the calibration board, fit the plane equation, and calculate the pose in the lidar coordinate system. The pose of the lidar to the zoom camera was calculated using the world coordinates of the calibration board. .

[0016] Gimbal motion compensation is used to calibrate the installation positions of the radar and camera in the gimbal coordinate system G. and The pitch angle is obtained through the gimbal encoder. and azimuth Calculate the rotation transformation of the gimbal coordinate system G relative to the world coordinate system W: Real-time positions of the lidar and zoom camera in the world coordinate system: ; .

[0017] radar point coordinates Transform to world coordinate system Then to the camera coordinate system Finally, the image pixel coordinates are projected onto the camera intrinsic parameters to obtain... : ; ; .

[0018] Ideally, the calibration parameters should be pre-stored for each focal length. ,storage When detecting a target, based on the current focal length f, intrinsic parameters are called or interpolated, and calculations are performed in real time as the gimbal rotates. and Ensure coordinate alignment during movement.

[0019] Furthermore, in step 2, the detection and control of low-speed, small targets includes the following steps: Step 21: LiDAR point cloud acquisition and preprocessing; raw LiDAR point cloud Each point contains coordinates (x, y, z) and reflection intensity. The point cloud is filtered, and statistical outlier filtering is used to remove noise. If a certain point... Average distance to neighboring points Greater than a threshold calculated based on the distribution of neighboring points , for If the standard deviation of the k nearest neighbors is given, then this point is identified as a noise point and removed. The noise-removed point cloud is subjected to target clustering. Euclidean distance is used to merge neighboring points into candidate target clusters. A threshold is set to fit the size of small, slow, and large targets. Clusters with sufficient volume are retained, while objects that are too large or too small are excluded.

[0020] Step 22: Gimbal-based collaborative control; Based on the target's 3D point cloud position detected by the lidar, control the azimuth and pitch angles of the two-degree-of-freedom gimbal in real time to ensure that the target is always located in the center of the camera's field of view, providing stable target imaging for subsequent image recognition.

[0021] Step 23: Dynamic focusing of the zoom camera; target distance based on radar point cloud computing. In addition, the ideal focal length is calculated to meet imaging requirements, ensuring the target's resolution in the image. Dynamic corrections are then performed, and relevant focal length calculations are conducted. Given that the drone's altitude is approximately 0.5m, the required focal length is calculated. Where h is the sensor height and H is the actual target height, which in this case is the drone's height. The radar measures the target distance; the parameters corresponding to the current focal length f are retrieved from the pre-stored calibration lookup table: the intrinsic parameter matrix K(f); if there is no pre-stored point for f, linear interpolation is used for calculation.

[0022] Step 24: Real-time optimization; Point cloud downsampling, using voxel filtering to reduce point cloud density and reduce computation; Traditional serial processes have high cumulative latency, so gimbal control and camera focusing are designed to be executed in parallel, target position information, distance, and timestamp are transmitted through shared memory, and a circular buffer is used to avoid read and write conflicts; Through hardware acceleration, the point cloud algorithm is deployed on FPGA, making the latency <10ms.

[0023] In step 3, the recognition of low, slow, and small targets includes the following steps: Step 31: In the image feature extraction module, a ResNet-50 network is used, and conditional parameterized convolutions are embedded in each residual block; an FPN network is used, in which focal length-aware feature options are introduced, and the fusion weights of features at each scale are input. When using telephoto lenses, the P5 layer, i.e., high semantic features, is strengthened, and when using wide-angle lenses, the P2 layer, i.e., detail features, is emphasized.

[0024] Step 32: In the radar point cloud feature extraction module, the point cloud is first voxelized and converted into a voxel mesh. VoxelNet is used to generate a Bev feature map to extract the target's height, density, and reflection intensity features, which are then concatenated with the learned features.

[0025] Step 33: Perform deformable BEV projection.

[0026] Step 34: Let the radar BEV characteristic be... The BEV features of the image are Cross-modal attention fusion computation is performed using a learnable weight matrix.

[0027] Step 35: Real-time optimization of small targets by improving the detection head.

[0028] Furthermore, in step 31, the parameterized convolution is represented as: ;in, The focal length-related weights are generated by a lightweight MLP: ; The pre-defined expert convolutional kernels are used; the fusion weight input FPN network employs a dynamic feature fusion algorithm. After the FPN features are constructed, dynamic weights are injected during multi-scale feature fusion, and weights are applied using element-wise multiplication. When the system is in telephoto mode (f>100mm), the P5 layer is strengthened in the following way: weight allocation, MLP output... The weight is increased by 0.6-0.8; when the system is in a wide-angle state (f<50mm), the P2 layer is strengthened in the following way: MLP output The weight increases by 0.5-0.7.

[0029] Furthermore, in step 33, the deformable BEV projection includes: for each BEV grid point (x,y), predicting its sampling offset in the image based on the scale factor S. :

[0030] s; where z is the depth provided by the radar, and f is the current focal length; dynamic extrinsic parameter compensation is performed, and calibration parameters are integrated in the view transformation. Corrected projected coordinates: ;in To correct the initial depth value of the foreground point in the camera coordinate system.

[0031] Furthermore, in step 34, the cross-modal attention fusion calculation includes: using a learnable weight matrix Generate Q, K, and V matrices, where Q = K= V= ,F L For lidar point cloud features, F I For image features, d k The dimension of the key vector K: B(f) is generated by a lightweight MLP based on the current focal length f and is used to adjust the mode weights. Learnable scaling parameters: At long distances (telephoto), increase the weight of image details; at close distances (wide-angle), increase the weight of point cloud features.

[0032] Ideally, in step 35, the dynamic dimensions of the anchor are designed as follows: for a focal length range of 20–50 mm, the anchor size is 1.0 m × 1.0 m; for a focal length range of 50–100 mm, the anchor size is 0.5 m × 0.5 m; and for a focal length range of 100–200 mm, the anchor size is 0.3 m × 0.3 m.

[0033] Point cloud voxelization and BEV generation are deployed on FPGA using a parallel architecture, processing point cloud data once per clock cycle, resulting in a latency of <5ms. The expert convolutional kernels of CondConv are preloaded into the NPU cache, requiring only updates to the weight coefficients π(f) when the focal length changes, with a communication volume of <1KB. The model is lightweighted, channels are pruned, and L1 regularization is added to the convolutional layers during training. During the evaluation phase, channels with weights below the threshold are removed for real-time optimization.

[0034] Ideally, in step 4, achieving a successful target strike includes the following steps: S41: Establish a strike decision module; based on the category confidence level and distance to the gimbal obtained from the target identification module, perform a target threat assessment: ;in, For category confidence, 1 represents the weighting coefficient for the class confidence score. 1 represents the weighting coefficient of the distance penalty term. This represents the weighting coefficient for the speed threat term.

[0035] = D is the distance parameter.

[0036] = Speed ​​threat standardization.

[0037] A threshold judgment is performed. If the threat index S is greater than 0.7, the following strike steps are taken.

[0038] S42: Laser emission control; precise gimbal aiming, adjusting pitch / azimuth angles based on predicted position, activating the laser emitter, gimbal control algorithm, motion prediction results: .

[0039] = Communication delay + gimbal response time; a is the acceleration obtained by the Kalman filter, x, y, z are the three-dimensional spatial coordinates at the current moment, and v is the velocity at the current moment.

[0040] If the laser fails to destroy the target on its first strike, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated based on the offset.

[0041] S43: Strike effect assessment; UAV-side detection, including laser receiver and visual verification device; Laser receiver: photodiode array is set to 16×160; Trigger threshold: >500W / The duration is 100ms; Visual verification device: A smoke alarm device is installed on the drone. If the hit is successful, the smoke system is triggered.

[0042] S44: Real-time performance assurance; FPGA-accelerated decision-making, parallel scoring calculation, SIMD architecture, threat scoring completed in a single cycle, reducing computational latency, real-time control loop, gimbal calibration using dual-loop control, communication optimization, low-latency data transmission, protocol stack optimization, data compression, and differential encoding for strike commands.

[0043] Beneficial Effects: Compared with existing technologies, the beneficial effects of this invention are: 1. Multimodal dynamic fusion significantly improves recognition accuracy. Traditional target recognition methods often use a single sensor, making it difficult to balance long-range detection accuracy and short-range recognition capability. This invention innovatively combines the precise three-dimensional ranging of LiDAR with the multi-scale visual features of a zoom camera, achieving optimal fusion of cross-modal features through a focal length-adaptive dynamic convolutional network and deformable BEV projection technology.

[0044] 2. Intelligent zoom-coordinated processing ensures rapid response. Addressing the dynamic characteristics of low-altitude, slow-moving, and small targets, this invention designs a distance-based adaptive zoom control algorithm and a hardware-accelerated processing pipeline. Point cloud voxelization (<5ms) and NPU-accelerated dynamic convolution calculation (0.5ms) are implemented using FPGA, enabling the entire process latency from target detection to strike decision to be controlled within 30ms.

[0045] 3. Robust in complex environments, stable operation in all weather conditions. The system integrates multispectral sensor data (visible light / infrared / LiDAR) and, through a deep learning-assisted clutter suppression algorithm, maintains a high detection rate even in harsh environments such as fog (visibility <100m), strong light (100,000 lux), rain, and snow. A unique focal length compensation mechanism ensures continuous and stable tracking during zooming, solving the performance degradation problem caused by environmental changes in traditional systems.

[0046] 4. Modular design for flexible and convenient deployment. Utilizing standardized hardware interfaces and the ROS communication framework, it supports plug-and-play compatibility with various LiDAR models and zoom cameras. Innovative dynamic calibration technology allows for pre-calibration parameters to cover the entire focal length range, shortening deployment time. The system can be flexibly integrated into various platforms such as vehicle-mounted and fixed stations to meet diverse operational needs.

[0047] 5. It has a high cost-effectiveness ratio and the potential for large-scale application. Through lightweight model design (INT8 quantization + channel pruning) and heterogeneous computing architecture optimization, it can achieve 4K@30fps real-time processing on edge devices such as Jetson Orin, reducing hardware costs.

[0048] In summary, this invention effectively solves the core problems of "unclear visibility, inaccurate tracking, and inability to hit" in the detection of low, slow, and small targets through an innovative "perception-fusion-decision-strike" end-to-end technical solution. Attached Figure Description

[0049] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are illustrative and should not be construed as limiting the invention in any way.

[0050] Figure 1 A schematic diagram for improving the BEV network architecture.

[0051] Figure 2 The flowchart for the closed-loop logic control. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] The present invention will be further illustrated below with reference to specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Modifications to the present invention in various equivalent forms all fall within the scope defined by the appended claims.

[0054] This invention provides a method for detecting and recognizing small, slow targets based on radar and a zoom camera, including the following steps: Step 1: Joint calibration and parameter pre-storage of lidar and zoom camera.

[0055] The lidar and the zoom camera are mounted on the same two-degree-of-freedom gimbal, with no relative motion between them. Joint calibration of the two sensors yields the intrinsic and extrinsic parameter matrices for each sensor.

[0056] Coordinate system conventions: W is the world coordinate system; G is the gimbal coordinate system; L is the lidar coordinate system; C is the camera coordinate system.

[0057] 1.1 Camera intrinsic parameter calibration.

[0058] Focal distance divergence, within the focal length range [ , Select N calibration points within a range of 20mm to 200mm (e.g., ...). =20mm, =30mm,……, =200mm). Fixed focal length to Multiple sets of chessboard images with different poses were captured. The intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( ), where α is the weighting coefficient, which calculates the contribution ratio of the target focal length in the known focal length range.

[0059] Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: .

[0060] 1.2 Radar and camera relative extrinsic parameter calibration.

[0061] Place the AprilTag calibration board in a static scene, ensuring it appears in the field of view of both the radar and the camera. Accurately measure the coordinates of the corner points of the calibration board in the world coordinate system W. The pose of the calibration board in the camera coordinate system is obtained by detecting the AprilTag. Extract the point cloud from the calibration board, fit the plane equation, and calculate the pose in the radar coordinate system. World coordinates via calibration plate Calculate the camera pose in the world coordinate system: .

[0062] 1.3 Gimbal Motion Compensation.

[0063] Calibrate the installation positions of the radar and camera in the gimbal coordinate system G. and (Measurement requires the gimbal to be in the zero position). The pitch angle is obtained via the gimbal encoder. and azimuth Calculate the rotation transformation of the gimbal coordinate system G relative to the world coordinate system W: Real-time positions of radar and camera in the world coordinate system: ; .

[0064] radar point Projected onto image pixels : ; ; .

[0065] 1.4 Pre-storage and recall of multi-focal length parameters.

[0066] Pre-stored calibration parameters for each focal length ,storage When detecting a target, intrinsic parameters are called or interpolated based on the current focal length f. Real-time calculations are performed as the gimbal rotates. and Ensure coordinate alignment during movement.

[0067] Step 2: Radar target detection and camera dynamic focus control.

[0068] Potential low-speed, small targets (drones, hot air balloons) are extracted from raw radar point cloud data, and their precise azimuth and distance are calculated. The gimbal is adjusted based on the target's azimuth to ensure the target remains centered in the sensor's field of view. The camera focal length is automatically adjusted based on the target's distance to guarantee optimal target resolution in the image.

[0069] 2.1 Radar point cloud acquisition and preprocessing.

[0070] Original point cloud of lidar Each point contains coordinates (x, y, z) and reflection intensity I. The point cloud is filtered, and statistical outlier filtering is used to remove noise. If a certain point... Average distance to neighboring points Greater than a threshold calculated based on the distribution of neighboring points , for If the standard deviation of the k nearest neighbors is given, then this point is identified as a noise point and removed. .

[0071] The noise-removed point cloud is subjected to target clustering. Using Euclidean distance, neighboring points are merged into candidate target clusters. A threshold is set to fit the size of small, slow, and large targets. Clusters with sufficient volume are retained, while objects that are too large or too small are excluded.

[0072] 2.2 Gimbal-based collaborative control.

[0073] Based on the target's 3D point cloud position detected by the lidar, the azimuth and pitch angles of the two-degree-of-freedom gimbal are controlled in real time to ensure that the target is always located in the center of the camera's field of view, providing stable target imaging for subsequent image recognition.

[0074] 2.3 Camera dynamic focus adjustment.

[0075] Target distance based on radar point cloud computing To calculate the ideal focal length based on imaging requirements and ensure the target maintains optimal resolution in the image, dynamic corrections are performed, and relevant focal length calculations are conducted: Given that the drone's altitude is approximately 0.5m, the required focal length is calculated as follows: Where h is the sensor height and H is the actual target height, which is 0.5m for the drone in this case. To determine the target distance using radar.

[0076] Finally, the parameter corresponding to the current focal length f is retrieved from the pre-stored calibration lookup table: the intrinsic parameter matrix K(f). If no pre-stored point exists for f, linear interpolation is used for calculation. .

[0077] 2.4 Real-time performance optimization.

[0078] Point cloud downsampling is achieved by using voxel filtering to reduce point cloud density and computational load. Traditional serial processes have high cumulative latency, so gimbal control and camera focusing are designed to be executed in parallel. Target position information, distance, and timestamp are transmitted through shared memory, and a circular buffer is used to avoid read / write conflicts. Hardware acceleration is achieved by deploying the point cloud algorithm on an FPGA, resulting in latency of <10ms.

[0079] Step 3: Low-speed, small target recognition method based on multimodal fusion.

[0080] This invention proposes an improved BEV fusion framework, which solves the challenge of identifying small, slow, and low-speed targets through the following core innovations. For example... Figure 1 For the overall network architecture.

[0081] 3.1 Multi-scale feature fusion network.

[0082] In the image feature extraction module, traditional CNNs have limited ability to extract features from zoom camera images, resulting in the loss of details of small targets at a distance. This invention embeds conditionally parameterized convolutions (CondConv) into each residual block of the ResNet-50: ; The focal length-related weights are generated by a lightweight MLP: ; The default expert convolution kernel.

[0083] In FPN, a focal length-aware feature option is introduced, with fusion weights for features at various scales. For telephoto lenses, the P5 layer (high semantic features) is strengthened, while for wide-angle lenses, the P2 layer (detail features) is emphasized. When the system is in telephoto mode (f>100mm), the P5 layer is strengthened using the following methods: weight allocation, and MLP output... The weights are increased by 0.6-0.8. When the system is in a wide-angle state (f<50mm), the P2 layer is enhanced in the following way: MLP output The weight increases by 0.5-0.7.

[0084] In the point cloud feature extraction module, the point cloud is first voxelized and converted into a voxel mesh. VoxelNet is then used to generate a Bev feature map to extract features such as the height, density, and reflection intensity of the target, which are then concatenated with the learned features.

[0085] 3.2 Deformable BEV projection.

[0086] For each BEV grid point (x, y), predict its sampling offset in the image based on the scale factor S. : s; where z is the depth provided by the radar, and f is the current focal length.

[0087] Perform dynamic extrinsic parameter compensation and integrate calibration parameters during view transformation. The corrected projected coordinates, where To correct the initial depth value of the foreground point in the camera coordinate system: .

[0088] 3.3 Cross-modal attention fusion.

[0089] Radar BEV characteristics are The BEV features of the image are Through learnable weight matrices Generate Q / K / V. Q= K= V= Perform attention calculations: B(f) is generated by a lightweight MLP based on the current focal length f and is used to adjust the mode weights. At long distances (telephoto), image details are more reliable, so their weight is increased; at close distances (wide-angle), radar spatial accuracy is higher, so the weight of point cloud features is increased.

[0090] 3.4 Optimize the detection head for small targets.

[0091] Improved CenterPoint header, input For small targets, optimization is performed using high-resolution BEV feature maps. Anchor dynamic sizes are designed: 1.0m x 1.0m for focal lengths of 20-50mm; 0.5m x 0.5m for focal lengths of 50-100mm; and 0.3m x 0.3m for focal lengths of 100-200mm.

[0092] 3.5 Real-time performance optimization.

[0093] Point cloud voxelization and BEV generation are deployed on an FPGA using a parallel architecture, processing point cloud data once per clock cycle, resulting in latency <5ms. CondConv's expert convolutional kernels are preloaded into the NPU cache, requiring only updates to the weight coefficients π(f) when the focal length changes, resulting in communication overhead <1KB. The model is lightweighted, and channel pruning is implemented. During training, L1 regularization is added to the convolutional layers; during evaluation, channels with weights less than a threshold (threshold = 0.01 × max_weight) are removed.

[0094] Step 4: Target strike control and execution.

[0095] After target identification, it is necessary to engage low-lying, slow-moving, and small targets. Figure 2 This refers to the overall strike process.

[0096] 4.1 Strike Decision Module.

[0097] Based on the category confidence score derived from the target recognition module and the distance to the gimbal, a target threat assessment is performed. ; Category confidence level. Drones = 1.0, Birds = 0.1.

[0098] = D is the distance parameter. = Speed ​​threat standardization; threshold judgment is performed, and if the threat index S is greater than 0.7, the following strike steps are taken.

[0099] 4.2 Laser emission control.

[0100] The gimbal precisely aims and adjusts the pitch / azimuth angles based on the predicted position. The laser emitter is activated. The gimbal control algorithm predicts motion. , = Communication delay + gimbal response time; 'a' is estimated by a Kalman filter; x, y, z are the current three-dimensional spatial coordinates, and v is the current velocity.

[0101] If the laser fails to destroy the target on its first strike, the system automatically triggers parameter adjustments, and the gimbal aiming point is dynamically compensated based on the offset.

[0102] 4.3 Evaluation of the effectiveness of the strike.

[0103] Unmanned aerial vehicle (UAV) inspection: laser receiving device, visual verification device.

[0104] Laser receiver: photodiode array (16*160); trigger threshold: >500W / , lasting 100ms.

[0105] Visual verification: A smoke alarm device is installed on the drone. If the drone is hit, the smoke system is triggered.

[0106] 4.4 Real-time performance guarantee.

[0107] Decision-making is accelerated by FPGA, employing parallel scoring computation and a SIMD architecture to complete threat scoring in a single cycle, reducing computational latency. Real-time control loops and dual-loop control are used for gimbal calibration. Communication optimization includes low-latency data transmission, protocol stack optimization, data compression, and differential encoding for strike commands.

[0108] The technical objectives of this invention are reflected in the following multiple dimensions: First, at the sensor collaboration level, this invention aims to establish an optimal cooperation mechanism between LiDAR and a zoom camera. In traditional solutions, these two sensors often operate independently, leading to the accumulation of spatiotemporal registration errors, especially during camera zooming, where the continuity of target tracking is difficult to guarantee. This invention develops a real-time calibration system based on zoom camera feedback to achieve dynamic parameter compensation when the focal length changes, ensuring that the registration error between the radar point cloud and the optical image is always controlled within a certain pixel range within the 20-200mm zoom range. Encoder feedback is used to update the sensor pose in real time, solving the coordinate offset caused by motion. This innovation enables the system to adapt to changes in target distance, maintaining high resolution during long-range detection and a large field of view coverage during close-range observation.

[0109] Secondly, at the target recognition level, single sensors such as radar require high point cloud density for target identification. When detecting small aerial targets at long distances, the point cloud often fails to meet the requirements of the target recognition algorithm. Cameras, on the other hand, rely heavily on image semantic information and cannot track targets within range in real time. This invention innovatively constructs an intelligent recognition architecture based on multimodal data collaborative perception. By deeply integrating the precise spatial perception capabilities of lidar with the rich semantic information of a zoom camera, it achieves high-precision recognition of low-flying, slow-moving, and small targets.

[0110] Finally, at the feature extraction level, this invention designs a dedicated multi-scale fusion network specifically for the characteristics of small, slow-moving targets. Traditional BEV fusion methods often result in the loss of small target information due to feature projection with fixed parameters when processing zoom camera data. This invention innovatively introduces a focal length conditional convolutional layer, enabling the network weights to be dynamically adjusted according to the real-time focal length value. This strengthens local texture feature extraction in telephoto mode and emphasizes global context awareness in wide-angle mode.

[0111] In summary, this invention integrates multidisciplinary technologies to construct a complete low-speed, small target detection solution. Its core value lies in: breaking through the physical limitations of a single sensor and achieving a leap in perception capability (1+1>2) through intelligent fusion; establishing an adaptive processing framework that enables the system to dynamically respond to various complex scenarios; and providing reliable real-time performance to meet the stringent requirements of critical security applications.

Claims

1. A radar and variable zoom camera based low, slow and small target detection and recognition method, characterized in that The method comprises the following steps: Step 1: joint calibration and parameter pre-storage of laser radar and variable focus camera; The laser radar and the variable focus camera are installed on the same two-degree-of-freedom holder without relative motion between them, and joint calibration is performed on the laser radar and the variable focus camera to obtain internal parameter matrices and external parameter matrices of the two; Step 2: laser radar target detection and variable focus camera dynamic focusing control; Potential low, slow and small targets are extracted from original radar point cloud data, and target azimuth and distance are calculated, the holder is adjusted according to the target azimuth to ensure that the target is always located at the center of the field of view of the variable focus camera, and the variable focus camera automatically adjusts the camera focal length according to the target distance to ensure that the resolution of the target in the image meets the requirements; Step 3: identification of low, slow and small targets based on multi-modal fusion; In the image feature extraction module, a conditional parameterized convolution is embedded in each residual block; in the radar point cloud feature extraction module, the height, density and reflectivity features of the target are extracted and spliced with the learning features; Cross-modal attention fusion calculation is performed through deformable BEV projection to realize real-time optimization of the small target; Step 4: target attack control and execution; According to the category confidence and the distance from the holder obtained by the target identification module, the threat of the target is evaluated, then the laser is attacked, and the attack effect is evaluated; if the first attack does not achieve the target destruction, the system automatically triggers parameter adjustment, the holder aiming point is dynamically compensated according to the offset until the target is successfully shot down and the attack is stopped.

2. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 1, characterized in that, In step 1, coordinate systems are established, the world coordinate system is W, the holder coordinate system is G, the laser radar coordinate system is L, and the variable focus camera coordinate system is C; the joint calibration includes variable focus camera internal parameter calibration, laser radar and variable focus camera relative external parameter calibration, and holder motion compensation.

3. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 2, characterized in that, The intrinsic parameter calibration of a zoom camera is to perform focal length divergence on the zoom camera within the focal length range [ , Select N calibration points within [the specified area]. , ... Fixed focal length to Multiple sets of chessboard images with different poses were captured, and the intrinsic parameter K was calculated using OpenCV's calibrateCamera() function. ) and distortion D ( Establish an intrinsic parameter interpolation model for arbitrary focal lengths. If not directly calibrated, linear interpolation is used: ; Wherein, α is a weight coefficient, and the contribution proportion of the target focal length in the known focal length is calculated; The laser radar and the variable-focus camera are calibrated with respect to an external parameter. In a static scene, an AprilTag calibration board is placed to ensure that it is simultaneously in the field of view of the laser radar and the variable-focus camera, and the coordinates of the corner points of the calibration board in a world coordinate system W are measured , i.e. the transformation from the world coordinate system to the calibration board is calculated ; Detecting AprilTag to get the pose of the calibration board in the variable focal camera coordinate system ; Extract the calibration board point cloud, fit the plane equation, and calculate the pose in the laser radar coordinate system ; Through the calibration board world coordinates, the laser radar reaches the variable focus camera pose: ; Gimbal motion compensation is, calibrate the installation position of radar and camera in the gimbal coordinate system G and , get the pitch angle and azimuth angle by gimbal encoders , calculate the rotation transformation of the gimbal coordinate system G relative to the world coordinate system W: ; The real-time positions of the laser radar and the variable focus camera in the world coordinate system are as follows: ; ; Convert radar point coordinates to world coordinate system , to camera coordinate system , and finally project to image pixel coordinate system I via camera intrinsics : ; ; 。 4. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 3, characterized in that, Pre-store calibration parameters, for each focal length , store ; when detecting the target, call or interpolate the internal parameters according to the current focal length f, and calculate and in real time when the holder rotates, to ensure the alignment of coordinates in motion.

5. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 1, characterized in that, In step 2, the detection and control of the low, slow and small target comprises the following steps: Step 21: laser radar point cloud acquisition and preprocessing; Laser radar raw point cloud , each point contains coordinates (x, y, z) and reflection intensity, the point cloud is filtered, statistical outlier filtering is adopted to remove noise, if a point to the average distance of adjacent points is greater than a threshold value calculated based on the distribution of adjacent points , is the standard deviation of k adjacent points of the point, then the point is determined as a noise point and is removed: ; The target clustering processing is performed on the noise-removed point cloud, the Euclidean distance is used to combine adjacent points into candidate target clusters, the threshold suitable for the size of the low, slow and small target is set, the clusters meeting the volume requirement are retained, and the objects that are too large or too small are excluded; Step 22: holder cooperative control; The azimuth angle and the pitch angle of the two-degree-of-freedom holder are controlled in real time according to the three-dimensional point cloud position of the target detected by the laser radar, so that the target is always located at the center of the field of view of the camera, and stable target imaging is provided for subsequent image recognition; Step 23: dynamic focusing of the variable focus camera; Calculating target distance according to radar point cloud And calculating ideal focal length according to imaging requirement, ensuring resolution of target in image, dynamic correction, and relevant focal length calculation: The height of the unmanned aerial vehicle is about 0.5 m, and the required focal length is calculated through the focal length calculation: ; where h is the sensor height, H is the actual height of the target, which is the height of the UAV in this case, is the distance of the target measured by the radar; The parameters corresponding to the current focal length f, i.e., the internal parameter matrix K(f), are called from the pre-stored calibration lookup table; if f does not exist in the pre-stored points, linear interpolation is used for calculation; Step 24: real-time optimization; Point cloud downsampling, using voxel filtering to reduce the density of point cloud, reduce the amount of calculation; The traditional serial process has high cumulative delay, the design of the gimbal control and camera focusing is executed in parallel, the target position information, distance, timestamp is transmitted through shared memory, and the ring buffer is used to avoid read-write conflict; Through hardware acceleration, the point cloud algorithm is deployed on FPGA, so that the delay is <10ms.

6. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 1, characterized in that, In step 3, the identification of low and slow small targets includes the following steps: Step 31: In the image feature extraction module, a ResNet-50 network is used to embed a conditional parameterized convolution in each residual block; A FPN network is used to introduce a focal length aware feature selection option, input the fusion weight of each scale feature, and strengthen the P5 layer, i.e. high semantic feature, when zooming in, and focus on the P2 layer, i.e. detail feature, when zooming out; Step 32: In the radar point cloud feature extraction module, first voxelize the point cloud and convert it into a voxel grid, use VoxelNet to generate a Bev feature map, and extract the height, density, and reflectivity features of the target, and then splice them with the learned features; Step 33: Perform deformable BEV projection; Step 34: set the radar BEV feature as , the image BEV feature as , and perform cross-modal attention fusion calculation through a learnable weight matrix; Step 35: Real-time optimization of small targets by improving the detection head.

7. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 6, characterized in that, In step 31, the focal length aware feature selection option is a multi-scale feature fusion mechanism based on dynamic weight distribution, and the network structure is a conditional parameterized convolution, which is represented as: ; where, is a focal length dependent weight, generated by a light-weight MLP: ; preset expert convolution kernel; The fusion weight input FPN network uses a dynamic feature fusion algorithm, which injects dynamic weights when performing multi-scale feature fusion after the FPN feature is constructed, and applies weights in an element-by-element multiplication manner; When the system is in long focus state (f>100mm), the P5 layer is reinforced in the following way: weight assignment, MLP output of weight increase 0.6-0.8; When the system is in wide-angle state (f<50mm), the P2 layer is reinforced in the following way: MLP output of weight increase 0.5-0.

7.

8. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 6, characterized in that, In step 33, the deformable BEV projection includes: For each BEV grid point (x, y), predict its sampling offset in the image according to scale factor S : s; Where z is the depth provided by the radar, and f is the current focal length; Compensation of dynamic external parameters, integration of calibration parameters in view transformation corrected projection coordinates: ; wherein is the initial depth value of the corrected point in the camera coordinate system; In step 34, the cross-modal attention fusion calculation includes: through learnable weight matrices generate Q, K, V matrices, where Q= , K= , V= , F L is a lidar point cloud feature, F I is an image feature, d k is the dimension of the key vector K: ; B(f) is generated by a light-weight MLP from the current focal length f, for adjusting the modality weights, are learnable scaling parameters: ; At a long distance, i.e. zooming in, increase the image detail weight; At a short distance, i.e. zooming out, increase the weight of the point cloud feature.

9. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 6, characterized in that, In step 35, the Anchor dynamic size is designed, the focal length is 20-50mm, the Anchor size is 1.0m x 1.0m; the focal length is 50-100mm, the Anchor size is 0.5m x 0.5m; the focal length is 100-200mm, the Anchor size is 0.3 x 0.3m; The point cloud voxelization and BEV generation are deployed on FPGA, using a parallel architecture that processes point cloud data once per clock cycle, resulting in a delay of <5ms, the expert convolution kernel of CondConv is preloaded into NPU cache, and only the weight coefficient π(f) needs to be updated when the focal length changes, the communication volume is <1KB; The model is lightened, the channels are pruned, and L1 regularization is added to the convolution layer during the training phase; During the evaluation phase, channels with weights less than a threshold are removed, thereby achieving real-time optimization.

10. The radar and variable focal camera based low, slow and small target detection and recognition method according to claim 1, characterized in that, In step 4, the implementation of successful target strike includes the following steps: S41: Establish a strike decision module; According to the category confidence obtained by the target identification module and the distance from the gimbal, the threat of the target is evaluated: ; wherein, is a category confidence, 1 is a weight coefficient for the category confidence, 1 is a weight coefficient for the distance penalty term, is a weight coefficient for the speed threat term; = D - D0 D is a distance parameter; = Speed threat standardization; Threshold judgment is performed, when the threat index S is greater than 0.7, the following strike steps are performed; S42: Laser emission control; Gimbal accurate aiming, according to the predicted position adjustment pitch / azimuth, start laser emitter, gimbal control algorithm, motion prediction results: , = communication delay + gimbal response time; a is the acceleration derived from the Kalman filter, x, y, z are the three-dimensional spatial coordinates at the current time, and v is the current speed; When the laser first strike did not achieve the target destruction, the system automatically triggers parameter adjustment, gimbal aiming point according to the offset dynamic compensation; S43: strike effect evaluation; UAV end detection, including laser receiving device, visual verification device; Laser receiving device: photodiode array 16 x 160; trigger threshold: > 500 W , 100 ms Visual verification device: UAV end installation smoke alarm device, if hit successfully, trigger smoke system; S44: real-time guarantee; Decision FPGA acceleration, using parallel score calculation, SIMD architecture, single cycle threat score, reduce the delay of calculation, real-time control loop, gimbal correction using correction double loop control, communication optimization, low delay data transmission, protocol stack optimization, data compression, strike instruction using difference coding.

Citation Information

Patent Citations

  • Target detection and identification device and method based on multi-fusion sensor

    CN110428008A

  • Low, small and slow target identification and positioning method and system based on mixed vision

    CN115035470A

  • Multi-sensor self-adaptive tug autonomous accompanying awareness method and device

    CN118778055A

  • Low-slow small target identification method based on radar and visual data fusion

    CN119516495A

  • Systems and methods for camera-lidar fused object detection

    WO2022086739A2