A high-precision turntable cooperative multi-modal low, slow and small target positioning and identification system

CN122836722APending Publication Date: 2026-09-29LUOYANG AIR ROUTE ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611093695.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

首先,飞行高度低,容易混杂于地物背景(建筑物、山体、树林等)之中,雷达和光电探测均受到强烈的背景杂波干扰;其次,飞行速度慢且运动轨迹多变(悬停、急停、变速机动),传统的基于多普勒频率的动目标检测方法易将其与慢速地物杂波混淆;再者,雷达散射截面小,回波信噪比极低,常规雷达难以稳定建航;同时,在远距离或恶劣天气条件下,目标在光电图像中仅占据几个至几十个像素,缺乏充足的形状、纹理等判别性特征,导致视觉识别极为困难

Benefits of technology

1、本发明通过协同控制模块的目标引导单元与动态跟踪单元,先基于雷达探测得到的可疑目标空间坐标初始估计,驱动转台快速粗指向目标区域;当目标进入视野中心后,利用脱靶量闭环反馈结合卡尔曼滤波轨迹预测与PID控制算法,生成精确的伺服补偿指令,实现从粗指向到精细跟踪的无缝衔接。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122836722A_ABST
    Figure CN122836722A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of low-altitude security detection, and discloses a high-precision rotating table cooperative multi-modal low, slow and small target positioning and identifying system. The application calculates the spatial coordinate initial estimation of a suspicious target by collecting the radar echo signal of the target airspace, and simultaneously sends the spatial coordinate initial estimation to a cooperative control module. A multi-waveband photoelectric imaging module comprises a high-precision two-degree-of-freedom servo rotating table and a visible light imaging sensor and an infrared thermal imaging sensor. The servo rotating table is used for real-time tracking of the suspicious target, and the imaging sensor is used for real-time collection of image data in the field of view. The cooperative control module calculates the rotating table pointing deviation based on the spatial coordinate initial estimation of the suspicious target, generates compensation instructions based on the real-time image sequence, and sends the compensation instructions to the servo rotating table to control the rotating table to turn. A multi-modal fusion processing module is used for feature fusion of the infrared thermal image and the visible light image. A target identifying module identifies the low, slow and small target based on the multi-modal fusion feature map obtained through fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of low-altitude security detection technology, specifically a high-precision turntable-coordinated multimodal low-altitude, slow-moving, small target positioning and identification system. Background Technology

[0002] Low-altitude, slow-moving, and small targets typically refer to targets with low flight altitudes (generally below 1000 meters), slow speeds (generally less than 50 meters per second), and small radar cross-sections (generally less than 2 square meters). Typical examples include small drones, birds, powered parachutes, and hot air balloons. With the rapid popularization of unmanned aerial vehicle (UAV) technology, low-altitude, slow-moving, and small UAVs are widely used in aerial photography, logistics, and inspection. However, frequent security incidents such as unauthorized flights, illegal intrusion into sensitive areas, and carrying dangerous goods pose serious challenges to public safety, national defense security, and civil aviation safety. Therefore, developing efficient and reliable detection, positioning, and identification systems for low-altitude, slow-moving, and small targets has become a critical technical issue that urgently needs to be addressed in the field of low-altitude security.

[0003] The inherent characteristics of low-altitude, slow-moving, and small targets make their detection far more difficult than that of conventional aerial targets. First, their low flight altitude makes them easily blend into the background of ground features (buildings, mountains, forests, etc.), resulting in strong background clutter interference for both radar and electro-optical detection. Second, their slow flight speed and varied trajectories (hovering, sudden stops, and variable-speed maneuvers) make them easy targets for traditional Doppler frequency-based moving target detection methods easily confused with slow-moving ground clutter. Third, their small radar cross-section and extremely low echo signal-to-noise ratio make it difficult for conventional radars to establish stable navigation. At the same time, at long distances or in adverse weather conditions, the target occupies only a few to tens of pixels in the electro-optical image, lacking sufficient discriminative features such as shape and texture, making visual identification extremely difficult. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a high-precision turntable-coordinated multimodal low-slow-small target positioning and identification system. It features the ability to collect multidimensional data from a servo turntable, send it to a collaborative control module, and then generate control commands to control the turntable's direction. This allows the turntable to track suspicious targets within its field of view in real time during the monitoring period, and simultaneously determine whether the current suspicious target is a low-slow-small target. This system solves the aforementioned technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a high-precision turntable-coordinated multimodal low-speed small target positioning and recognition system, comprising a radar detection module, a multi-band photoelectric imaging module, a cooperative control module, a multimodal fusion processing module, and a target recognition module; The radar detection module is used to perform an all-round scan of the target airspace, acquire radar echo signals, process the radar echo signals to obtain the range, azimuth, radial velocity and radar cross section information of the suspected target, generate an initial estimate of the spatial coordinates of the suspected target, and send it to the cooperative control module. The multi-band optoelectronic imaging module includes a high-precision two-degree-of-freedom servo turntable and a visible light imaging sensor and an infrared thermal imaging sensor mounted on the high-precision two-degree-of-freedom servo turntable. The multi-band optoelectronic imaging module points to the direction of the suspected target based on the turntable pointing control, and at the same time collects real-time image sequences within the field of view. The collaborative control module is communicatively connected to the radar detection module and the high-precision two-degree-of-freedom servo turntable. The collaborative control module includes a target guidance unit and a dynamic tracking unit. The target guidance unit calculates the turntable pointing deviation based on the initial estimation of the spatial coordinates of the suspected target and sends it to the high-precision two-degree-of-freedom servo turntable. The dynamic tracking unit generates compensation commands based on the real-time image sequence acquired by the multi-band photoelectric imaging module and sends them to the high-precision two-degree-of-freedom servo turntable. The multimodal fusion processing module is used to perform spatiotemporal registration of the visible light image acquired by the visible light imaging sensor and the infrared thermal image acquired by the infrared thermal imaging sensor, extract visible light modal features and infrared modal features, and perform feature-level fusion of the extracted multimodal features to obtain a multimodal fusion feature map. The target recognition module is used to input the multimodal fusion feature map into a pre-trained deep learning target detection and recognition model, perform target detection and classification recognition on low, slow and small targets, output target category information and target confidence, and determine whether a suspicious target is a low, slow and small target.

[0006] As a preferred embodiment of the present invention, the collaborative control module includes the following execution flow: Step A1: The cooperative control module receives the initial estimate of the spatial coordinates of the suspicious target sent by the radar detection module; Step A2: Based on the initial estimation of the spatial coordinates of the suspected target, the target guidance unit calculates the azimuth angle of the target relative to the turntable using the position of the high-precision two-degree-of-freedom servo turntable as the origin and through the conversion relationship between polar coordinates and rectangular coordinates. and pitch angle Then the turning point pointing deviation can be calculated. Step A3: The target guidance unit generates a pointing drive command based on the turntable pointing deviation obtained in step A2, drives the turntable to rotate rapidly at a set rate, and at the same time ensures that the pointing deviation is reduced to within a set threshold. Step A4: When the pointing deviation decreases to within the set threshold, the multi-band photoelectric imaging module begins to acquire images within the field of view, and the dynamic tracking unit extracts the pixel position of the suspicious target in the image plane. Computation and Image Center The pixel deviation between them is used as the off-target amount; Step A5: The dynamic tracking unit obtains a motion command based on the miss distance and drives the turntable to adjust its direction, so that the target gradually moves to the center area of ​​the image. Step A6: After the target enters the continuous tracking stage, the dynamic tracking unit uses the Kalman filter algorithm to predict the target's trajectory based on the target's miss distance sequence in multiple consecutive frames of images, and obtains the predicted value of the target position in the next frame. Step A7: The dynamic tracking unit calculates the compensation command using a PID control algorithm based on the current miss distance and the prediction result; Step A8: The dynamic tracking unit sends the compensation command to the turntable, driving the turntable to track the target in real time.

[0007] As a preferred embodiment of the present invention, the expression for the turntable pointing deviation in step A2 is as follows: ; ; in, This refers to the azimuth deviation; This refers to pitch angle deviation; The target azimuth angle; Current azimuth angle of the turntable; Pitch angle for the target; This is the current pitch angle of the turntable.

[0008] As a preferred embodiment of the present invention, the expression for the off-target amount in step A4 is as follows: ; ; in, This represents the horizontal miss distance. This refers to the vertical miss distance. The pixel coordinates of the target in the image coordinate system; These are the pixel coordinates of the image center.

[0009] As a preferred embodiment of the present invention, the expression of the compensation instruction in step A7 is as follows: ; in, Indicates the first Frame compensation instructions; Indicates the first The off-target distance of a frame; This refers to the inter-frame time interval. Indicates the first The off-target distance of a frame; This indicates the frame number corresponding to the current image; This represents the summation operation; This is the proportional gain coefficient; This is the integral gain coefficient; This is the differential gain coefficient.

[0010] As a preferred embodiment of the present invention, the expression for the fusion weights involved in the feature-level fusion of the extracted multimodal features is as follows: ; in, Indicates the visible light feature at position The fusion weight at the location; Indicates the visible light feature map at position The characteristic response intensity at that location; For infrared feature map at location The characteristic response intensity at that location; This is the Sigmoid function.

[0011] As a preferred embodiment of the present invention, the expression of the multimodal fusion feature map is as follows: ; in, For multimodal fusion feature maps at location The eigenvector at that location; For visible light feature maps at location The eigenvector at that location; For infrared feature map at location The eigenvector at that location; Indicates the visible light feature at position The fusion weight at the location.

[0012] As a preferred embodiment of the present invention, the total loss function of the deep learning object detection and recognition model is as follows: ; in, This represents the total loss value. For bounding box regression loss; For classification loss; For confidence loss; , and This is the balance coefficient for all losses.

[0013] As a preferred technical solution of the present invention, the training set used by the deep learning target detection and recognition model contains a large number of low-speed small target images collected under different lighting conditions, different weather conditions, and different background types.

[0014] As a preferred technical solution of the present invention, the step of determining whether a suspicious target is a low, slow and small target specifically means: when the target confidence level is higher than a set threshold, the current target is a low, slow and small target, and the precise position of the target in the image coordinate system is output.

[0015] Compared with existing technologies, this invention provides a high-precision turntable-coordinated multimodal low-speed small target localization and recognition system, which has the following beneficial effects: 1. This invention uses the target guidance unit and dynamic tracking unit of the collaborative control module to first estimate the spatial coordinates of the suspicious target obtained by radar detection, and drive the turntable to quickly and coarsely point to the target area; when the target enters the center of the field of view, it uses the closed-loop feedback of the miss amount combined with Kalman filter trajectory prediction and PID control algorithm to generate precise servo compensation commands, so as to achieve a seamless connection from coarse pointing to fine tracking.

[0016] 2. This invention employs a weighted fusion strategy based on differences in feature response intensity to calculate the fusion weights of visible light and infrared features pixel-by-pixel, generating a multimodal fusion feature map. This fusion method can dynamically adjust the contribution ratio of visible light and infrared modes according to scene lighting, weather, occlusion, and other conditions. It prioritizes the use of visible light texture information in bright light areas and automatically switches to infrared thermal radiation features in low light or smoky environments, significantly improving the feature representation capability and robustness of small, slow-moving targets in complex environments.

[0017] 3. This invention utilizes a lightweight YOLO architecture and designs a multi-task weighted loss function specifically for small, slow-moving targets. By balancing the weights of bounding box regression, classification, and confidence losses, the model can simultaneously optimize both localization accuracy and classification accuracy. Furthermore, the training set covers typical small, slow-moving targets such as drones, birds, and paragliders under various lighting, weather, and background conditions, and data augmentation is employed to expand sample diversity, effectively reducing the false negative and false alarm rates for small targets. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the system framework of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 A high-precision turntable-coordinated multimodal low-speed small target localization and recognition system includes: a radar detection module, a multi-band photoelectric imaging module, a cooperative control module, a multimodal fusion processing module, and a target recognition module; The radar detection module is used to perform an all-round scan of the target airspace, acquire radar echo signals, process the radar echo signals to obtain the range, azimuth, radial velocity and radar cross section information of the suspected target, and generate an initial estimate of the spatial coordinates of the suspected target, which is then sent to the cooperative control module. The multi-band optoelectronic imaging module includes a high-precision two-degree-of-freedom servo turntable and a visible light imaging sensor and an infrared thermal imaging sensor mounted on the high-precision two-degree-of-freedom servo turntable. Based on the turntable pointing control, the multi-band optoelectronic imaging module points to the direction of the suspected target and collects real-time image sequences within the field of view. The collaborative control module is communicatively connected to the radar detection module and the high-precision two-degree-of-freedom servo turntable. The collaborative control module includes a target guidance unit and a dynamic tracking unit. The target guidance unit calculates the turntable pointing deviation based on the initial estimation of the spatial coordinates of the suspected target and sends it to the high-precision two-degree-of-freedom servo turntable. The dynamic tracking unit generates compensation commands based on the real-time image sequence acquired by the multi-band photoelectric imaging module and sends them to the high-precision two-degree-of-freedom servo turntable. The collaborative control module includes the following execution flow: Step A1: The cooperative control module receives the initial estimate of the spatial coordinates of the suspicious target sent by the radar detection module; Step A2: Based on the initial estimation of the spatial coordinates of the suspected target, the target guidance unit calculates the azimuth angle of the target relative to the turntable using the position of the high-precision two-degree-of-freedom servo turntable as the origin and through the conversion relationship between polar coordinates and rectangular coordinates. and pitch angle Then, the turntable pointing deviation is calculated, and its expression is as follows: ; ; in, This refers to the azimuth deviation; This refers to pitch angle deviation; The target azimuth angle; Current azimuth angle of the turntable; Pitch angle for the target; This is the current pitch angle of the turntable; Step A3: The target guidance unit generates a pointing drive command based on the turntable pointing deviation obtained in step A2, drives the turntable to rotate rapidly at a set rate, and at the same time ensures that the pointing deviation is reduced to within a set threshold. Step A4: When the pointing deviation decreases to within the set threshold, the multi-band photoelectric imaging module begins to acquire images within the field of view, and the dynamic tracking unit extracts the pixel position of the suspicious target in the image plane. Computation and Image Center The pixel deviation between the two is used as the off-target amount, and its expression is as follows: ; ; in, This represents the horizontal miss distance. This refers to the vertical miss distance. The pixel coordinates of the target in the image coordinate system; The coordinates of the center pixel of the image; Step A5: The dynamic tracking unit obtains a motion command based on the miss distance and drives the turntable to adjust its direction, so that the target gradually moves to the center area of ​​the image. Step A6: After the target enters the continuous tracking stage, the dynamic tracking unit uses the Kalman filter algorithm to predict the target's trajectory based on the target's miss distance sequence in multiple consecutive frames of images, and obtains the predicted value of the target position in the next frame. Step A7: Based on the current miss distance and the prediction result, the dynamic tracking unit calculates the compensation command using a PID control algorithm, the expression of which is as follows: ; in, Indicates the first Frame compensation instructions; Indicates the first The off-target distance of a frame; This refers to the inter-frame time interval. Indicates the first The off-target distance of a frame; This indicates the frame number corresponding to the current image; This represents the summation operation; This is the proportional gain coefficient; This is the integral gain coefficient; The differential gain coefficient; The magnitude and direction of the values ​​determine how the turntable corrects the current pointing deviation; the proportional term makes the turntable return to center at a speed proportional to the current miss distance, and the larger the miss distance, the stronger the return; the integral term continuously accumulates the tiny residual error of each frame, and gradually increases the output when the target deviates from the center for a long time, eventually pulling the target back to the center; the differential term senses the instantaneous change trend of the miss distance, and if the target is accelerating away from the center, the differential term outputs the reverse torque in advance to avoid excessive overshoot; the three terms work together to enable the photoelectric turntable to respond quickly to sudden deviations, eliminate steady-state deviations delicately, and predict motion trends in advance, thereby firmly locking low, slow, and small targets in the center of the field of view; Step A8: The dynamic tracking unit sends compensation commands to the turntable, driving the turntable to track the target in real time; The multimodal fusion processing module is used to perform spatiotemporal registration of visible light images acquired by visible light imaging sensors and infrared thermal images acquired by infrared thermal imaging sensors, extract visible light modal features and infrared modal features, and perform feature-level fusion of the extracted multimodal features to obtain a multimodal fusion feature map. The expression for the fusion weights involved in feature-level fusion of the extracted multimodal features is as follows: ; in, Indicates the visible light feature at position The fusion weight at the location; Indicates the visible light feature map at position The characteristic response intensity at that location; For infrared feature map at location The characteristic response intensity at that location; This is the Sigmoid function; Essentially, it's a modal confidence index used to determine whether visible light or infrared information should be trusted more at each pixel location in an image. In real-world scenarios, if a local area is well-lit and has clear texture (such as the exterior wall of a building), the visible light feature response intensity... Larger Approaching 1, the system primarily utilizes the color and detail features provided by visible light; if the area is in shadow, at night, or obscured by smoke (such as a drone under a tree), the infrared signature response intensity... Larger When the value approaches zero, the system switches to relying on infrared thermal radiation features for identification. This mechanism of "using the stronger modal signal more often" is equivalent to dynamically assigning modal weights to each pixel, so that the fused feature map always retains the most discriminative information in complex environments such as strong light, weak light, fog, and occlusion, thereby significantly improving the overall robustness of low, slow and small target identification. The expression for the multimodal fusion feature map is as follows: ; in, For multimodal fusion feature maps at location The eigenvector at that location; For visible light feature maps at location The eigenvector at that location; For infrared feature map at location The eigenvector at that location; Indicates the visible light feature at position The fusion weight at the location; It is the comprehensive feature descriptor that the system ultimately uses for target recognition; it represents the location in the image. This is a unified representation combining visible light texture information and infrared thermal radiation information. Taking a drone flying at dusk as an example: in areas with moving parts such as the fuselage and propellers, the visible light image may be blurry due to insufficient light. Smaller Mainly derived from infrared features The system uses thermal signals to determine the presence of a target; in the edge region where the fuselage meets the sky, the visible light image still retains a certain outline, and the infrared image may also show edge gradients. Close to 0.5 Equivalent to the average of the two features, edge information is doubly enhanced; while in a completely dark background with no thermal difference, the intensity of infrared features is low. Approaching 1, It degenerates into visible light features (although visible light is also weak, it at least preserves the noise distribution and avoids interference from invalid features); in summary, this expression is weighted... The contribution ratio of the two modalities is dynamically adjusted so that the fused feature map can automatically select the most reliable modal information at each point, thereby providing a robust, information-rich and dimensionally compact input for subsequent target detection and recognition. The target recognition module is used to input the multimodal fusion feature map into the pre-trained deep learning target detection and recognition model, perform target detection and classification recognition on low, slow and small targets, output target category information and target confidence, and determine whether a suspicious target is a low, slow and small target. The total loss function of the deep learning object detection and recognition model is as follows: ; in, This represents the total loss value. For bounding box regression loss; For classification loss; For confidence loss; , and This is the balance coefficient for all losses; Total loss It is a scalar value that quantifies the total error between the model's current prediction and the ground truth label. If the predicted bounding box deviates from the ground truth ( If the target class is misclassified (large), the model will be forced to adjust the weights of the regression branch to reduce bias; if the target class is misclassified (large), the model will be forced to adjust the weights of the regression branch to reduce bias. (Large), the weights of the classification branches will be updated to enhance inter-class separability; if the model falsely reports a target in a background region without a target ( (For large numbers of cases), confidence branches will be suppressed to reduce false alarms; The deep learning object detection and recognition model adopts a lightweight network model based on the improved YOLO architecture, which mainly consists of three parts: the backbone network, the neck network, and the detection head. The backbone network is used to extract multi-level semantic features from the multi-modal fusion feature map. The neck network performs multi-scale fusion of multi-level semantic features based on the feature pyramid structure and the path aggregation network. The detection head is used to output the target's class probability, bounding box coordinates, and confidence score at multiple scales. The training set used by the deep learning target detection and recognition model contains a large number of low-speed small target images collected under different lighting conditions, weather conditions, and background types (urban, suburban, mountain, and water), covering a variety of typical low-speed small target types such as drones, birds, and paragliders. Data diversity is further expanded through data augmentation (random cropping, rotation, color transformation, noise addition, etc.). To determine whether a suspicious target is a low, slow, or small target, the following steps are taken: when the target confidence level is higher than a set threshold, the current target is considered a low, slow, or small target, and the precise position of the target in the image coordinate system is output.

[0021] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A high-precision turntable-coordinated multimodal low-speed small target localization and recognition system, characterized in that: It includes a radar detection module, a multi-band optoelectronic imaging module, a collaborative control module, a multi-modal fusion processing module, and a target recognition module; The radar detection module is used to perform an all-round scan of the target airspace, acquire radar echo signals, process the radar echo signals to obtain the range, azimuth, radial velocity and radar cross section information of the suspected target, generate an initial estimate of the spatial coordinates of the suspected target, and send it to the cooperative control module. The multi-band optoelectronic imaging module includes a high-precision two-degree-of-freedom servo turntable and a visible light imaging sensor and an infrared thermal imaging sensor mounted on the high-precision two-degree-of-freedom servo turntable. The multi-band optoelectronic imaging module points to the direction of the suspected target based on the turntable pointing control, and at the same time collects real-time image sequences within the field of view. The collaborative control module is communicatively connected to the radar detection module and the high-precision two-degree-of-freedom servo turntable. The collaborative control module includes a target guidance unit and a dynamic tracking unit. The target guidance unit calculates the turntable pointing deviation based on the initial estimation of the spatial coordinates of the suspected target and sends it to the high-precision two-degree-of-freedom servo turntable. The dynamic tracking unit generates compensation commands based on the real-time image sequence acquired by the multi-band photoelectric imaging module and sends them to the high-precision two-degree-of-freedom servo turntable. The multimodal fusion processing module is used to perform spatiotemporal registration of the visible light image acquired by the visible light imaging sensor and the infrared thermal image acquired by the infrared thermal imaging sensor, extract visible light modal features and infrared modal features, and perform feature-level fusion of the extracted multimodal features to obtain a multimodal fusion feature map. The target recognition module is used to input the multimodal fusion feature map into a pre-trained deep learning target detection and recognition model, perform target detection and classification recognition on low, slow and small targets, output target category information and target confidence, and determine whether a suspicious target is a low, slow and small target.

2. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The collaborative control module includes the following execution flow: Step A1: The cooperative control module receives the initial estimate of the spatial coordinates of the suspicious target sent by the radar detection module; Step A2: Based on the initial estimation of the spatial coordinates of the suspected target, the target guidance unit calculates the azimuth angle of the target relative to the turntable using the position of the high-precision two-degree-of-freedom servo turntable as the origin and through the conversion relationship between polar coordinates and rectangular coordinates. and pitch angle Then the turning point pointing deviation can be calculated. Step A3: The target guidance unit generates a pointing drive command based on the turntable pointing deviation obtained in step A2, drives the turntable to rotate rapidly at a set rate, and at the same time ensures that the pointing deviation is reduced to within a set threshold. Step A4: When the pointing deviation decreases to within the set threshold, the multi-band photoelectric imaging module begins to acquire images within the field of view, and the dynamic tracking unit extracts the pixel position of the suspicious target in the image plane. Computation and Image Center The pixel deviation between them is used as the off-target amount; Step A5: The dynamic tracking unit obtains a motion command based on the miss distance and drives the turntable to adjust its direction, so that the target gradually moves to the center area of ​​the image. Step A6: After the target enters the continuous tracking stage, the dynamic tracking unit uses the Kalman filter algorithm to predict the target's trajectory based on the target's miss distance sequence in multiple consecutive frames of images, and obtains the predicted value of the target position in the next frame. Step A7: The dynamic tracking unit calculates the compensation command using a PID control algorithm based on the current miss distance and the prediction result; Step A8: The dynamic tracking unit sends the compensation command to the turntable, driving the turntable to track the target in real time.

3. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 2, characterized in that: The expression for the turntable pointing deviation in step A2 is as follows: ; ; in, This refers to the azimuth deviation; This refers to pitch angle deviation; The target azimuth angle; Current azimuth angle of the turntable; Pitch angle for the target; This is the current pitch angle of the turntable.

4. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 2, characterized in that: The expression for the off-target amount in step A4 is as follows: ; ; in, This represents the horizontal miss distance. This refers to the vertical miss distance. The pixel coordinates of the target in the image coordinate system; These are the pixel coordinates of the image center.

5. A high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 2, characterized in that: The expression for the compensation instruction in step A7 is as follows: ; in, Indicates the first Frame compensation instructions; Indicates the first The off-target distance of a frame; This refers to the inter-frame time interval. Indicates the first The off-target distance of a frame; This indicates the frame number corresponding to the current image; This represents the summation operation; This is the proportional gain coefficient; This is the integral gain coefficient; This is the differential gain coefficient.

6. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The expression for the fusion weights involved in the feature-level fusion of the extracted multimodal features is as follows: ; in, Indicates the visible light feature at position The fusion weight at the location; Indicates the visible light feature map at position The characteristic response intensity at that location; For infrared feature map at location The characteristic response intensity at that location; This is the Sigmoid function.

7. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The expression for the multimodal fusion feature map is as follows: ; in, For multimodal fusion feature maps at location The eigenvector at that location; For visible light feature maps at location The eigenvector at that location; For infrared feature map at location The eigenvector at that location; Indicates the visible light feature at position The fusion weight at the location.

8. The high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The total loss function of the deep learning object detection and recognition model is as follows: ; in, This represents the total loss value. For bounding box regression loss; For classification loss; For confidence loss; , and This is the balance coefficient for all losses.

9. A high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The training set used by the deep learning target detection and recognition model contains a large number of low-speed, small target images collected under different lighting conditions, weather conditions, and background types.

10. A high-precision turntable-coordinated multimodal low-speed small target localization and recognition system according to claim 1, characterized in that: The determination of whether a suspicious target is a low, slow, and small target specifically involves: when the target confidence level is higher than a set threshold, the current target is a low, slow, and small target, and the precise position of the target in the image coordinate system is output.