Moving small target detection method and system based on infrared image
By combining AI detection model with traditional optical flow field algorithms, the problems of low detection rate and high error detection rate of drone moving small targets in complex backgrounds are solved, and high precision and stable small target detection are achieved.
Patent Information
- Application Number
- CN202510374850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-22
AI Technical Summary
The existing drone motion small-object detection technology has low detection rate and high error detection rate in complex backgrounds, making it difficult to adapt to multi-scenario applications, especially insensitive to small targets.
The detection method based on infrared images is adopted, combined with AI detection model and traditional optical flow field algorithm, and the YOLO model is trained through self-built data sets to carry out target feature learning and noise distribution analysis, and the data correlation is used by Hungarian algorithm, and the target motion attribute judgment is used for lightweight CNN model, and the weight allocation strategy is dynamically adjusted for decision-making.
It realizes high-precision detection and stable tracking of small targets in complex contexts, adapts to the optimal detection decisions in different scenarios, and improves the accuracy and robustness of detection.
Smart Images

Figure CN120355980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target automatic detection, and particularly relates to a method and system for detecting moving small targets based on infrared images. Background Art
[0002] Drones themselves have great value. For example, the aerial photography characteristics of drones: drones collect the situation of towns and map the future development trends of planned towns; drones collect data such as Gobi mountains, which brings great convenience to the placement and deployment of new energy; the transportation characteristics of drones: drone express delivery; in dangerous operations such as places that are difficult for modern transportation to reach, drone transportation has also become an important part. Drones also pose many threats, including drones being used as drug smuggling tools, intruding drones causing them to explode or using them to steal sensitive information, and drones may also be used by spies mixed in public places at key locations to carry explosives or conduct illegal surveillance. These events pose a huge threat to security and privacy.
[0003] Drones are widely used due to their advantages such as small size, low cost, convenient use, low requirements for the combat environment, and strong battlefield survival ability. The traditional target detection algorithms designed based on these characteristics are mainly divided into target activity analysis, target characteristic analysis, and the combination of both. However, this design has great limitations, such as too much human intervention, resulting in weak generalization ability in different scenarios. In recent years, with the rapid development of AI technology, based on big data, it can learn target features, so as to find all the targets (objects) of interest in the image and determine their categories and positions. One of the technical difficulties is that it is very difficult to obtain a dataset with sufficient data volume, complete categories, and all scenarios. At the same time, the commonly used model designs on the market are not sensitive to small targets, with low detection rate and high false detection rate. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the existing automatic detection technology for moving small targets, and provide a method and system for detecting moving small targets based on infrared images.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A method for detecting moving small targets based on infrared images, the steps include:
[0007] S1. Training and prediction of the AI detection model: Construct a customized dataset containing multiple scenarios and multiple sizes, train the YOLO deep learning model to learn target features and noise distribution; input the preprocessed infrared image, and output the target position, category, and confidence through the trained model;
[0008] S2. Traditional moving target detection and data association: Use optical flow field or modeling algorithm to detect the target's subtle displacement, extract motion features and determine the target position; construct a cost matrix based on the IOU of adjacent frame detection results, use the Hungarian algorithm to match the target trajectory, and assign a unique ID;
[0009] S3. AI attribute judgment model training and prediction: Build a lightweight dataset with multiple frames of continuous input and train a model with few parameters; input multiple frame image sequences and output target motion attributes, including speed and direction;
[0010] S4. Decision mechanism fusion results: Normalize the confidence of AI detection results and traditional trajectories;
[0011] According to the scene weight distribution strategy, the target with the highest confidence is selected as the final detection result, and the location, ID and attribute information are output.
[0012] Furthermore, the dataset construction in step 1 includes: covering scenes with simple to complex backgrounds; and the target size distribution is biased towards the proportion of small targets in practical applications.
[0013] Furthermore, the data association in step 2 includes: calculating the IOU distance between the current frame and the historical frame detection frame; and completing the trajectory integration by minimizing the matching cost through the Hungarian algorithm.
[0014] Furthermore, the attribute judgment model in step 3 is inputted as a two-dimensional image sequence after dimensionality reduction, and background interference is reduced by adjusting the number of channels.
[0015] Furthermore, the weight allocation strategy in step 4 is dynamically adjusted according to the complexity of the scene, with complex backgrounds focusing on AI models and dynamic scenes focusing on traditional algorithms.
[0016] Furthermore, when training the AI detection model, the self-built data set needs to cover the transition from simple background to complex background, and select the target size distribution that is biased towards the actual application scenario.
[0017] Furthermore, the input of the AI attribute judgment model is a multi-frame continuous image sequence, which realizes the time series analysis of the target motion attributes by reducing the three-dimensional spatiotemporal data into a two-dimensional image and adjusting the number of channels, including using the stacked time dimension as the channel.
[0018] Furthermore, the data association module constructs a cost matrix by calculating the IOU of adjacent frame detection boxes, and uses the Hungarian algorithm to optimize the solution to minimize the matching cost and complete the trajectory integration of the historical target and the current target.
[0019] Further, the policy decision-making module outputs the final detection result according to the decision rule that meets the actual requirements of the project by performing score normalization and weight assignment on the AI detection result and the traditional detection result.
[0020] A moving small target detection system based on infrared images is provided, and the system is used for detecting moving small targets.
[0021] The beneficial effects of the present invention are as follows:
[0022] (1) Through the fusion of AI and traditional algorithms, specifically through YOLO detection and optical flow field analysis, high-precision detection and stable tracking of small targets in complex backgrounds are achieved;
[0023] (2) Adopting a dynamic weight assignment strategy to achieve optimal detection decisions in different scenarios such as complex backgrounds and dynamic environments;
[0024] (3) Based on the analysis of multi-frame time series and lightweight CNN models, background interference filtering and accurate extraction of the speed and direction of the target motion attributes are achieved. Description of the Drawings
[0025] Figure 1 It is the overall flowchart of the moving small target detection method based on infrared images;
[0026] Figure 2 It is the inference flowchart of the AI detection model;
[0027] Figure 3 It is the processing flowchart of the AI attribute judgment model;
[0028] Figure 4 It is the data association and trajectory matching flowchart;
[0029] Figure 5 It is the fusion logic diagram of the policy decision-making module. Specific Embodiments
[0030] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] A detection method based on AI inference combined with traditional moving small targets is provided. Detection is completed through the AI detection algorithm. At the same time, traditional moving target detection and data association obtain moving targets with unique ID numbers. The detection results of the two branches are given target motion attributes. Policy decision-making is performed according to the results of the two branches after the attributes are given, and the final detection result of the tiny target is obtained according to different weight methods.
[0032] First, preprocess the collected infrared images, perform AI detection and traditional algorithm analysis, obtain results based on the analysis features of the two algorithms, send the analysis results into the AI attribute model for judgment respectively, perform attribute bundling on the detected targets, and use the decision-making mechanism to select the targets with higher credibility, so as to complete the small target detection.
[0033] To achieve the above purpose, the technical solutions adopted are as follows:
[0034] The detection method based on the combination of AI and traditional algorithms includes the following four steps:
[0035] AI detection model training and prediction;
[0036] Traditional moving target detection and data association;
[0037] AI attribute judgment model training;
[0038] Prediction and decision-making mechanism.
[0039] The function of AI detection model training and prediction is: in actual situations, since the drone is at an altitude of hundreds of meters or even higher, the most significant impact of this high-altitude environment is that the target signal becomes relatively weak, and at the same time, there are many types of drones. AI model training lies in building its own dataset, which can construct a customized dataset according to specific requirements to improve the accuracy and adaptability of the model. The purpose is to let the model learn the target features and data noise distribution, so as to fully model the target and learn more features beneficial to distinguishing the target. After the model has fully learned, save the model for the deployment of the AI detection model and small target prediction.
[0040] The function of traditional moving target detection and data association is: use traditional algorithms to detect low-altitude and slow-speed drones. Whether it is based on the optical flow field or modeling methods, the motion features in the image can present specific patterns in the target motion trajectory. These algorithms can keenly respond to these subtle changes, so as to determine the position and contour of the drone in the image, and then obtain the target detection result. Data association is actually a process of connecting signals from the same object at different times along the time axis. Data association is usually carried out before state estimation. Only by obtaining accurate data association processing results can the correctness of subsequent processing be guaranteed.
[0041] The flight of the drone is a continuous process. It analyzes various information such as the position, speed, and appearance characteristics of the target, and performs data association on the detected positions according to the analysis results to ensure that each target is stably connected during future monitoring and is assigned a unique ID number. This unique ID number assigned by data association is of extremely crucial significance for long-term tracking, target behavior analysis, and distinguishing different drones in complex environments. It enables the entire monitoring system to clearly grasp the dynamics of each drone and provides an accurate basis for subsequent decision-making and response measures.
[0042] The functions of AI attribute judgment model training and prediction are as follows: The environment and form of the drone are changeable. It may be at a clear high altitude, in a cloudy layer, or in a more complex mountain background. Whether it is AI inference or traditional detection, there is a possibility of false detection because the target is small enough, and this phenomenon will lead to too many selectable results. The function of this is to assign target motion attributes to these results. If it is known in advance that the target is in the flight process, then this attribute can express the current state of the target, thereby narrowing the selectable range and increasing the possibility of making the correct choice.
[0043] The function of the decision-making mechanism is to comprehensively consider the AI inference results and the results of traditional detection data association, and select the final detection result according to different situations.
[0044] Embodiment 1
[0045] Refer to Figure 1 , a detection method based on AI inference combined with traditional small moving targets. The implementation process of this method is as Figure 1 shown, and mainly includes the following steps: AI detection model training and prediction, AI attribute judgment model training and prediction, moving target detection, and policy decision-making. The target signal of a tiny drone belongs to weak signal detection and is easily interfered with or submerged by noise. It can be determined whether to perform data preprocessing on the image according to actual needs. In this process, AI training and verification can fully learn the target characteristics, thereby obtaining a very high accuracy. Traditional moving target detection algorithms can keenly respond to these subtle changes in target displacement, so as to determine the position and contour of the drone in the image. Therefore, this method can be given the learning function of deep learning and can also obtain the keen perception of traditional algorithms. The combination of the two can give full play to the ability of small target detection.
[0046] Refer to Figure 2 , specifically to Figure 1 which includes AI model inference. The specific implementation process of this process is as Figure 2 shown.
[0047] Model inference can use the mainstream YOLO framework on the market for object detection. When selecting a model, it is necessary to choose a suitable model version according to the background requirements. For the detection of small targets, many novel ideas have been proposed by scholars and practices have been carried out. When inputting data, it is necessary to select a suitable data augmentation strategy according to the actual situation to improve the generalization and accuracy of the model in scene applications. Another aspect that needs attention is the establishment of the test set. Build a self-built dataset according to the actual demand scenario. In terms of the scene: try to cover the transformation from a simple background to a complex background; in terms of the target size: it is necessary to select a suitable size distribution, which is more biased towards the size of the target in the actual application scenario.
[0048] See Figure 3 , specifically in Figure 1 which contains an AI attribute judgment model, and the specific process is as Figure 3 shown. This model is a small model with few parameters and takes less time. The thing to note about this module is how to build a self-built dataset. Fully consider the model input size according to the actual demand. The full map size is definitely not suitable for a small model. The input of this module is multi-frame continuous input. Because three-dimensional deployment is more difficult and dimensionality reduction is required for two-dimensional processing, the model input can work on the number of channels. Due to the construction of the dataset, this module also has the characteristics of being insensitive to the central region and small displacements, so it can greatly avoid the possibility of misjudgment caused by small displacements such as cloud movement. Thus, it can improve the judgment of the target motion attributes by learning background changes more.
[0049] See Figure 4 , specifically in Figure 1 which contains a data association module, and the specific process of this module is as Figure 4 shown. The purpose of data association is to associate the same target to form a complete motion trajectory information, and at the same time to assign a unique ID number to this target. When assigning the detection results to existing targets (previously detected targets), first calculate the assignment cost matrix, that is, the IOU distance between each current frame detection result and the bounding box of the historical target. The assignment uses the Hungarian algorithm for optimization. To achieve a match, certain criteria are required. The criterion based on by the Hungarian algorithm is "minimum loss". The loss is represented in the form of a loss matrix, and the loss matrix describes the cost of matching two elements in two sets. Therefore, the task of the Hungarian algorithm is to match the boxes in the t-th frame with the boxes in the (t - 1)-th frame pairwise, so as to complete the integration of the historical target and the current target into a complete motion trajectory.
[0050] See Figure 5 , specifically in Figure 1 which contains a policy decision-making module, and the specific process of this module is as Figure 5 shown. This module is also a decision-level fusion and can be implemented according to the actual needs of the project.
[0051] Example 2
[0052] AI Detection Model Training and Prediction: Optimize the training of deep learning models such as YOLO by building a customized dataset. First, collect an infrared image dataset containing small unmanned aerial vehicles (UAVs), covering simple to complex background scenarios such as clear high altitude, cloudy layers, and mountains, and ensure that the target size distribution is mainly based on the proportion of small targets in actual applications (such as wingspan of 0.5 - 2 meters). Data augmentation strategies include random rotation (±15°), brightness adjustment (±20%), Gaussian blur (σ = 1.5), etc., to improve the generalization ability of the model. When training, use the YOLOv5s model with an input image size of 640×640 to learn the target features and noise distribution. In the inference stage, input the pre - processed infrared image (denoised and contrast - enhanced), and the model outputs the target position (bounding box), category (UAV), and confidence level (0 - 1).
[0053] Traditional Moving Target Detection and Data Association: Achieve target trajectory matching through the combination of the optical flow field algorithm and the Hungarian algorithm. First, use the LK algorithm based on the optical flow field to detect the small displacements between adjacent frames, extract the target contour, and determine the position. Subsequently, calculate the Intersection over Union (IOU) value between the detection box in the current frame and the detection box in the historical frame to construct a cost matrix, and apply the Hungarian algorithm for optimal solution to minimize the matching cost, associate the historical trajectory of the same target with the current detection result, assign a unique ID (such as "UAV_001") to each target, and generate a continuous motion trajectory. This process ensures the stable tracking of the target in a complex environment and provides temporal information for subsequent decision - making.
[0054] The Hungarian algorithm is used in the data association module of this detection method. By calculating the Intersection over Union (IOU) of the detection boxes in adjacent frames, a cost matrix is constructed. In this detection method for small moving targets based on infrared images, the Intersection over Union (IOU) is a key metric. In the data association link, it is used to measure the similarity between the detection boxes in adjacent frames. Specifically, the IOU value is obtained by calculating the ratio of the intersection area to the union area of the detection box in the current frame and the detection box in the historical frame. The higher this value, the greater the overlap degree of the two detection boxes, indicating a higher possibility that they represent the same target. Using the IOU value as the basis for constructing the cost matrix, the Hungarian algorithm performs target matching based on this cost matrix, thereby achieving accurate association and tracking of the target trajectory, ensuring the continuous recognition of the target in different frames. Solve the global optimal matching based on the "minimum loss" criterion, associate the detection box in the current frame with the historical trajectory, assign a unique ID to the target, and update the motion trajectory. Its advantage lies in ensuring the global optimal solution, effectively coping with occlusion and noise interference, ensuring the continuity and accuracy of target tracking, and providing core support for subsequent behavior analysis and target discrimination in complex scenarios.
[0055] AI Attribute Judgment Model Training and Prediction: Temporal analysis of the target motion attributes is achieved through a lightweight CNN model. First, a lightweight dataset with continuous multi-frame input is constructed. The three-dimensional spatio-temporal data (time × height × width) is reduced to two-dimensional images (number of channels = 5, stacking the time dimension), and the input size is 128×128×5. The dataset design includes interferences such as cloud movement and small target displacements to train the model to learn the background change patterns. Lightweight models such as MobileNetV3 are used, and a sequence of 5 consecutive frames of images is input, and the target motion attributes are output, including speed (m / s) and direction (0 - 360°). This model reduces the computational time through a low-parameter design (such as depthwise separable convolution), and at the same time reduces background interference by adjusting the number of channels.
[0056] The lightweight dataset is mainly used for the training of the AI attribute judgment model. In this detection scheme, considering the requirements of practical applications for computing resources and detection speed, a lightweight dataset is constructed. This dataset adopts the method of continuous multi-frame input, and through dimensionality reduction processing, the three-dimensional spatio-temporal data is converted into two-dimensional images, and at the same time, the number of channels is adjusted to reduce the data volume. Although the dataset is optimized in terms of scale and complexity, it covers various situations that may occur in the actual scenario, such as small target movements and slow background changes. By using the lightweight dataset to train a low-parameter model, it not only ensures that the model can learn the key features of the target motion attributes, but also reduces the computational cost, improves the running efficiency of the model, and enables the entire detection system to work efficiently under resource constraints.
[0057] Continuous multi-frame input is applied to the training and prediction processes of the AI attribute judgment model. In this small moving target detection scheme, in order to more accurately analyze the motion attributes of the target, such as speed and direction, a sequence of continuous multi-frame infrared images is used as the input of the model. Since the information provided by a single frame of image is limited and it is difficult to fully reflect the motion state of the target, continuous multi-frame input can capture information such as the position change of the target over a period of time. At the same time, for ease of processing, the three-dimensional spatio-temporal data (time, height, width) is reduced to two-dimensional images, and by adjusting the number of channels, for example, stacking the time dimension as a channel, temporal analysis of the target motion attributes is achieved. This method helps the model learn the motion patterns of the target, reduces background interference, and improves the accuracy of judging the target motion attributes.
[0058] Decision-making mechanism fusion result: The detection results of two types are fused through a dynamic weight allocation strategy. First, the confidence levels of AI detection (0-1) and traditional trajectory confidence (calculated based on trajectory continuity, 0-1) are normalized to the range of [0, 100]. The weight allocation strategy is dynamically adjusted according to the scenario: in complex backgrounds (such as mountains), the weight of the AI model accounts for 70%, and the traditional algorithm accounts for 30%; in dynamic scenarios (such as strong winds), the weight of the traditional algorithm accounts for 60%, and the AI model accounts for 40%. The formula for the comprehensive confidence level is score=(AI_score×weight)+(traditional_score×weight). The target with the highest comprehensive score is selected as the final result, and its position, ID, speed, and direction information are output. This mechanism improves the detection accuracy through multi-modal fusion and adapts to the requirements of different application scenarios.
[0059] The dynamic weight allocation strategy is the core content in the steps of the decision-making mechanism fusion result. In this small moving target detection scheme, due to the respective advantages and disadvantages of the AI detection model and the traditional moving target detection method in different scenarios, a dynamic weight allocation strategy is adopted. For complex background scenarios, such as when there are interference factors like mountains and clouds, the AI model is given a higher weight due to its strong feature learning ability to improve the accuracy of target detection; while in dynamic scenarios, such as when the target moves rapidly or is affected by strong winds, the sensitivity of the traditional algorithm to motion features increases its weight. After normalizing the confidence levels of the AI detection results and the traditional trajectories, appropriate weights are allocated to the two methods according to the characteristics of the scenario, and finally the confidence levels of both are combined to select the final detection result, enabling the detection method to adapt to a variety of different practical application scenarios.
[0060] The solution described in the embodiment can be widely applied to the automatic detection of small drones in the security field. Customized collected data is used for model training and verification to obtain a detection model and an attribute judgment model. The detection results of the AI branch are obtained using the obtained AI detection model and attribute judgment model. At the same time, the detection results of this branch are obtained using the traditional algorithm detection and attribute judgment model. The results obtained from the two branches are used as inputs and sent into the decision-making mechanism, and finally the detection results of small moving targets are obtained.
[0061] By combining the feature learning ability of the AI model with the motion sensitivity of the traditional algorithm, the detection accuracy in complex scenarios is improved; the weight allocation strategy is automatically adjusted according to the scenario to optimize the detection performance in different environments; the attribute judgment model filters out interferences such as cloud movement through time series analysis to enhance robustness; and the combination of YOLOv5s and the lightweight attribute model takes into account both detection speed and resource consumption.
[0062] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, and should not be regarded as excluding other embodiments. Instead, it can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concepts described herein through the above teachings or the techniques or knowledge in related fields. As long as the changes and variations made by those skilled in the art do not depart from the spirit and scope of the present invention, they should all be within the protection scope of the appended claims of the present invention.
Claims
1. A method for detecting small moving targets based on infrared images, characterized in that, The method steps include: S1. AI detection model training and prediction: Build a customized data set with multiple scenes and sizes, train the YOLO deep learning model, learn target features and noise distribution; input pre-processed infrared images, and output the target location, category and confidence through the trained model; S2. Traditional moving target detection and data association: Use optical flow field or modeling algorithm to detect the target's subtle displacement, extract motion features and determine the target position; construct a cost matrix based on the IOU of adjacent frame detection results, use the Hungarian algorithm to match the target trajectory, and assign a unique ID; S3. AI attribute judgment model training and prediction: Build a lightweight dataset with multiple frames of continuous input and train a model with few parameters; input multiple frame image sequences and output target motion attributes, including speed and direction; S4. Decision mechanism fusion results: Normalize the confidence of AI detection results and traditional trajectories; According to the scene weight distribution strategy, the target with the highest confidence is selected as the final detection result, and the location, ID and attribute information are output.
2. The method for detecting a small moving target based on an infrared image according to claim 1, wherein The dataset construction in step 1 includes: covering scenes with simple to complex backgrounds; and the target size distribution is biased towards the proportion of small targets in practical applications.
3. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The data association in step 2 includes: calculating the IOU distance between the current frame and the historical frame detection frame; minimizing the matching cost through the Hungarian algorithm to complete the trajectory integration.
4. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The attribute judgment model input in step 3 is a two-dimensional image sequence after dimensionality reduction, and background interference is reduced by adjusting the number of channels.
5. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The weight allocation strategy in step 4 is dynamically adjusted according to the complexity of the scene. Complex backgrounds focus on AI models, while dynamic scenes focus on traditional algorithms.
6. The method for detecting small moving targets based on infrared images according to claim 1, wherein When training the AI detection model, the self-built data set needs to cover the transition from simple background to complex background, and select the target size distribution that is biased towards the actual application scenario.
7. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The input of the AI attribute judgment model is a multi-frame continuous image sequence. By reducing the three-dimensional spatiotemporal data into a two-dimensional image and adjusting the number of channels, including using the stacked time dimension as the channel, the temporal analysis of the target motion attributes is achieved.
8. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The data association module constructs a cost matrix by calculating the IOU of adjacent frame detection boxes, and uses the Hungarian algorithm to optimize the solution to minimize the matching cost and complete the trajectory integration of the historical target and the current target.
9. A method for detecting small moving targets based on infrared images according to claim 1, characterized in that, The strategy decision module normalizes the scores and assigns weights to the AI detection results and traditional detection results, and outputs the final detection results according to the decision rules that meet the actual needs of the project.
10. A moving small target detection system based on infrared images, characterized in that, The system is used for detecting small moving targets, and the system uses a small moving target detection method based on infrared images as described in claims 1 to 9.
Citation Information
Cited By
Conveyor belt wine bottle opening positioning method and system combining single light source and optical flow algorithm
CN121600238A
Infrared small target detection method based on adaptive multi-scale filtering
CN122336425A