Method for detecting complex background infrared dark and weak moving target based on motion perception

By constructing a DAMF-Net network, the background motion and target motion difference of multi-frame images are compensated and enhanced, and the detection problem of space-based infrared detection systems is solved under complex dynamic backgrounds, and high-precision detection of weak infrared targets is achieved.

CN120388044APending Publication Date: 2025-07-29SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437603.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing space-based infrared detection system is difficult to effectively detect weak infrared targets under complex dynamic backgrounds. Traditional algorithms are prone to false alarms when the background and targets move at the same time, and deep learning algorithms have limited effects in single-frame detection.

Method used

A DAMF-Net network is constructed, including a motion background compensation module, a multi-frame difference enhancement module, a residual feature extraction module and a multi-attention-guided feature fusion module, which uses the background motion and target motion differences of multi-frame images to improve detection accuracy through feature fusion.

Benefits of technology

It realizes high-precision detection of weak infrared targets under complex dynamic backgrounds, effectively suppresses background clutter, improves detection rate and accuracy rate, and enhances feature characterization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388044A_ABST
    Figure CN120388044A_ABST
Patent Text Reader

Abstract

The invention relates to a method for detecting a complex background infrared dim moving target based on motion perception, which comprises the following steps of: firstly, constructing a complex dynamic background infrared dim moving target data set, and then constructing an overall detection framework which specifically comprises a motion background compensation module, a multi-frame difference enhancement module, a feature extraction module and a feature fusion module, and finally, training, testing and evaluating the model. On the basis of a space-based infrared detection system, under the condition that dynamic clutters are similar to targets and infrared weak and small targets cannot be accurately detected, the background is compensated by utilizing the difference between background motion and target motion, the influence of the motion on target detection is eliminated, the weak and small targets are enhanced through multi-frame difference enhancement, and the target detection accuracy is improved. And finally, realizing feature adaptive fusion by using a feature fusion module. According to the method, time-space domain information is combined, an end-to-end infrared weak and small moving target detection network is provided, and better detection performance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and target detection, and relates to a detection method for infrared dim moving targets with complex backgrounds based on motion perception. Background Art

[0002] Space-based infrared detection systems play a significant role in space exploration and early warning due to their strong concealment and long-distance detection capabilities. However, the space-based detection platform is often more than 300 km away from the target. Due to the long distance from the detection target, complex background clutter, and atmospheric radiation interference, the target has characteristics such as low gray level, no shape, and small size, and is easily submerged by clutter. In complex ground backgrounds, it is difficult to distinguish the background and the target only by gray difference. Therefore, existing algorithms have begun to consider multi-frame joint detection of infrared small targets. However, in the case where both the background and the target are moving, it is difficult to effectively suppress the moving background only by general multi-frame joint detection, and a large number of false alarms are easily generated. Traditional algorithms are often ineffective in the face of this type of complex situation. Although deep learning algorithms can adapt to background changes and target size changes, most current algorithms focus on single-frame detection and are aimed at small targets with high signal-to-noise ratios. Therefore, it is still a difficult point to effectively detect infrared small targets in complex dynamic backgrounds.

[0003] Under space-based infrared detection systems, the target has the characteristics of being weak, small, and textureless, and is easily submerged in complex background clutter. It is difficult to distinguish the target and the background only by gray information. General deep learning algorithms that use multi-frame information to achieve target detection use the motion characteristics of the target to distinguish the target and the background. However, when the background is also dynamic, relying solely on this motion to distinguish is not enough. Summary of the Invention

[0004] The purpose of the present invention is to provide a detection method for infrared dim moving targets with complex backgrounds based on motion perception, which overcomes the problem of insufficient detection ability of existing algorithms for infrared dim moving targets in complex dynamic scenarios.

[0005] To achieve the above purpose, the technical solution of the present invention is:

[0006] A detection method for infrared dim moving targets with complex backgrounds based on motion perception includes the following steps:

[0007] S1, constructing a multi-frame infrared dim moving target image dataset with complex backgrounds;

[0008] S2, dividing the dataset into a training set and a test set;

[0009] S3. Construct a complex background infrared dim moving target detection network DAMF-Net based on motion perception in complex dynamic scenarios. The DAMF-Net network consists of an input end, a motion background compensation module, a multi-frame difference enhancement module, a residual feature extraction module, a multi-attention guided feature fusion module, and an output end.

[0010] S4. Use the dataset constructed in step S1 to train the DAMF-Net model according to the training set segmented in S2.

[0011] S5. Input the dataset constructed in step S1 and the test set segmented in S2 into the DAMF-Net after being trained in step S4 to test the detection performance of the DAMF-Net.

[0012] S6. Evaluate the detection effect of the DAMF-Net network on low signal-to-noise ratio infrared dim moving targets.

[0013] In step S1, use the real-shot dataset to construct A infrared dim moving target image sequences. Each image has a resolution of R1×R2 and a channel number of C, and label the infrared dim targets on each frame of the image. The labeling format is mask.

[0014] In step S2, segment the A infrared dim moving target image sequences constructed in step S1. Use B infrared dim moving target image sequences and their corresponding labels as the training set X, and use the remaining A - B infrared dim moving target image sequences and their corresponding labels as the validation set V.

[0015] In step S3, the steps to construct DAMF-Net are as follows:

[0016] S301. Sample the infrared dim moving target image sequences by the method of sliding window sampling, select the current frame and the N frames before and after it to ensure the temporal continuity of the input images to the DAMF-Net.

[0017] S302. The input end reads the infrared dim moving target image sequences in time series, reads in the current frame and its adjacent N frames before and after to ensure the temporal order of the input images, and performs preprocessing such as randomly cropping, randomly enhancing, or unifying the input size on the input images.

[0018] S303. The motion background compensation module consists of a convolutional layer, local similarity calculation, and a Softmax layer, and uses the input multi-frame image sequences to calculate the dynamic background similarity.

[0019] S304. The multi-frame differential enhancement module consists of multiple feature mapping modules and differential operations. The multi-frame images after multi-background compensation are subjected to feature mapping registration operations, and then the current frame feature layer is differentially operated using the front and back frames to achieve dynamic background suppression and target enhancement.

[0020] S305. The residual feature extraction module consists of multiple residual modules to obtain feature layers of different scales. The feature layers of different scales respectively enter the background compensation module in S303 and the differential enhancement module in S304 to enhance the feature layers.

[0021] S306. The multi-attention-guided feature fusion module is divided into two parts. The first part increases the spatial feature correlation through the group attention gating mechanism EAG, focusing on the spatial key information. The second part separates the strong and weak features through the hierarchical feature refinement HAR, further refines the strong features, and enhances the weak features.

[0022] S307. The feature layers of different scales obtained from the network constructed according to steps S303, S304, and S305 are fused by gradually upsampling and using the multi-attention-guided feature fusion module in S306 in sequence to obtain the final segmentation prediction map.

[0023] S308. Calculate the training loss by comparing the prediction map obtained in step S307 with the corresponding ground truth map obtained during S1 data collection. The loss function uses the HPM loss, and the specific calculation is as follows:

[0024]

[0025] Among them, b p represents the number of positive samples, and α ∈ [0, 1] is the first type of weight factor. By adjusting the ratio of different types, the imbalance between positive and negative samples can be alleviated.

[0026] The object detection loss function in S308 can also use SoftIou loss, or FocalLoss, or OHEMloss.

[0027] In step S4, the training of the DAMF-Net model includes the following steps:

[0028] S401. Start training the DAMF-Net model constructed in step S3 from 0.

[0029] S402. Retain the best weights obtained in the training in step S401 for model verification and evaluation.

[0030] In step S5, the detection performance of the DAMF-Net is tested, including the following steps:

[0031] S501. Input the best weights obtained after training in step S4 into the test data set segmented in step S2, and perform prediction on the data to obtain the target prediction result;

[0032] S502. In the prediction result, determine the result with a threshold greater than the set threshold T1 as the correct result.

[0033] In step S6, the evaluation of the DAMF-Net network's detection effect on complex dynamic dim infrared targets includes the following steps:

[0034] S601. Use the detection accuracy rate to evaluate the network's ability to find the right objects, and the calculation formula is as follows:

[0035]

[0036] S602. Use the false alarm rate to evaluate the network's ability to detect all targets, and the calculation formula is as follows:

[0037]

[0038] where N target represents the number of targets predicted as positive, N all represents the total number of all targets, N f represents the number of positive class targets misjudged as negative class targets. H represents the image length, W represents the image width, and m represents the number of test images.

[0039] The advantages of the present invention are as follows: 1. The present invention utilizes the differences between the background motion and the target motion in multiple frames of images to compensate for the dynamic background, and then performs differential enhancement on the multiple frames of images after background compensation, so as to enhance the target and suppress the background. Finally, the final feature fusion is performed through a multi-attention-guided feature fusion module to achieve high-precision detection of infrared small and weak targets in complex dynamic scenes; 2. It can effectively enhance the target and suppress the background by utilizing the differences between the background and the target motion in multiple frames of data, and can make full use of spatio-temporal features to improve the detection performance of the network for infrared small and weak targets in complex dynamic scenes; 3. In view of the situation that in the complex dynamic background of space-based infrared images for ground targets, it is difficult to distinguish moving targets from dynamic clutter, and it is difficult to effectively detect infrared small and weak targets only relying on simple target motion information, a motion background compensation module and a multi-frame differential enhancement module are proposed. The motion background compensation module uses multiple frames of data to register and compensate the dynamic background, suppressing the dynamic clutter similar to the target motion generated by the background motion, while the multi-frame differential enhancement module uses multiple frames of data for differential enhancement, effectively enhancing small and weak targets and suppressing dynamic clutter. The two enhance the target and suppress the background by utilizing the differences between the background motion and the target motion, and the combination of the two can be carried on any single-frame detection network to improve the detection rate of small and weak targets; 4. A multi-attention-guided feature fusion module is constructed to perform grouped attention enhancement and hierarchical strong and weak feature refinement on the feature layer. By separating the strong and weak features, the weak features are enhanced through the spatial attention mechanism, the strong features are refined, and then fully fused. This module can achieve adaptive fusion of multi-scale strong and weak features, effectively enhancing the feature representation ability of the target and improving the detection rate of infrared small and weak targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is the step flowchart of the detection method of the present invention;

[0041] Figure 2 is the network structure diagram of the motion background compensation module of the present invention;

[0042] Figure 3 is the network structure diagram of the motion multi-frame differential enhancement module of the present invention;

[0043] Figure 4 is the overall network structure diagram of the present invention;

[0044] Figure 5 is the multi-attention-guided feature fusion module of the present invention;

[0045] Figure 6 is the detection effect diagram of infrared dim and weak targets in complex dynamic background in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0046] The present invention will be further described below with reference to the accompanying drawings. The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent.

[0047] To more concisely describe this embodiment, some components that are well-known to those skilled in the art but not relevant to the main content of this creation will be omitted in the drawings or descriptions. Additionally, for ease of expression, some components in the drawings will be omitted, enlarged, or reduced, but this does not represent the size or entire structure of the actual product.

[0048] The present invention discloses a method for detecting infrared dim moving targets with complex backgrounds based on motion perception, as Figure 1 shown, including the following steps:

[0049] S1, construct a multi-frame infrared dim moving target image dataset with a complex background; use the actual shooting dataset to construct A infrared dim moving target image sequences, each with an image resolution of R1×R2 and a channel number of C, and label the infrared dim targets on each frame of the image, with the labeling format being mask.

[0050] In this embodiment, A = 100, R1 = 256, R2 = 320, and C = 3.

[0051] S2, divide the dataset into a training set and a test set; divide the A infrared dim moving target image sequences constructed in step S1, use B infrared dim moving target image sequences and their corresponding labels as the training set X, and use the remaining A - B infrared dim moving target image sequences and corresponding labels as the validation set V.

[0052] In this embodiment, A = 100 and B = 70.

[0053] S3, construct a detection network DAMF-Net for infrared dim moving targets with complex backgrounds based on motion perception in complex dynamic scenes; as Figure 4 shown, the DAMF-Net network consists of an input end, a motion background compensation module, a multi-frame differential enhancement module, a residual feature extraction module, a multi-attention guided feature fusion module, and an output end;

[0054] The steps for constructing DAMF-Net are as follows:

[0055] S301, sample the infrared dim moving target image sequences by the method of sliding window sampling, select the current frame and the N frames before and after, to ensure the temporal continuity of the input images to DAMF-Net; in this embodiment, N = 1.

[0056] S302. The input end reads the infrared dim moving target image sequence in time series, reads in the current frame and its adjacent N frames before and after to ensure the time series of the input images, and performs preprocessing such as randomly cropping, randomly enhancing, or unifying the input size on the input images.

[0057] S303. The motion background compensation module, as Figure 2 shown, consists of a convolutional layer, local similarity calculation, and a Softmax layer, and uses the input multi-frame image sequence to calculate the dynamic background similarity.

[0058] The motion background compensation module consists of an ordinary two-dimensional convolutional layer, a local similarity calculation module, and a Softmax layer. For the two input images, after extracting features through the convolutional layer, the images are grouped into local regions with a size of L×L, and the local region similarity is calculated using matrix multiplication and Softmax to calculate the dynamic background similarity.

[0059] S304. The multi-frame difference enhancement module consists of multiple feature mapping modules and difference operations. The multi-frame images after multi-background compensation are subjected to feature mapping registration operations, and then the current frame feature layer is subjected to difference operations using the front and back frames to achieve dynamic background suppression and target enhancement.

[0060] As Figure 3 shown, in this embodiment, the multi-frame difference enhancement module consists of 2 feature mapping modules, namely the Warp module and the feature layer difference operation. The background compensation features obtained through background similarity calculation are registered through Warp to achieve the registration of dynamic features, and then the current frame is subjected to difference operations using the front and back frames. In this way, the dynamic background clutter is suppressed through dynamic background compensation, and the background is further suppressed and the target is enhanced through two difference operations.

[0061] S305. The residual feature extraction module consists of multiple residual modules, obtains feature layers of different scales, and the feature layers of different scales respectively enter the background compensation module of S303 and the difference enhancement module of S304 to enhance the feature layers.

[0062] S306. The multi-attention-guided feature fusion module is divided into two parts, as Figure 5 shown. The first part increases the spatial feature correlation through the grouped attention gating mechanism EAG, focusing on the spatial key information. The second part separates the strong and weak features through the hierarchical feature refinement HAR, further refines the strong features, and enhances the weak features, which can significantly improve the model's ability to distinguish between targets and backgrounds and improve the network's ability to extract features of small targets.

[0063] S307. The feature layers of different scales obtained from the network constructed according to steps S303, S304, and S305 are fused by gradually upsampling and using the multi-attention-guided feature fusion module in S306 to obtain the final segmentation prediction map.

[0064] S308. Calculate the training loss by comparing the prediction map obtained in step S307 with the corresponding ground truth map obtained during S1 data collection. The loss function uses HPM loss, and the specific calculation is as follows:

[0065]

[0066] Among them, b p represents the number of positive samples, and α ∈ [0, 1] is the first type of weight factor. By adjusting the ratio of different types, the imbalance between positive and negative samples can be alleviated.

[0067] In S308, the object detection loss function can also use SoftIou loss, or FocalLoss, or OHEMloss.

[0068] S4. Use the dataset constructed in step S1 to train the DAMF-Net model according to the training set segmented in S2. The steps for training the DAMF-Net model are as follows:

[0069] S401. Train the DAMF-Net model constructed in step S3. The training environment is NVIDIA GeForce RTX3090 GPU, the network training is Epoch = 100, the initial learning rate Lr = 0.05, the learning rate optimizer is SGD, Batchsize = 4, and start training from 0.

[0070] S402. Retain the best weights obtained from the training in step S401 for model validation and evaluation.

[0071] S5. Input the test set segmented in S2 of the dataset constructed in step S1 into the DAMF-Net after being trained in step S4 to test the detection performance of the DAMF-Net. The steps for testing the detection performance of the DAMF-Net are as follows:

[0072] S501. Input the test dataset segmented in step S2 into the best weights obtained after training in step S4 to predict the data and obtain the target prediction results.

[0073] S502. In the prediction results, determine the results with a threshold greater than the set threshold T1 as the correct results.

[0074] S6. Evaluate the detection effect of the DAMF-Net network on low signal-to-noise ratio infrared dim moving targets; the evaluation of the detection effect of the DAMF-Net network on complex dynamic background infrared dim targets includes the following steps:

[0075] S601. Use the detection accuracy rate to evaluate the precision ability of the network. The calculation formula is as follows:

[0076]

[0077] S602. Use the false alarm rate to evaluate the ability of the network to detect all targets. The calculation formula is as follows:

[0078]

[0079] Among them, N target represents the number of targets predicted as positive, N all represents the total number of all targets, N f represents the number of positive class targets misjudged as negative class targets. H represents the image length, W represents the image width, and m represents the number of test images.

[0080] The partial experimental effects of this embodiment on real infrared image sequences are as Figure 6 shown.

[0081] In summary, the above are only the preferred embodiments of the present invention, and are not used to limit the scope of implementation of the present invention. That is, all equivalent changes and modifications made according to the content of the patent application scope of the present invention should fall within the technical scope of the present invention.

Claims

1. A detection method for infrared dim moving targets in complex backgrounds based on motion perception, characterized in that, It includes the following steps: S1. Construct a multi-frame infrared dim moving target image dataset with complex backgrounds; S2. Split the dataset into a training set and a test set; S3. Construct a complex background infrared dim moving target detection network DAMF-Net based on motion perception for complex dynamic scenarios; The DAMF-Net network consists of an input end, a motion background compensation module, a multi-frame difference enhancement module, a residual feature extraction module, a multi-attention guided feature fusion module, and an output end; S4. Use the dataset constructed in step S1 to train the DAMF-Net model according to the training set split in S2; S5. Input the dataset constructed in step S1 into the DAMF-Net trained in step S4 according to the test set split in S2 to test the detection performance of DAMF-Net; S6. Evaluate the detection effect of the DAMF-Net network on low signal-to-noise ratio infrared dim moving targets.

2. The detection method according to claim 1, characterized in that: In step S1, A infrared dim moving target image sequences are constructed using a real-shot dataset. Each image has a resolution of R1×R2 and a channel number of C, and the infrared dim targets on each frame of the image are labeled with a mask annotation format.

3. The detection method according to claim 1, characterized in that: In step S2, the A infrared dim moving target image sequences constructed in step S1 are split. B infrared dim moving target image sequences and their corresponding labels are used as the training set X, and the remaining A - B infrared dim moving target image sequences and corresponding labels are used as the validation set V.

4. The detection method according to claim 1, characterized in that: In step S3, the steps to construct DAMF-Net are as follows: S301. Sample the infrared dim moving target image sequences by means of sliding window sampling, select the current frame and the N frames before and after it to ensure the temporal continuity of the input images to DAMF-Net; S302. The input end reads the infrared dim moving target image sequences in time series, reads in the current frame and its adjacent N frames before and after to ensure the temporal order of the input images, and performs preprocessing such as randomly cropping, randomly augmenting, or unifying the input size on the input images; S303. The motion background compensation module consists of a convolutional layer, local similarity calculation, and a Softmax layer, and calculates the dynamic background similarity using the input multi-frame image sequences; S304. The multi-frame difference enhancement module consists of multiple feature mapping modules and difference operations. Feature mapping registration operations are performed on the multi-frame images after multi-background compensation, and then the current frame feature layer is differentially operated using the front and back frames to achieve dynamic background suppression and target enhancement; S305. The residual feature extraction module consists of multiple residual modules, obtains feature layers of different scales, and the feature layers of different scales enter the background compensation module in S303 and the difference enhancement module in S304 respectively to enhance the feature layers; S306. The multi-attention guided feature fusion module is divided into two parts. The first part increases the spatial feature correlation through a grouped attention gating mechanism EAG, focusing on spatial key information. The second part separates strong and weak features through hierarchical feature refinement HAR, further refines the strong features, and enhances the weak features; S307. The feature layers of different scales obtained from the network constructed according to steps S303, S304, and S305 are fused by gradually upsampling and using the multi-attention-guided feature fusion module of S306 in sequence to obtain the final segmentation prediction map. S308. Calculate the training loss by computing the prediction map obtained in step S307 and the corresponding ground truth map obtained during S1 data collection. The loss function uses HPM loss, and the specific calculation is as follows: HPM(p t ) = ∑-αW(1 - p t ) γ log(p t ) Among them, b p represents the number of positive samples, α ∈ [0, 1] is the first type of weight factor, and by adjusting the proportions of different types, the imbalance between positive and negative samples can be alleviated.

5. The detection method according to claim 4, characterized in that: In S308, the object detection loss function can also use SoftIou loss, or FocalLoss, or OHEMloss.

6. The detection method according to claim 1, wherein: In step S4, the training of the DAMF-Net model includes the following steps: S401. Train the DAMF-Net model constructed in step S3 from scratch. S402. Retain the best weights obtained in the training of step S401 for the verification and evaluation of the model.

7. The detection method according to claim 1, characterized in that: In step S5, the detection performance of the DAMF-Net test includes the following steps: S501. Input the segmented test data set in step S2 into the best weights obtained after training in step S4 to predict the data and obtain the target prediction result. S502. In the prediction result, determine the result with a threshold greater than the set threshold T1 as the correct result.

8. The detection method according to claim 1, wherein: In step S6, the evaluation of the DAMF-Net network's detection effect on complex dynamic background infrared dim targets includes the following steps: S601. Use the detection accuracy rate to evaluate the network's ability to find accurate targets. The calculation formula is as follows: S602. Use the false alarm rate to evaluate the network's ability to detect all targets. The calculation formula is as follows: Among them, N target represents the number of positive target predictions, and N all represents the total number of all targets, and N f represents the number of positive-class targets misjudged as negative-class targets. H represents the image length, W represents the image width, and m represents the number of test images.