Feature enhancement modulation infrared small target detection method
By constructing a feature-enhanced modulation infrared small target detection model, the problems of lack of deep texture features and information loss in infrared small target detection are solved, and more efficient target detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202511004111.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing infrared small target detection methods are limited in efficiency and accuracy in complex environments, lack deep texture features and suffer from severe information loss, making it difficult to effectively extract and fuse target information.
A feature-enhanced and modulated infrared small target detection model is constructed, including a feature extraction module, a feature pyramid, an adaptive feature modulation module, and a dynamic wavelet feature enhancement module. Through multi-scale feature extraction, adaptive pooling, and feature modulation processing, combined with Haar wavelet filters and multilayer perceptrons, shallow features are enhanced and deep feature textures are supplemented. Target detection is achieved using an anchorless detection head network.
It improves the accuracy and robustness of infrared small target detection, effectively solves the problems of missing texture features in deep features and loss of small target information caused by upsampling, and improves the accuracy and recall of detection.
Smart Images

Figure CN120912860A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and deep learning, and particularly relates to an infrared small target detection method based on feature enhancement modulation. BACKGROUND
[0002] Due to the unique penetration characteristics, infrared imaging technology can effectively capture the thermal energy information of the target in complex environments such as fog, haze and smoke. This advantage makes it play an irreplaceable role in key fields such as military monitoring, maritime search and rescue, and intelligent transportation. The traditional infrared small target detection is mainly through image preprocessing to suppress the interference of background noise, and then through the algorithm of machine learning to realize target positioning. Although this method is simple in calculation, low in model complexity, and the detection accuracy depends on the selection of human features. In actual complex application scenarios, the detection efficiency and accuracy of the method are limited.
[0003] With the development of deep learning technology, the infrared small target detection method based on deep learning effectively extracts image features by constructing a specific network structure, which greatly improves the detection accuracy. However, due to the lack of texture features of deep features and the loss of target information in the process of deep network feature fusion, these technical problems together constitute the key obstacles that need to be broken through in the field of infrared small target detection, and therefore it is urgent to develop new feature enhancement technology. SUMMARY
[0004] The present application provides an infrared small target detection method based on feature enhancement modulation to overcome the above technical problems.
[0005] In order to achieve the above purpose, the technical scheme of the present application is:
[0006] An infrared small target detection method based on feature enhancement modulation, comprising the steps of:
[0007] Step 1: acquiring an infrared image for small target detection;
[0008] The infrared image is sequentially subjected to size adjustment and xml format conversion, and the training set and test set are divided according to a preset ratio;
[0009] Step 2: constructing an infrared small target detection model based on feature enhancement modulation;
[0010] The infrared small target detection model comprises a feature extraction module, a feature pyramid, an adaptive feature modulation module, a dynamic wavelet feature enhancement module and a detection head network;
[0011] The feature extraction module extracts a multi-scale feature map of the infrared image, and the multi-scale feature map comprises a plurality of feature maps with gradually reduced detail information; the detail information at least includes target edge and texture features;
[0012] The fusion feature map is obtained by using multi-scale feature maps through a feature pyramid;
[0013] The adaptive feature modulation module is used to perform adaptive pooling processing on the first fusion feature map to the fourth fusion feature map, to obtain adaptive pooling feature maps of different scales, and perform feature modulation processing on the adaptive pooling feature maps to obtain modulation feature maps, and the generated modulation feature maps are fused with the corresponding fusion feature maps to obtain the first enhanced modulation feature map to the fourth enhanced modulation feature map;
[0014] The dynamic wavelet feature enhancement module is used to perform feature enhancement processing on the first layer of the multi-scale feature map using a Haar wavelet filter to obtain an enhanced shallow layer feature map, and the fifth fusion feature map is fused with the enhanced shallow layer feature map to obtain an optimized enhanced feature map; and the first layer of the multi-scale feature map is the feature map with the most detailed information; and the fifth fusion feature map is the feature map with the least detailed information;
[0015] The detection head network uses the optimized enhanced feature map and the first enhanced modulation feature map to the fourth enhanced modulation feature map to realize detection of small targets in an infrared image;
[0016] Step 3: Use the training data set and the test data set to train and test the infrared small target detection model to obtain an optimal detection model; and use the optimal detection model to determine the position of small targets in the infrared image to be detected.
[0017] Further, the dynamic wavelet feature enhancement module constructed in step 2 includes a Haar wavelet filter, a multi-layer perception, a softmax activation function layer, and a feature processing module composed of a concatenation layer, a convolution layer, a batch normalization layer, and a ReLU activation function layer connected in sequence;
[0018] The Haar wavelet filter is used to extract high-frequency feature components and low-frequency feature components in the first layer of the multi-scale feature map; and the high-frequency feature components include horizontal high-frequency features, vertical high-frequency features, and diagonal high-frequency features;
[0019] The multi-layer perception is used to perform a nonlinear mapping operation on the first layer of the multi-scale feature map;
[0020] After the softmax activation function layer is used to perform an activation operation on the output of the multi-layer perception, a set of dynamic normalized weights is obtained, and the normalized weights are evenly divided in the channel dimension to obtain four divided weights W1-W4;
[0021] The feature processing module is used to obtain an enhanced shallow layer feature map from the enhanced feature map;
[0022] And the enhanced feature map is a feature map obtained by performing an element-wise multiplication operation on the segmentation weights W1-W4 and the low-frequency feature component, the horizontal high-frequency feature, the vertical high-frequency feature, and the diagonal high-frequency feature, respectively.
[0023] Further, the adaptive feature modulation module constructed in step 2 includes a channel segmentation layer, an adaptive pooling layer, a convolution layer, a bilinear interpolation upsampling layer, and a splicing layer.
[0024] The first fusion feature map to the fourth fusion feature map output by the feature pyramid are split along the channel by the channel segmentation layer, and four independent channel segmentation feature maps are obtained.
[0025] The adaptive pooling layer uses a preset scaling factor to perform adaptive maximum pooling operation on any three channel segmentation feature maps, and obtains feature maps of different spatial sizes.
[0026] The convolution layer performs convolution operation on part of the first fusion feature map to the fourth fusion feature map output by the feature pyramid to compress the channel, and uses the channel-compressed feature map to obtain the average variance information of the feature map. The average variance information is respectively summed with the feature maps of different spatial sizes to obtain the variance feature map.
[0027] The splicing layer performs splicing operation on the deep convolution feature map to obtain the modulation feature map.
[0028] The deep convolution feature map is a feature map obtained by performing deep convolution operation on the variance feature map using a deep convolution block, and then performing bilinear interpolation operation with different factors by the bilinear interpolation upsampling layer.
[0029] Further, the method for obtaining the optimal detection model in step 3 includes the following steps:
[0030] Step 31: using the training data set to train the constructed infrared small target detection model to obtain a trained infrared small target detection model:
[0031] Step 32: based on the constructed model loss function, using the test set to evaluate the trained infrared small target detection model, and determining whether the output of the trained infrared small target detection model converges.
[0032] If yes, the trained infrared small target detection model is the optimal detection model.
[0033] Otherwise, based on the adaptive optimization algorithm, the model parameter weight of the trained infrared small target detection model is adaptively adjusted, and step 31 is repeatedly executed.
[0034] Further, the model loss function constructed in step 32 is the sum of the classification loss function and the regression loss function, and the classification loss function is measured by Quality Focal Loss, and the regression loss function is measured by GIoUloss.
[0035] Further, the feature extraction module in step 2 includes but is not limited to residual network ResNet-50.
[0036] The detection head network adopts an anchor-free detection head network, including but not limited to a CornerNet target detector.
[0037] The application provides an infrared small target detection method based on feature enhancement modulation, and has the following beneficial effects: (1) through the designed dynamic wavelet feature enhancement module, the high-frequency texture features in the feature map are extracted by using the Haar wavelet filter in the shallow feature, that is, the multi-scale feature map S1, and further, the dynamic weight information of the feature map is extracted by using the multi-layer perception, the weight information is combined with the high-frequency texture features, and then added to the deep feature, so as to make up for the problem of missing target texture features in the deep feature; (2) through the designed adaptive feature modulation module, different spatial size feature maps are generated by using adaptive pooling of different factors to operate the features after the feature pyramid, different spatial size feature information is extracted, and different feature information is combined to make the target information more prominent. The application constructs the infrared small target detection model based on feature enhancement modulation, effectively solves the problems of missing texture features in the deep feature and small target information loss caused by up-sampling in the feature pyramid, and further improves the precision of infrared small target detection. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0039] Figure 1 The application provides an infrared small target detection method based on feature enhancement modulation, and has the following beneficial effects: (1) through the designed dynamic wavelet feature enhancement module, the high-frequency texture features in the feature map are extracted by using the Haar wavelet filter in the shallow feature, that is, the multi-scale feature map S1, and further, the dynamic weight information of the feature map is extracted by using the multi-layer perception, the weight information is combined with the high-frequency texture features, and then added to the deep feature, so as to make up for the problem of missing target texture features in the deep feature; (2) through the designed adaptive feature modulation module, different spatial size feature maps are generated by using adaptive pooling of different factors to operate the features after the feature pyramid, different spatial size feature information is extracted, and different feature information is combined to make the target information more prominent. The application constructs the infrared small target detection model based on feature enhancement modulation, effectively solves the problems of missing texture features in the deep feature and small target information loss caused by up-sampling in the feature pyramid, and further improves the precision of infrared small target detection.
[0040] Figure 2 The application provides an infrared small target detection method based on feature enhancement modulation, and has the following beneficial effects: (1) through the designed dynamic wavelet feature enhancement module, the high-frequency texture features in the feature map are extracted by using the Haar wavelet filter in the shallow feature, that is, the multi-scale feature map S1, and further, the dynamic weight information of the feature map is extracted by using the multi-layer perception, the weight information is combined with the high-frequency texture features, and then added to the deep feature, so as to make up for the problem of missing target texture features in the deep feature; (2) through the designed adaptive feature modulation module, different spatial size feature maps are generated by using adaptive pooling of different factors to operate the features after the feature pyramid, different spatial size feature information is extracted, and different feature information is combined to make the target information more prominent. The application constructs the infrared small target detection model based on feature enhancement modulation, effectively solves the problems of missing texture features in the deep feature and small target information loss caused by up-sampling in the feature pyramid, and further improves the precision of infrared small target detection.
[0041] Figure 3 The application provides an infrared small target detection method based on feature enhancement modulation, and has the following beneficial effects: (1) through the designed dynamic wavelet feature enhancement module, the high-frequency texture features in the feature map are extracted by using the Haar wavelet filter in the shallow feature, that is, the multi-scale feature map S1, and further, the dynamic weight information of the feature map is extracted by using the multi-layer perception, the weight information is combined with the high-frequency texture features, and then added to the deep feature, so as to make up for the problem of missing target texture features in the deep feature; (2) through the designed adaptive feature modulation module, different spatial size feature maps are generated by using adaptive pooling of different factors to operate the features after the feature pyramid, different spatial size feature information is extracted, and different feature information is combined to make the target information more prominent. The application constructs the infrared small target detection model based on feature enhancement modulation, effectively solves the problems of missing texture features in the deep feature and small target information loss caused by up-sampling in the feature pyramid, and further improves the precision of infrared small target detection.
[0042] Figure 4A network structure schematic diagram of the adaptive feature modulation module of the present application;
[0043] Figure 5 A result schematic diagram of infrared small target detection on SIRST data set by applying the present application;
[0044] Figure 6 A result schematic diagram of infrared small target detection on ISATD by applying the present application. DETAILED DESCRIPTION
[0045] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in a clear and complete manner with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0046] The present embodiment provides a feature-enhanced modulation infrared small target detection method, as shown in Figure 1 The method specifically comprises the following steps:
[0047] Step 1: acquiring an infrared image for small target detection;
[0048] The infrared images in the present embodiment are from two public infrared data sets SIRST and ISATD, wherein the infrared images in the infrared data sets are labeled with a boundary box, and the size of each infrared image in the two data sets is adjusted to 512*512 pixels, and the label file is converted into xml format, and then the infrared images are divided into a training set and a test set in a ratio of 3:1;
[0049] Step 2: constructing a feature-enhanced modulation infrared small target detection model;
[0050] The infrared small target detection model comprises a feature extraction module, a feature pyramid, an adaptive feature modulation module, a dynamic wavelet feature enhancement module and a detection head network.
[0051] Step 200: extracting a multi-scale feature map of the infrared image through the feature extraction module, wherein the multi-scale feature map comprises a plurality of feature maps S1-S4 with gradually reduced detail information; the detail information at least comprises target edge and texture features; and the multi-scale feature map S1 with the most detail information is defined as the first layer of the multi-scale feature map.
[0052] Specifically, as shown in Figure 2As shown, for the input image, a feature extraction can be performed by using an arbitrary deep backbone network, and in this embodiment, a residual network ResNet-50 is taken as an example for illustration. The residual network ResNet-50 can preliminarily extract a multi-scale feature map from the input image, and only the first layer S1 of the multi-scale feature map is taken as the input original feature of the dynamic wavelet feature enhancement module, while the multi-scale feature maps S1-S4 are taken as the input of the feature pyramid to further extract features.
[0053] Step 201: Obtain fusion feature maps P2-P6 by using the multi-scale feature maps S1-S4 through the feature pyramid, and after obtaining the fusion feature maps P2-P5, input them into the respective corresponding adaptive feature modulation modules
[0054] Specifically, the feature pyramid network is a computer vision technology for improving the performance of object detection and image segmentation. The feature pyramid network constructs feature levels from bottom to top and fuses features of different levels through horizontal connection to improve detection performance. Its core advantage lies in the ability to balance high resolution and strong semantic information, which is particularly important for small target detection, and the versatility and efficiency of the feature pyramid network make it a standard module of many modern detection frameworks. The feature pyramid extracts features at different image scales to capture object information of different sizes and resolutions. The differences between the feature maps output by the feature pyramid are as follows: (1) Resolution: Different levels of the feature pyramid represent information of different scales. Generally, high-level feature maps contain more semantic information; while low-level feature maps have higher resolution. This is because high-level feature maps have undergone more downsampling operations and capture more abstract features, while low-level feature maps retain more detailed information. (2) Semantic information: High-level feature maps can better recognize and locate objects, while low-level feature maps are more suitable for capturing local information such as object edges and textures. In summary, the feature pyramid can represent features at multiple scales, enabling the model to effectively detect and recognize objects at different scales, which helps to improve the model's ability to detect objects of different sizes, and by fusing multi-scale information, the model's robustness and accuracy are enhanced;
[0055] Step 202: The adaptive feature modulation module is used to perform adaptive pooling processing on the fusion feature maps P2-P5 (first fusion feature map to fourth fusion feature map), and a modulation feature map is obtained by using deep convolution and bilinear interpolation up-sampling, and the modulation feature map is fused with the corresponding fusion feature map P2-P5 to obtain the corresponding enhanced modulation feature map, i.e., the first enhanced modulation feature map to the fourth enhanced modulation feature map; in this example, the adaptive feature modulation module uses adaptive pooling in the module to process the fusion features to obtain feature maps of different scales, and target information is extracted from the feature maps of different scales, thereby improving the network's ability to capture target information in the feature map.
[0056] In specific embodiments, the adaptive feature modulation module includes a channel segmentation layer, an adaptive pooling layer, a convolution layer, a bilinear interpolation up-sampling layer, and a splicing layer.
[0057] The fusion feature maps P2-P5 output by the feature pyramid are split along the channel by the channel segmentation layer to obtain four independent channel segmentation feature maps.
[0058] The adaptive pooling layer uses a preset scaling factor to perform adaptive maximum pooling on any three channel segmentation feature maps to obtain feature maps of different spatial sizes.
[0059] The convolution block performs convolution on the fusion feature maps P2-P5 output by the feature pyramid to compress the channels, and the average variance information of the feature maps is obtained using the compressed feature maps; the element-wise sum operation is performed on the average variance information and the feature maps of different spatial sizes to obtain a variance feature map; the splicing layer is used to splice the deep convolution feature maps to obtain a modulation feature map; and the deep convolution feature map is obtained by performing deep convolution on the variance feature map using the deep convolution block, and then performing bilinear interpolation operation with different factors using the bilinear interpolation up-sampling layer.
[0060] In this embodiment, the adaptive feature modulation module is inserted after the feature pyramid, i.e., the adaptive feature modulation module is used to obtain feature maps of different spatial sizes by performing adaptive pooling operation with different factors, and the target information in these feature maps is extracted to improve the network's ability to capture target information. Specifically, the adaptive feature modulation module is as follows Figure 4As shown, after the fusion of the feature maps P2-P5, an adaptive feature modulation module is inserted, and the original features are split into four independent parts along the channel by using a channel segmentation technique and configured with channel coding. Among them, the feature map of the first channel is kept unchanged in space size, while the feature maps of the second, third and fourth channels are respectively subjected to adaptive max-pooling operations with scaling factors of 2, 4 and 8, thereby generating feature maps of different spatial sizes. Meanwhile, in order to better capture the target information in the feature map, the original features are compressed in channels through a 1*1 convolution operation, and then the variance information of the feature map is obtained. The method for obtaining the variance information of the feature map is a known technique, and will not be described in detail here. The variance information is fused into the feature maps of the second, third and fourth channels, which helps to highlight the feature information of the target region. Subsequently, the feature maps of the second, third and fourth channels are respectively subjected to 3*3 deep separable convolution processing, and different factors of bilinear interpolation upsampling are adopted to unify the dimensions of the features considering the size difference of the features in each channel. At the same time, after the feature map of the first channel is subjected to 3*3 deep separable convolution processing, it is combined with the feature maps of the second, third and fourth channels described above to generate the final output feature, i.e., the modulation feature map.
[0061] Step 203: Through the dynamic wavelet feature enhancement module, only the first layer S1 of the multi-scale feature map is subjected to feature enhancement processing by using the Haar wavelet filter to obtain an enhanced shallow layer feature map, and the fifth fusion feature map is fused with the enhanced shallow layer feature map to obtain an optimized enhanced feature map; and the first layer S1 of the multi-scale feature map is the feature map with the most detailed information; and the fifth fusion feature map P6 is the fusion feature map with the least detailed information.
[0062] In specific embodiments, the constructed dynamic wavelet feature enhancement module includes a Haar wavelet filter, a multi-layer perception, a softmax activation function layer, and a feature processing module composed of a concatenation layer, a convolution layer, a batch normalization layer, and a ReLU activation function layer connected in sequence; in this embodiment, the high-frequency information of the first layer S1 of the multi-scale feature map is enhanced by using the Haar wavelet filter, and the enhanced features are fused into the deep features to compensate for the lack of texture features in the deep features, which specifically includes
[0063] The high-frequency feature components and low-frequency feature components in the first layer S1 of the multi-scale feature map are extracted by the Haar wavelet filter; and the high-frequency feature components include horizontal high-frequency features, vertical high-frequency features, and diagonal high-frequency features;
[0064] The multi-layer perception performs a non-linear mapping operation on the first layer S1 of the multi-scale feature map;
[0065] After performing an activation operation on the output of the multi-layer perceptron through a softmax activation function layer, a set of dynamic normalized weights is obtained, and the normalized weights are evenly divided in the channel dimension and four segmentation weights W1-W4 are obtained;
[0066] The feature processing module uses the enhanced feature map to obtain the enhanced shallow layer feature map; and the enhanced feature map is a feature map obtained by performing an element-by-element multiplication operation on the segmentation weights W1-W4 and the low-frequency feature component, the horizontal high-frequency feature, the vertical high-frequency feature, and the diagonal high-frequency feature, respectively.
[0067] The dynamic wavelet feature enhancement module in the embodiment is as shown in Figure 3 The output feature map S1 of the residual network ResNet-50 is used as the input original feature of the module. First, the original feature is filtered by the Haar wavelet filter to generate four different high-frequency components and low-frequency components from the horizontal and vertical directions, that is, the diagonal high-frequency information, the vertical high-frequency information, the horizontal high-frequency information, and the low-frequency information. The method of obtaining the high-frequency component and the low-frequency component from the original feature through the Haar wavelet filter is a known technology, and will not be described in detail. After the original feature is nonlinearly mapped by the multi-layer perceptron, a set of dynamic normalized weight information is generated by the Softmax activation function. Then, the dynamic weight information is evenly divided into four weight information, that is, the segmentation weights W1-W4, according to the number of channel dimensions. For example, the weight information generated by the multi-layer perceptron from the original feature is in the form of 1024 weights of all channels in four directions, which is then converted into the form of 4*256 (each part is 256 channels, and the normalized weights are evenly divided according to the direction dimension), and the four weight information is multiplied with the high-frequency and low-frequency components generated by the Haar wavelet filter to realize dynamic enhancement of the feature. Finally, the high-frequency information and the low-frequency information are fused through a concatenation layer, and the feature map is adjusted in the channel through a 1*1 convolution, batch normalization, and a ReLU activation function to obtain the final output feature, that is, the enhanced shallow layer feature map. The feature map is combined into the fusion feature map P6 to supplement the texture information of the deep layer feature. In this example, the high-frequency information of the shallow layer feature map is extracted by the Haar wavelet filter, and the dynamic weight information is generated by the multi-layer perceptron. The enhanced feature and the dynamic weight information are combined and integrated into the deep layer feature to compensate for the lack of texture features in the deep layer feature.
[0068] Step 204: detecting the small target in the infrared image by detecting the head network using the optimized enhanced feature map and the first to fourth enhanced modulation feature maps; specifically, the detection head network adopts an anchor-free detection head network, including but not limited to a CornerNet target detector, wherein the technical principle of the detection head network for detecting the small target in the infrared image is a known technology, and will not be described in detail here.
[0069] Step 3: training and testing the infrared small target detection model using the training data set and the test data set to obtain an optimal detection model; and determining the position of the small target in the infrared image to be detected by the optimal detection model.
[0070] The method for obtaining the optimal detection model specifically includes the following steps:
[0071] Step 31: training the constructed infrared small target detection model using the training data set to obtain a trained infrared small target detection model:
[0072] Step 32: based on the constructed model loss function, evaluating the trained infrared small target detection model using the test set to determine whether the output of the trained infrared small target detection model converges.
[0073] The model loss function constructed in this embodiment is the sum of the classification loss function and the regression loss function, and the classification loss function is measured by Quality Focal Loss, and the regression loss function is measured by GIoU loss.
[0074] If yes, the trained infrared small target detection model is the optimal detection model.
[0075] Otherwise, based on the adaptive optimization algorithm, the model parameter weight of the trained infrared small target detection model is adaptively adjusted, and step 31 is repeatedly executed.
[0076] The infrared small target detection model constructed in this embodiment uses the adaptive optimization algorithm for end-to-end optimization in training, so that the model maintains high detection accuracy while significantly improving the recall rate of small targets; and after training is completed, the test image is input, and the optimal detection model obtained in this embodiment is used to determine the position of the target in the image. At the same time in the testing process, any image is input, and the accuracy of the infrared small target detection can be directly obtained according to the trained model parameters. Figures 5 to 6 As shown in the figure, the trained optimal detection model can accurately extract the infrared small target description in the image and give the current accuracy.
[0077] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting small infrared targets by feature enhancement modulation, characterized in that, The method comprises the steps of: Step 1: acquiring an infrared image for small target detection; The infrared image is sequentially subjected to size adjustment and xml format conversion, and the training set and the test set are divided according to a preset ratio; Step 2: constructing a feature enhancement modulation infrared small target detection model; The infrared small target detection model comprises a feature extraction module, a feature pyramid, an adaptive feature modulation module, a dynamic wavelet feature enhancement module and a detection head network; The feature extraction module extracts a multi-scale feature map of the infrared image, and the multi-scale feature map comprises a plurality of feature maps with gradually reduced detail information; the detail information at least comprises target edge and texture features; The feature pyramid utilizes the multi-scale feature map to obtain a fusion feature map; The adaptive feature modulation module performs adaptive pooling processing on the first to fourth fusion feature maps to obtain adaptive pooling feature maps of different scales, and performs feature modulation processing on the adaptive pooling feature maps to obtain modulation feature maps, and fuses the generated modulation feature maps with the corresponding fusion feature maps to obtain the first to fourth enhanced modulation feature maps; The dynamic wavelet feature enhancement module only performs feature enhancement processing on the first layer of the multi-scale feature map by using a Haar wavelet filter to obtain an enhanced shallow layer feature map, and fuses the fifth fusion feature map with the enhanced shallow layer feature map to obtain an optimized enhanced feature map; the first layer of the multi-scale feature map is the feature map with the most detail information, and the fifth fusion feature map is the fusion feature map with the least detail information; The detection head network utilizes the optimized enhanced feature map and the first to fourth enhanced modulation feature maps to realize detection of small targets in the infrared image; Step 3: training and testing the infrared small target detection model by using the training set and the test set to obtain an optimal detection model; the optimal detection model is used to determine the position of small targets in a to-be-detected infrared image.
2. The method of claim 1, wherein the feature enhanced modulation is a binary phase shift keying (BPSK) modulation. The dynamic wavelet feature enhancement module constructed in step 2 comprises a Haar wavelet filter, a multi-layer perception, a softmax activation function layer and a feature processing module composed of a splicing layer, a convolution layer, a batch normalization layer and a ReLU activation function layer connected in sequence; The Haar wavelet filter extracts high-frequency feature components and low-frequency feature components in the first layer of the multi-scale feature map; the high-frequency feature components comprise horizontal high-frequency features, vertical high-frequency features and diagonal high-frequency features; The multi-layer perception performs a nonlinear mapping operation on the first layer of the multi-scale feature map; After the softmax activation function layer performs an activation operation on the output of the multi-layer perception, a set of dynamic normalized weights are obtained, and the normalized weights are evenly divided in the channel dimension to obtain four divided weights W1-W4; The feature processing module utilizes the enhanced feature map to obtain the enhanced shallow layer feature map; The enhanced feature map is a feature map obtained by performing an element-wise multiplication operation on the divided weights W1-W4 and the low-frequency feature components, the horizontal high-frequency features, the vertical high-frequency features and the diagonal high-frequency features, respectively.
3. The method of claim 1, wherein the feature enhanced modulation is a binary phase shift keying (BPSK) modulation. The adaptive feature modulation module constructed in step 2 comprises a channel division layer, an adaptive pooling layer, a convolution layer, a bilinear interpolation up-sampling layer and a splicing layer; The first fusion feature map to the fourth fusion feature map output by the feature pyramid is split along the channel by the channel segmentation layer and four independent channel segmentation feature maps are obtained; Any three channel segmentation feature maps are subjected to adaptive maximum pooling operation by the adaptive pooling layer using a preset scaling factor to obtain feature maps of different spatial sizes; The first fusion feature map to the fourth fusion feature map output by the feature pyramid is subjected to convolution operation by the convolution layer to perform channel compression, and the average variance information of the feature map is obtained using the channel compressed feature map; the element-wise summation operation is performed on the average variance information and the feature maps of different spatial sizes respectively to obtain the variance feature map; The modulation feature map is obtained by the splicing layer for splicing operation on the deep convolution feature map. The deep convolution feature map is obtained by performing deep convolution operation on the variance feature map using the deep convolution block, and then performing bilinear interpolation operation of different factors by the bilinear interpolation up-sampling layer.
4. The method of claim 1, wherein the feature enhanced modulation is a binary phase shift keying (BPSK) modulation. The method for obtaining the optimal detection model in step 3 comprises the steps of: Step 31: using the training data set to train the constructed infrared small target detection model to obtain a trained infrared small target detection model: Step 32: based on the constructed model loss function, using the test set to evaluate the trained infrared small target detection model, and determining whether the output of the trained infrared small target detection model converges; If yes, the trained infrared small target detection model is the optimal detection model; Otherwise, based on the adaptive optimization algorithm, the model parameter weight of the trained infrared small target detection model is adaptively adjusted, and step 31 is repeatedly executed.
5. The method of claim 4, wherein the feature enhanced modulation is a binary phase shift keying (BPSK) modulation. The model loss function constructed in step 32 is the sum of the classification loss function and the regression loss function, the classification loss function is measured by Quality Focal Loss, and the regression loss function is measured by GIoUloss.
6. The method of claim 1, wherein the feature enhanced modulation is a binary phase shift keying (BPSK) modulation. The feature extraction module in step 2 includes but is not limited to residual network ResNet-50; The detection head network adopts an anchor-free detection head network, including but not limited to CornerNet target detector.