Infrared small target detection method based on multi-scale feature extraction and reconstruction

By adopting multi-scale feature extraction and reconstruction methods in infrared small object detection, the problem of easy loss of target information in infrared small object detection is solved, and higher detection accuracy and completeness are achieved.

CN119919644APending Publication Date: 2025-05-02EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510095103.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing infrared small target detection methods are difficult to effectively identify and locate small infrared targets. They are mainly due to the extremely small target size, dim brightness and lack of obvious texture information, which leads to low signal-to-noise ratio and easy loss of target information.

Method used

Using detection methods based on multi-scale feature extraction and reconstruction, the multi-scale feature extraction module, multi-dimensional feature reconstruction module and detail attention fusion module are used to extract and reconstruct feature information at different scales and levels, reduce redundant information, and enhance target feature representation.

Benefits of technology

It improves the detection effect of small infrared targets, reduces the possibility of missing target information, enhances the feature representation ability of the model, and improves the accuracy and completeness of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919644A_ABST
    Figure CN119919644A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence computer vision, and provides an infrared small target detection method based on multi-scale feature extraction and reconstruction, and the method comprises the following steps: 1, obtaining an infrared image data set, and carrying out the data preprocessing; 2, constructing a multi-scale feature extraction module; step 3, constructing a multi-dimensional feature reconstruction module; 4, constructing a detail attention fusion module; and 5, continuously training the constructed model, and obtaining a final infrared small target detection effect after the model obtains a convergence state. The infrared small target detection effect has relatively high detection precision and intersection-to-parallel ratio, and the problems of small target loss, high false detection rate and the like in the existing infrared small target detection can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence computer vision, and in particular to an infrared small target detection method based on multi-scale feature extraction and reconstruction. Background Art

[0002] In today's imaging technology, infrared imaging technology has shown important application value and broad application prospects with its unique advantages. The infrared imaging system generates images by keenly capturing the thermal radiation emitted by the target. This imaging method has significant advantages over the traditional visible light imaging mode. Specifically, the excellent penetration ability of infrared imaging enables it to easily penetrate various obstructions such as smoke, dust and fog, thereby ensuring stable all-weather imaging capabilities in complex environments, providing reliable visual information support for various practical applications. At the same time, compared with radar imaging, infrared imaging is a passive detection method that does not require active signal transmission. This feature gives it a strong concealment advantage, enabling it to effectively avoid being detected by the target in key tasks such as military reconnaissance, thereby greatly improving the success rate and safety of reconnaissance operations. Therefore, it is highly favored in the military field, and it also plays an indispensable role in many fields such as forest fire monitoring, traffic safety monitoring and industrial detection. Its importance is increasing day by day and has received widespread attention and in-depth research.

[0003] However, infrared imaging technology also faces a series of severe challenges in practical applications. Since its imaging distance is often long, small infrared targets appear extremely small in the image, usually occupying only a few pixels or even less than a pixel, which makes it very easy to miss detection during the detection process, seriously affecting the accuracy and completeness of the detection. In addition, small infrared targets are not only small in size, but also very dim in brightness, with almost no obvious texture information, which directly leads to a low signal-to-noise ratio, making these small targets easily submerged in the complex and changing clutter background, difficult to be accurately identified and located, further increasing the difficulty and complexity of detection.

[0004] Although deep learning technology has made rapid progress in recent years, and the performance of convolutional neural networks in many target detection tasks has also been significantly improved, due to the target characteristics caused by the unique imaging principle of infrared imaging itself, traditional deep learning-based target detection methods are difficult to effectively solve the problem of low detection rate of infrared small targets. General target detection networks usually rely on rich texture information of the target to extract features effectively, and gradually separate the detection target from the complex background through multi-level downsampling operations, so as to achieve accurate recognition and positioning of the target. However, due to the inherent characteristics of small infrared targets, which are extremely small in size and dim in color, after multiple downsampling processes, it is very easy to cause the risk of serious loss of target information, which makes it difficult for the detection model to accurately capture the key features of the target. At the same time, the infrared image itself lacks obvious texture information and has low contrast, which makes it difficult for small targets to stand out in complex backgrounds, further exacerbating the possibility of being submerged by clutter backgrounds, making the existing general target detection network incapable of facing the task of infrared small target detection, and unable to meet the urgent needs of practical applications for high detection rate and accuracy. Therefore, for the challenging task of infrared small target detection, there is an urgent need to conduct in-depth research and development of specialized and efficient detection algorithms and models to overcome the current difficulties, improve the performance and reliability of infrared small target detection, and thus promote the wider and deeper application and development of infrared imaging technology in various fields. Summary of the invention

[0005] In view of the above problems existing in the existing infrared small target detection, the present invention proposes an infrared small target detection method based on multi-scale feature extraction and reconstruction. Facing the problems that small targets in infrared images are dim in color, small in size and easy to lose target information in multiple downsampling, this method can use different scales to extract feature information at different levels, and reconstruct it to reduce redundant information, enhance the feature representation of the target, and improve the detection effect of infrared small targets.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] An infrared small target detection method based on multi-scale feature extraction and reconstruction, the infrared small target detection method comprises the following steps:

[0008] Step 1: Obtain infrared image dataset and perform data preprocessing, including normalizing the image, randomly cropping image blocks of fixed size, and flipping and enhancing the cropped image blocks;

[0009] Step 2: Construct a multi-scale feature extraction module, input the preprocessed and enhanced image data into the multi-scale feature extraction and reconstruction infrared small target detection network based on the U-Net framework, and use the branches of different scales to extract feature information of different scales and levels through the multi-scale feature extraction module;

[0010] Step 3: Construct a multi-dimensional feature reconstruction module. In the downsampling bottom layer of the infrared small target detection network with multi-scale feature extraction and reconstruction, the multi-dimensional feature reconstruction module is used to separate redundant spatial features, reduce the influence of irrelevant background on infrared small target detection, and enhance feature representation.

[0011] Step 4: Construct a detail attention fusion module. By fusing the features in the downsampling and the features after upsampling through the detail attention block, focus on the important areas in each channel, fuse more useful information, reduce the possibility of forgetting key information during the downsampling process, and improve the detection effect;

[0012] Step 5: Compare the detected infrared small target image after training the infrared small target detection network with multi-scale feature extraction and reconstruction with the real infrared small target image label to obtain the loss, continuously train the data set, and after the model reaches a convergence state, obtain the final infrared small target detection result.

[0013] Furthermore, the data preprocessing in step 1 specifically includes normalizing images of different data sets, converting mask pixel values ​​from a range of 0-255 to a range of 0-1, retaining only the channel of the first image, and then randomly cropping image blocks and mask blocks of 256×256 size to increase data diversity, and finally randomly flipping the image blocks horizontally and vertically.

[0014] Furthermore, the multi-scale feature extraction module in step 2 is the main feature extraction module in the infrared small target detection network of multi-scale feature extraction and reconstruction. Small-scale information and large-scale information are calculated by linear transformation of patch block sizes of different dimensions, and the original features of the infrared small target are retained by residual calculation. Feature information of different scales is calculated by controlling the size of the patch block. When the size of the patch block is 2, small-scale information is calculated, and when the size of the patch block is 4, large-scale information is calculated. Parallel general branch sampling residual blocks are used for calculation. The specific process is as follows:

[0015] F local =LGA(X,2)

[0016] F global =LGA(X,4)

[0017] F original =ResBlock(X)

[0018] Where X is the input feature, LGA(·,·) represents the linear transformation of aggregating and shifting non-overlapping patches in the spatial dimension by controlling the size of the patch block, and F local and F global represents the small-scale information and large-scale information of the patch block size of 2 and 4 respectively, ResBlock(·) represents the residual calculation of the residual block, and F original Represents the original information after residual calculation.

[0019] After obtaining feature information of different scales, the information of each scale is spliced ​​and the spatial attention mechanism and channel attention mechanism are used to perform adaptive feature enhancement. Finally, the feature information after multi-scale feature extraction is output. The specific process is as follows:

[0020] F X =Cat(F local ,F global ,F original )

[0021] MDF(X)=Relu(BN(CASA(F X )))

[0022] Among them, Cat(·) means concatenating feature information, CASA(·) means channel and spatial attention mechanism, BN(·) means batch normalization, Relu(·) means Relu activation function, and MDF(·) means the output after multi-scale feature extraction.

[0023] Furthermore, the multi-dimensional feature reconstruction module of step 3 reconstructs the features of the spatial dimension to highlight the local spatial features of the small target and suppress background interference. The spatial dimension feature reconstruction divides the spatial features into important information and unimportant information and recombines them, which helps to locate the small target more accurately. First, the input feature map is group normalized to obtain the normalized weights reflecting the importance of different feature maps. The specific process is as follows:

[0024] γ i =GN(X)

[0025]

[0026] Among them, GN(·) is group normalization, γ i represents the trainable parameters after group normalization, C represents the number of channels of the input image, and w γ represents the normalized correlation weight.

[0027] After obtaining the standard normalized weight, map the normalized weight to the range of (0,1), and filter the set threshold 0.5 through the gating function. Set the weight W1 for weights greater than the threshold, and set the weight W2 for weights less than the threshold. The specific process is as follows:

[0028] G i =Sigmoid(w γ GN(X)

[0029]

[0030] Among them, Sigmoid(·) represents the Sigmoid activation function, G i represents the probability value after group normalization and weight multiplication, Gate(·) represents the gating function, and the weights greater than or equal to the threshold 0.5 are set to 1 to obtain the weights W1 corresponding to the information, and the weights less than the threshold 0.5 are set to 0 to obtain the weights W2 corresponding to the non-information.

[0031] Finally, in order to reduce spatial redundancy, a cross reconstruction operation is used. The overall process of spatial dimension feature reconstruction can be expressed by the following formula:

[0032]

[0033] in, and It means that the division along the channel dimension is obtained and Cat(·) means concatenating the feature information, and MFR(·) means that the divided features are finally cross-superimposed and concatenated to obtain the output after spatial feature reconstruction.

[0034] Furthermore, the detail attention fusion module of step 4 performs detail fusion on the upsampled features of the layer and the corresponding downsampled features, and obtains the fusion weights by using space, channel and pixel attention calculations, which can more effectively extract and integrate small targets and background information in infrared images. The overall process of detail attention fusion can be expressed by the following formula:

[0035] W casa =CA(X+Y)+SA(X+Y)

[0036] W pa =Sigmoid(PA(W casa ,X+Y))

[0037] DAB(X,Y)=Conv(X+Y+W pa *X+(1-W pa )*Y)

[0038] Among them, X and Y represent up-sampled features and down-sampled features respectively, CA(·) represents the channel attention mechanism, SA(·) represents the spatial attention mechanism, and W case represents the features extracted after performing channel and spatial attention on the input features respectively, PA(·,·) represents the pixel attention mechanism, Sigmoid(·) represents the Sigmoid activation function, Conv(·) represents 1×1 convolution, and DAB(·,·) represents the fusion result after the detail attention fusion block.

[0039] Furthermore, in step 5, the preprocessed data set image is sent to the infrared small target detection network for multi-scale feature extraction and reconstruction for training, and after the model reaches a convergence state, the final infrared small target detection effect is obtained.

[0040] Compared with the prior art, the present invention has the following advantages:

[0041] (1) The present invention proposes a multi-scale feature extraction module, which extracts feature information of different scales through a multi-branch strategy of different scales and a dilation convolution with different expansion rates. It can retain key information during multiple downsampling processes and reduce the possibility of information loss of small infrared targets.

[0042] (2) The present invention proposes a multi-dimensional feature reconstruction module, which utilizes group normalization to evaluate the information content of feature maps, can accurately separate informative and uninformative feature maps, and further adopts a cross-reconstruction operation to effectively suppress spatial redundancy and enhance the feature representation capability of the model.

[0043] (3) The present invention proposes a detail attention fusion module, which fully mixes the channel attention weights and the spatial attention weights, ensuring the effective interaction of information, enabling the model to comprehensively consider the information of the channel and spatial dimensions, and fully integrate the feature information adopted by the downsampling and the corresponding upsampling. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of the infrared small target detection method based on multi-scale feature extraction and reconstruction proposed by the present invention;

[0045] Figure 2 It is a schematic diagram of the network structure of the infrared small target detection method based on multi-scale feature extraction and reconstruction proposed by the present invention;

[0046] Figure 3 It is a schematic diagram of the structure of the multi-scale feature extraction module proposed in the present invention;

[0047] Figure 4 It is a schematic diagram of the module structure of the multi-dimensional feature reconstruction proposed by the present invention;

[0048] Figure 5 It is a schematic diagram of the structure of the detail attention fusion module proposed in the present invention;

[0049] Figure 6 This is a diagram showing the effect of the infrared small target detection method proposed by the present invention;

[0050] Figure 7 It is a three-dimensional effect diagram detected by the infrared small target detection method proposed in the present invention. DETAILED DESCRIPTION

[0051] In order to more clearly illustrate the purpose, technical solutions and advantages of the present invention, the present invention is described in detail in combination with the accompanying drawings and the following embodiments. In the accompanying drawings, the same reference numerals represent the same or similar components. The specific embodiments described below are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0052] The present invention proposes an infrared small target detection method based on multi-scale feature extraction and reconstruction, referring to the flowchart Figure 1 , the specific steps of this method are as follows:

[0053] Step 1: Obtain an infrared image dataset and perform preprocessing operations on the dataset images, including normalization, fixed-size random cropping, and random vertical or horizontal flipping.

[0054] Step 2: Construct an overall detection network structure for infrared small target detection based on multi-scale feature extraction and reconstruction. The overall detection structure is constructed based on a 4-layer U-Net basic framework. The main feature extraction uses the constructed multi-scale feature extraction module to extract multi-scale features.

[0055] Step 3: At the bottom of the overall network structure, a multi-dimensional feature extraction module is constructed to reconstruct the features of the spatial dimension, highlight the local spatial features of the small target, and suppress background interference.

[0056] Step 4: During the upsampling process, a detail attention fusion module is constructed to replace the skip connection, which fuses the upsampled features with the corresponding downsampled features, effectively integrates the feature information of small infrared targets, and reduces the possibility of small target information loss.

[0057] Step 5: After building the overall network structure, perform model training on the preprocessed data set until the model converges. Finally, use the converged model to detect infrared small targets and obtain the detected infrared small target image.

[0058] Figure 2The schematic diagram of the network structure of the infrared small target detection method based on multi-scale feature extraction and reconstruction is shown in Figure 1. The overall network structure is based on a 4-layer U-Net. In downsampling, each layer uses a multi-scale feature extraction module to extract the feature information of the image; at the connection between downsampling and upsampling, a multi-dimensional feature reconstruction module is used to reduce the interference of the background on the target; in the upsampling process, the feature information of the upsampling and the corresponding downsampling is fused through the detail attention fusion module to reduce the probability of infrared small target information loss; finally, after four layers of upsampling, the image of the detected infrared small target is output.

[0059] Figure 3 The schematic diagram of the multi-scale feature extraction module is the main feature extraction module of the infrared small target detection method based on multi-scale feature extraction and reconstruction. The specific implementation method is to take the image information X as the input of the module, and extract the feature information of the image in three different scales. The middle part is the extraction of the original feature information, which is extracted by a residual block ResBlock; the upper and lower parts are used to extract small-scale local information and large-scale global information respectively, and different extractions are performed by controlling the block size in the linear transformation, where the block size p=2 is used to extract small-scale information, and the block size p=4 is used to extract large-scale information; after extraction, the image information needs to be restored to the same size as the original information for subsequent image processing. The specific operation is to perform average pooling, convolution and transposition in sequence to restore the image information of different sizes to the same size information. Then, through the splicing of the channel dimension, the small-scale information, large-scale information and original feature information are spliced ​​together and input into the channel and spatial attention module for further feature extraction. Finally, after Sigmoid activation function and batch normalization processing, the result of multi-scale feature extraction is output. The overall process of the multi-scale feature extraction module can be expressed by the following formula:

[0060] F X =Cat(LGA(X,2),LGA(X,4),ResBlock(X))

[0061] MDF(X)=Relu(BN(CASA(F X )))

[0062] Among them, Cat(·) means concatenating feature information, LGA(·,·) means linear transformation of aggregating and displacing non-overlapping patches in the spatial dimension by controlling the size of the patch block, LGA(X,2) and LGA(X,4) respectively represent the small-scale information and large-scale information of patch block sizes of 2 and 4, ResBlock(·) represents the original image information after residual calculation of the residual block, CASA(·) represents the channel and spatial attention mechanism, BN(·) represents batch normalization, Relu(·) represents the Relu activation function, and MDF(·) represents the output after multi-scale feature extraction.

[0063] Figure 4 The schematic diagram of the module structure for multi-dimensional feature reconstruction is as follows: the input feature X is first group-normalized, and the normalization-related weights are calculated, γ i =GN(X) represents the process of group normalization, γ i represents the trainable parameters after group normalization, represents the process of calculating the normalized related weights, C represents the number of channels of the input image, and W γ Represents the normalized correlation weight. After obtaining the normalized correlation weight, G i =Sigmoid(w γ GN(X)) activation function is used to map the normalized weights to the range of (0,1) so that the corresponding information weight W1 and the non-corresponding information weight W2 can be obtained through gated convolution. The specific process can be expressed by the following formula:

[0064]

[0065] Wherein, Gate(·) represents a gating function, which sets weights greater than or equal to a threshold of 0.5 to 1 to obtain weights W1 corresponding to information, and sets weights less than the threshold of 0.5 to 0 to obtain weights W2 corresponding to non-information.

[0066] Finally, in order to reduce spatial redundancy, a cross reconstruction method is used. After matrix multiplication of the original image information and weight information W1 and W2, the obtained feature information is divided according to the channel dimension, and then the feature information of the corresponding dimension is added and the channel is spliced. The specific process can be expressed by the following formula:

[0067]

[0068] in, and It means that the division along the channel dimension is obtained and Cat(·) means concatenating the feature information, and MFR(·) means that the divided features are finally cross-superimposed and concatenated to obtain the output after spatial feature reconstruction.

[0069] Figure 5 The following is a schematic diagram of the structure of the detail attention fusion module. First, two feature information X and Y are input. The low-level features and high-level features are added together for preliminary feature fusion. Then, the fused information is used to calculate the channel attention and spatial attention in parallel. Then, the information after the channel and spatial attention calculations is added to the feature information after preliminary fusion. Then, the added feature information is input into the pixel attention and activation function for calculation. The specific process can be expressed by the following formula:

[0070] W casa =CA(X+Y)+SA(X+Y)

[0071] W pa =Sigmoid(PA(W casa ,X+Y))

[0072] Among them, X and Y represent up-sampled features and down-sampled features respectively, CA(·) represents the channel attention mechanism, SA(·) represents the spatial attention mechanism, and W casa represents the features extracted after performing channel and spatial attention on the input features respectively, PA(·,·) represents the pixel attention mechanism, and Sigmoid(·) represents the Sigmoid activation function.

[0073] Then, in order to balance the contribution of low-level features and high-level features, W is used pa and 1-W pa Weighted processing of low-level and high-level features, multiplying the low-level feature X by its contribution W pa And multiply the high-level feature Y by its contribution 1-W pa Finally, the product of their different contributions and the result of the pixel attention output are added and input into the convolution to get the fused result. The specific process can be expressed by the following formula:

[0074] DAB(X,Y)=Conv(X+Y+W pa *X+(1-W pa )*Y)

[0075] Among them, Conv(·) represents a 1×1 convolution, W pa Represents the contribution of low-level features, 1-W pa represents the contribution of high-level features, and DAB(·,·) represents the fusion result after the detail attention fusion block.

[0076] Figure 6The figure shows the detection effect of the infrared small target detection method proposed by the present invention, which shows the detection effect of infrared small target detection on the NUAA-SIRST data set. It can be seen that in the original infrared image, the size of the small target is very small and the color is dim, and the background is also very complex, with interference from clouds and fog, but the information of the image detected by the detection method of the present invention is very clear and the detection rate is high. Specifically, the infrared small targets are successfully detected, and there is no false detection. This result reflects the effectiveness and superiority of the infrared small target detection method based on multi-scale feature extraction and reconstruction of the present invention.

[0077] Figure 7 for Figure 6 The three-dimensional image effect diagram of infrared small target detection proposed in the present invention can better show that there is a lot of interference from clutter background in the original image data according to the effect of the three-dimensional image, but after processing and detection by the model, the infrared small target immersed in the clutter background can be effectively separated and effectively detected.

Claims

1. A method for detecting small infrared targets based on multi-scale feature extraction and reconstruction, characterized in that: The following steps are involved: Step 1: Obtain infrared image dataset and perform data preprocessing, including normalizing the image, randomly cropping image blocks of fixed size, and flipping and enhancing the cropped image blocks; Step 2: Construct a multi-scale feature extraction module, input the preprocessed and enhanced image data into the multi-scale feature extraction and reconstruction infrared small target detection network based on the U-Net framework, and use the branches of different scales to extract feature information of different scales and levels through the multi-scale feature extraction module; Step 3: Construct a multi-dimensional feature reconstruction module. In the downsampling bottom layer of the infrared small target detection network with multi-scale feature extraction and reconstruction, the multi-dimensional feature reconstruction module is used to separate redundant spatial features, reduce the influence of irrelevant background on infrared small target detection, and enhance feature representation. Step 4: Construct a detail attention fusion module. By fusing the features in the downsampling and the features after upsampling through the detail attention module, focus on the important areas in each channel, fuse more useful information, reduce the possibility of forgetting key information during the downsampling process, and improve the detection effect; Step 5: Compare the detected infrared small target image after training the infrared small target detection network with multi-scale feature extraction and reconstruction with the real infrared small target image label to obtain the loss, continuously train the data set, and after the model reaches a convergence state, obtain the final infrared small target detection result.

2. The infrared small target detection method based on multi-scale feature extraction and reconstruction according to claim 1 is characterized in that: The data preprocessing in step 1 specifically includes normalizing the images in the data set, converting the mask pixel values ​​from the range of 0-255 to the range of 0-1, and retaining only the channel of the first image, and then randomly cropping image blocks and mask blocks of specified sizes with a certain probability to increase the diversity of the data, and finally randomly flipping the image blocks horizontally and vertically.

3. The infrared small target detection method based on multi-scale feature extraction and reconstruction according to claim 1 is characterized in that: The multi-scale feature extraction module in step 2 is the main feature extraction module in the infrared small target detection network for multi-scale feature extraction and reconstruction. It calculates small-scale information and large-scale information through linear transformation of patch block sizes of different dimensions, and retains the original features of the infrared small target through residual calculation. Feature information of different scales is calculated by controlling the size of the patch block. When the size of the patch block is 2, small-scale information is calculated, and when the size of the patch block is 4, large-scale information is calculated. Parallel general branch sampling residual blocks are used for calculation. The specific process is as follows: F local =LGA(X,2) F global =LGA(X,4) F original =ResBlock(X) Where X is the input feature, LGA(·,·) represents the linear transformation of aggregating and shifting non-overlapping patches in the spatial dimension by controlling the size of the patch block, and F local and F global represents the small-scale information and large-scale information of the patch block size of 2 and 4 respectively, ResBlock(·) represents the residual calculation of the residual block, and F original Represents the original information after residual calculation. After obtaining feature information of different scales, the information of each scale is spliced ​​and the spatial attention mechanism and channel attention mechanism are used to perform adaptive feature enhancement. Finally, the feature information after multi-scale feature extraction is output. The specific process is as follows: F X =Cat(F local ,F global ,F original ) MDF(X)=Relu(BN(CASA(F X ))) Among them, Cat(·) means concatenating feature information, CASA(·) means channel and spatial attention mechanism, BN(·) means batch normalization, Relu(·) means LGA activation function, and MDF(·) means the output after multi-scale feature extraction.

4. The infrared small target detection method based on multi-scale feature extraction and reconstruction according to claim 1 is characterized in that: The multi-dimensional feature reconstruction module in step 3 reconstructs the features of the spatial dimension to highlight the local spatial features of the small target and suppress background interference. The spatial dimension feature reconstruction divides the spatial features into important information and unimportant information and recombines them, which helps to locate the small target more accurately. First, the input feature map is group normalized to obtain the normalized weights reflecting the importance of different feature maps. The specific process is as follows: γ i =GN(X) Among them, GN(·) is group normalization, γ i represents the trainable parameters after group normalization, C represents the number of channels of the input image, and w γ represents the normalized correlation weight. After obtaining the standard normalized weight, map the normalized weight to the range of (0,1), and filter the set threshold 0.5 through the gating function. Set the weight W1 for weights greater than the threshold, and set the weight W2 for weights less than the threshold. The specific process is as follows: G i =Sigmoid(w γ ·GN(X)) Among them, Sigmoid(·) represents the Sigmoid activation function, G i represents the probability value after group normalization and weight multiplication, Gate(·) represents the gating function, and the weights greater than or equal to the threshold 0.5 are set to 1 to obtain the weights W1 corresponding to the information, and the weights less than the threshold 0.5 are set to 0 to obtain the weights W2 corresponding to the non-information. Finally, in order to reduce spatial redundancy, a cross reconstruction operation is adopted. The overall process of spatial dimension feature reconstruction can be expressed by the following formula: in, and It means that the division along the channel dimension is obtained and Cat(·) means concatenating the feature information, and MFR(·) means that the divided features are finally cross-superimposed and concatenated to obtain the output after spatial feature reconstruction.

5. The infrared small target detection method based on multi-scale feature extraction and reconstruction according to claim 1 is characterized in that: The detail attention fusion module in step 4 performs detail fusion on the upsampled features of the layer and the corresponding downsampled features, and obtains the fusion weights by using spatial, channel and pixel attention calculations, which can more effectively extract and integrate small targets and background information in infrared images. The overall process of detail attention fusion can be expressed by the following formula: W casa =CA(X+Y)+SA(X+Y) W pa =Sigmoid(PA(W casa ,X+Y)) DAB(X,Y)=Conv(X+Y+W pa *X+(1-W pa )*Y) Among them, X and Y represent up-sampled features and down-sampled features respectively, CA(·) represents the channel attention mechanism, SA(·) represents the spatial attention mechanism, and W casa represents the features extracted after performing channel and spatial attention on the input features respectively, PA(·,·) represents the pixel attention mechanism, Sigmoid(·) represents the Sigmoid activation function, Conv(·) represents 1×1 convolution, and DAB(·,·) represents the result after the detail attention fusion block.