Feature fusion method for infrared weak and small target detection based on neural network

Through the dynamic weighted fusion method of feature extraction, fusion and post-processing modules, the redundancy and information loss problems in deep and shallow feature fusion are solved, and the effect of infrared weak target detection is improved.

CN120388261APending Publication Date: 2025-07-29SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510496759.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing feature fusion method cannot adaptively realize the effective fusion of deep and shallow features, resulting in feature redundancy and information loss in infrared weak target detection.

Method used

The feature extraction module, feature fusion module and post-processing module are adopted to realize adaptive dynamic weighted fusion of deep and shallow features through upsampling module, dynamic weighted fusion module and residual convolution module to enhance feature expression.

Benefits of technology

It significantly improves the performance of infrared weak target detection, makes full use of effective information, avoids information redundancy, and improves detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388261A_ABST
    Figure CN120388261A_ABST
Patent Text Reader

Abstract

The invention relates to a feature fusion method for infrared weak and small target detection based on a neural network, and the method employs a feature extraction module, a feature fusion module and a post-processing module of the neural network to carry out the dynamic weighting of multi-scale context information, enhances the feature expression, and improves the infrared weak and small target detection performance. The problems of feature redundancy and information loss caused by the fact that an existing feature fusion method cannot achieve effective fusion of deep features and shallow features in a self-adaptive mode are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and target detection, and relates to a feature fusion method for infrared small and weak target detection based on a neural network. Background Art

[0002] Space-based infrared target detection technology, as a key technology of space-based infrared detection systems, plays a huge role in space exploration, obstacle avoidance, disaster prevention and warning, etc. However, targets in infrared images have always been difficult to detect due to characteristics such as low gray level, no texture, and small size. Existing target detection methods based on neural networks all use methods of fusing deep and shallow features to fully learn feature information and improve target detection performance. However, for infrared small and weak targets, excessive downsampling leads to target loss. Therefore, directly fusing deep and shallow feature layers will add a lot of redundant information, and the information of small and weak targets may also disappear with the number of downsampling times, resulting in unsatisfactory detection effects for small and weak targets. Therefore, how to dynamically process deep and shallow layer features, which can not only ensure the full utilization of effective information but also ignore useless information to achieve efficient feature fusion, is very important for small target detection. Summary of the Invention

[0003] The purpose of the present invention is to provide a feature fusion method for infrared small and weak target detection based on a neural network, mainly to solve the problem that existing feature fusion methods cannot adaptively achieve effective fusion of deep features and shallow features, resulting in feature redundancy and information loss.

[0004] To achieve the above purpose, the technical solution of the present invention is:

[0005] A feature fusion method for infrared small and weak target detection based on a neural network, the neural network is provided with a feature extraction module, a feature fusion module, and a post-processing module; wherein the feature fusion module includes an upsampling module, a dynamic weighted fusion module, and a residual convolution module;

[0006] The specific steps of using the neural network for feature fusion of infrared small and weak target detection are:

[0007] S1, input the data into the feature extraction module to obtain initial feature layers of different scales;

[0008] S2, input the initial feature layers obtained in step S1 into the feature fusion module to achieve adaptive dynamic weighted fusion of shallow detail information and deep semantic information;

[0009] S3, input the feature layer fused through step S2 into the post-processing module to obtain the prediction result of the target.

[0010] The upsampling module adjusts feature layers of different scales to the same resolution to facilitate subsequent dynamic feature dynamic fusion;

[0011] The dynamic weighted fusion module weights and refines the deep and shallow features of different scales for adaptive and efficient fusion;

[0012] The residual convolution module further processes the module after dynamic weighted feature fusion. The residual connection avoids losing detailed information and fuses context features while retaining information.

[0013] Step S1 inputs an image with a resolution of A×B×C into a feature extraction module composed of multiple residual convolution modules, and outputs four feature maps of different scales containing different feature information

[0014]

[0015] where A represents the image length, B represents the image width, and C represents the number of image channels.

[0016] In S2, the initial feature layer obtained in step S1 is input into the feature fusion module to achieve the adaptive dynamic weighted fusion of shallow layer detail information and deep layer semantic information; several upsampling modules, the dynamic weighted fusion module M (i,i+1) (2≤i≤4) and the convolutional residual module R i (2≤i≤4), M (i,i+1) represents the dynamic weighted fusion of the i-th feature layer and the i+1-th feature layer, so as to achieve progressive feature fusion of features with different resolutions, specifically including the following steps:

[0017] S201, the upsampling module upsamples the deep layer features to make their resolution consistent with that of the shallow layer features;

[0018] S202, the dynamic weighted fusion module consists of two parallel branches. Branch one consists of a convolution and an adaptive hierarchical refinement module H, that is:

[0019] F i_1 =H(conv(F i )+up(conv(F i+1 ))),(2≤i≤4) where H is the adaptive hierarchical refinement module, up represents the upsampled features, and F i represents the shallow layer features, and F i+1 represents the deep layer features.

[0020] S203, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the input deep and shallow layer features, first uses a simple channel attention mechanism to perform attention weighting on the deep and shallow layer features to obtain channel dimension weight information and pay more attention to effective information. The calculation is as follows: W D =CA(Fi+1 ), W S = CA(F i+1 ), where CA represents the channel attention mechanism, CA = sigmoig(conv(Relu(conv(pooling(F))))), sigmoid represents the sigmoid activation function, conv represents the convolution operation, Relu represents the activation function, and pooling represents the pooling operation;

[0021] S204. The adaptive hierarchical refinement module H in the dynamic weighted fusion module calculates the global weight W after adding the shallow and deep layer features through global pooling and sigmoid G = sigmoid(pooling(F i+1 + F i ))), ensuring the full utilization of global and local information;

[0022] S205. The adaptive hierarchical refinement module H in the dynamic weighted fusion module performs dynamic weighted fusion on the channel attention weights of the shallow and deep features obtained in step S203 to obtain the weighted weights ε is a constant;

[0023] S206. The adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the different feature attention weights and the global weight W G , adjusts the local weight according to the global weight to avoid local information dominating the network decision-making. The calculation is as follows:

[0024] S207. The adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the different feature attention weights obtained in step S206, performs refined weighted fusion on the shallow and deep layer features, and then enhances the spatial correlation and enriches the spatial information hierarchical information of the weighted fusion feature layer through the spatial attention feature, that is: F H represents the feature layer finally output by the adaptive hierarchical refinement module H;

[0025] S208. The dynamic weighted fusion module is composed of two parallel branches. Branch two is composed of a convolution and a grouped enhanced attention module E, that is: F i_2 = E(conv(F i ) + up(conv(F i+1 ))), (2 ≤ i ≤ 4) where E is the grouped enhanced attention module, up represents upsampling the feature, F i represents the shallow layer feature, Fi+1 Represents deep features;

[0026] S209. In the grouped enhancement attention module E in the dynamic weighted fusion module, the deep and shallow features are respectively used to extract features through grouped convolution, and then the deep and shallow features are fused by addition. Then, a global attention weight is calculated through a convolution and sigmoid operation. The important target information is captured by calculating the weight, and a residual connection is introduced to avoid losing detailed information. The calculation formula is: F E = sigmoid(conv(Groupconv(F i+1 )) + Groupconv(F i )))*F i + F i , where Groupconv represents grouped convolution, and F E represents the features after being fused by the grouped enhancement attention module;

[0027] S210. In the dynamic weighted fusion module, the finally fused features after dynamic weighted fusion are the sum of the fusion results of the two branches. The formula is expressed as: F M(i,i+1) = F H + F E ;

[0028] S211. In the dynamic weighted fusion module, the finally fused feature layer passes through a residual convolution module to further process and enhance the detailed information;

[0029] S212. In the dynamic weighted fusion module, through the adaptive hierarchical refinement module, according to steps S201 - S211, from deep to shallow, the deep and shallow features are progressively and adaptively fused to achieve efficient feature fusion and enhance the object detection performance.

[0030] The specific steps of step S3 are as follows:

[0031] S301. For the final fusion result obtained in step S2, through a simple object segmentation head, a predicted mask image is obtained;

[0032] S302. For the mask image obtained in step S301, threshold segmentation is performed with the threshold being T to obtain the final predicted binary image.

[0033] The advantages of the present invention are as follows: 1. The feature fusion method of the present invention makes full use of context and multi-scale information, and uses a multi-layer attention mechanism to dynamically weight and fuse shallow and deep features, making full use of effective information, avoiding information redundancy, achieving efficient fusion, and significantly improving the detection ability of infrared small and weak targets; 2. The present invention uses a deep residual network to extract features of different scales, and then adaptively weights and fuses the features of different scales. The efficient dynamic weighted fusion module consists of two parallel branches. Branch one is an adaptive hierarchical refinement module, which adaptively weights and refines the shallow and deep feature layers, avoids feature redundancy, and efficiently uses effective information. Branch two is a grouped attention gating mechanism, which realizes accurate and efficient feature fusion through the cooperation of the two branches. The feature fusion method of the present invention can dynamically weight multi-scale context information, enhance feature expression, and improve the performance of infrared small and weak target detection; 3. The present invention uses grouped enhanced attention weights, which can not only retain the detailed information of the target, but also realize the weighted fusion of multi-scale features, effectively improving the detection rate of infrared small and weak targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the flow chart of the steps of the feature fusion method of the present invention;

[0035] Figure 2 is the overall network structure diagram of the present invention;

[0036] Figure 3 is the structure diagram of the feature fusion method of the present invention;

[0037] Figure 4 is the structure diagram of the adaptive hierarchical refinement module of branch one of the feature fusion method of the present invention;

[0038] Figure 5 is the structure diagram of the grouped enhanced attention module of branch two of the feature fusion method of the present invention;

[0039] Figure 6 is the final detection result diagram in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] The present invention will be further described below with reference to the accompanying drawings. The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent.

[0041] In order to more concisely describe this embodiment, some components that are well known to those skilled in the art but not relevant to the main content of the present invention will be omitted in the drawings or the description. In addition, for the sake of clarity of expression, some components in the drawings will be omitted, enlarged or reduced, but this does not represent the size or all structures of the actual product.

[0042] The present invention discloses a feature fusion method for infrared small and weak target detection based on a neural network, asFigures 1 - 5 As shown, the neural network is provided with a feature extraction module, a feature fusion module and a post-processing module; wherein the feature fusion module includes an upsampling module, a dynamic weighted fusion module and a residual convolution module.

[0043] Preferably, the upsampling module adjusts the feature layers of different scales to the same resolution to facilitate the subsequent dynamic fusion of dynamic features;

[0044] Dynamic weighted fusion module performs weighted refinement and fusion of deep and shallow features of different scales to achieve adaptive and efficient fusion;

[0045] The residual convolution module further processes the module after the dynamic weighted feature fusion. The residual connection avoids the loss of detail information and integrates context features while retaining information.

[0046] The specific steps of feature fusion for infrared dim target detection using this neural network are:

[0047] S1, input the data into the feature extraction module to obtain the initial feature layers of different scales;

[0048] Specifically, step S1 is to input the image with a resolution of A×B×C into a feature extraction module composed of multiple residual convolution modules, and output the following four feature maps with different scales containing different feature information:

[0049]

[0050] Among them, A represents the image length, B represents the image width, and C represents the number of image channels.

[0051] In the combined embodiment, A=512, B=512, C1=64, c2=128, C3=256, C4=512.

[0052] S2, input the initial feature layer obtained in step S1 into the feature fusion module to achieve adaptive dynamic weighted fusion of shallow detail information and deep semantic information;

[0053] Specifically, in step S2, the initial feature layer obtained in step S1 is input into the feature fusion module to realize the adaptive dynamic weighted fusion of shallow detail information and deep semantic information; the feature fusion module includes several upsampling modules, dynamic weighted fusion module M (i,i+1) (2≤i≤4) and convolutional residual module R i (2≤i≤4), M (i,i+1) Represents the dynamic weighted fusion of the i-th feature layer and the i+1-th feature layer, so as to achieve progressive feature fusion of features of different resolutions, which specifically includes the following steps:

[0054] S201, the upsampling module upsamples the deep features to unify their resolution with that of the shallow features;

[0055] S202, the dynamic weighted fusion module consists of two parallel branches, as Figure 4 shown. Branch 1 consists of a convolution and an adaptive hierarchical refinement module H, that is:

[0056] F i_1 = H(conv(F i ) + up(conv(F i+1 ))), (2 ≤ i ≤ 4) where H is the adaptive hierarchical refinement module, up represents the upsampled features, F i represents the shallow features, and F i+1 represents the deep features.

[0057] S203, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the input deep and shallow features, first uses a simple channel attention mechanism to perform attention weighting on the deep and shallow features to obtain channel dimension weight information, paying more attention to the effective information. The calculation is as follows: W D = CA(F i+1 ), W S = CA(F i+1 ), where CA represents the channel attention mechanism, CA = sigmoig(conv(Relu(conv(pooling(F)))), sigmoid represents the sigmoid activation function, conv represents the convolution operation, Relu represents the activation function, and pooling represents the pooling operation;

[0058] S204, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, calculates the global weight W after adding the deep and shallow features through global pooling and sigmoid G = sigmoid(pooling(F i+1 + F i ))), ensuring the full utilization of global and local information;

[0059] S205, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the channel attention weights of the deep and shallow features obtained in step S203, performs dynamic weighted fusion on them to obtain the weighted weight ε is a constant;

[0060] S206, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the different feature attention weights and the global weight W G, adjust the local weight according to the global weight to avoid local information dominating the network decision, which is calculated as follows:

[0061] S207, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the different feature attention weights obtained in step S206 The deep and shallow features are refined and weighted fused, and then the weighted fused feature layer is passed through the spatial attention feature to enhance the spatial correlation and enrich the spatial information level information, namely: F H Represents the feature layer finally output by the adaptive hierarchical refinement module H;

[0062] S208, the dynamic weighted fusion module is composed of two parallel branches, such as Figure 5 As shown, branch 2 consists of a convolution and group-enhanced attention module E, namely: F i_2 =E(conv(F i )+up(conv(F i+1 ))),(2≤i≤4) where E is the group-enhanced attention module, up represents the up-sampled feature, F i Represents shallow features, F i+1 Represents deep features;

[0063] S209, the group enhanced attention module E in the dynamic weighted fusion module extracts features of the deep and shallow features respectively through group convolution, then fuses the deep and shallow features by addition, and then calculates the global attention weight through a convolution and sigmoid operation. The weight is calculated to capture important target information and introduce residual connection to avoid losing detail information. The calculation formula is: F E =sigmoid(conv(Groupconv(F i+1 )+Groupconv(F i ))))*F i +F i ,Groupconv means group convolution, F E Represents the features after fusion by the group-enhanced attention module;

[0064] S210, the dynamic weighted fusion module, the final feature after dynamic weighted fusion is the sum of the fusion results of the two branches, the formula is expressed as: F M(i,i+1) =F H +F E ;

[0065] S211, the dynamic weighted fusion module, the final fused feature layer passes through a residual convolution module to further process and enhance the detail information;

[0066] S212. The dynamic weighted fusion module, through the adaptive hierarchical refinement module, according to steps S201 - S211, from deep to shallow, progressively and adaptively fuses the deep and shallow features to achieve efficient feature fusion and enhance the object detection performance.

[0067] S3. Input the feature layer after being fused in step S2 into the post - processing module to obtain the prediction result of the object.

[0068] The specific content of step S3 is as follows:

[0069] S301. For the final fusion result obtained in step S2, through a simple object segmentation head, obtain the predicted mask image;

[0070] S302. For the mask image obtained in step S301, perform threshold segmentation with the threshold being T to obtain the final predicted binary image. As Figure 6 shown, it is the final detection result graph applying the method of the present invention.

[0071] Compare the method of the present invention with the original FPN method. Except for the different feature fusion methods, the other structures are exactly the same. As shown in the experimental table, it can be seen that our method is superior to FPN.

[0072]

[0073] The above is only the preferred embodiment of the present invention and is not used to limit the scope of implementation of the present invention. That is, all equivalent changes and modifications made according to the content of the patent application scope of the present invention should fall within the technical scope of the present invention.

Claims

1. Feature fusion method for detecting infrared dim small targets based on neural network, characterized in that: The neural network is provided with a feature extraction module, a feature fusion module and a post-processing module; The feature fusion module includes an upsampling module, a dynamic weighted fusion module and a residual convolution module; The specific steps for performing feature fusion for infrared small target detection using the neural network are as follows: S1, Input the data into the feature extraction module to obtain initial feature layers of different scales; S2, Input the initial feature layers obtained in step S1 into the feature fusion module to achieve adaptive dynamic weighted fusion of shallow detail information and deep semantic information; S3, Input the feature layer fused through step S2 into the post-processing module to obtain the prediction result of the target.

2. The feature fusion method according to claim 1, wherein: The upsampling module adjusts the feature layers of different scales to the same resolution for subsequent dynamic feature dynamic fusion; The dynamic weighted fusion module performs weighted refinement fusion on the deep and shallow layer features of different scales to achieve adaptive and efficient fusion; The residual convolution module further processes the module after dynamic weighted feature fusion. The residual connection avoids losing detail information and fuses context features while retaining information.

3. The feature fusion method according to claim 2, wherein: Step S1 is to input an image with a resolution of A×B×C into a feature extraction module composed of multiple residual convolution modules, and output the following four feature maps with different feature information at different scales Where, A represents the image length, B represents the image width, and C represents the number of image channels.

4. The feature fusion method according to claim 2, wherein: In step S2, the initial feature layer obtained in step S1 is input into the feature fusion module to achieve adaptive dynamic weighted fusion of shallow detail information and deep semantic information; several upsampling modules, dynamic weighted fusion module M (i,i+1) (2 ≤ i ≤ 4) and convolutional residual module R i (2 ≤ i ≤ 4), M (i,i+1) represents the dynamic weighted fusion of the i-th feature layer and the (i + 1)-th feature layer, so as to achieve progressive feature fusion of features with different resolutions, which specifically includes the following steps: S201, The upsampling module upsamples the deep layer features to make their resolution consistent with that of the shallow layer features; S202, The dynamic weighted fusion module is composed of two parallel branches. Branch one consists of a convolution and an adaptive hierarchical refinement module H, that is: F i_1 = H(conv(F i ) + up(conv(F i+1 )), (2 ≤ i ≤ 4) where H is an adaptive hierarchical refinement module, up represents upsampling features, F i represents shallow features, and F i+1 represents deep features. S203, in the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the input shallow and deep layer features, first uses a simple channel attention mechanism to perform attention weighting on the shallow and deep layer features to obtain channel dimension weight information, paying more attention to the effective information, and the calculation is as follows: W D = CA(F i+1 ), W S = CA(F i+1 ), where CA represents the channel attention mechanism, CA = sigmoig(conv(Relu(conv(pooling(F)))), sigmoid represents the sigmoid activation function, conv represents the convolution operation, Relu represents the activation function, and pooling represents the pooling operation; S204, the adaptive hierarchical refinement module H in the dynamic weighted fusion module calculates the global weight WG = sigmoid(pooling(F i+1 +F i )) through global pooling and sigmoid, ensuring the full utilization of global and local information; In S205, the adaptive hierarchical refinement module H in the dynamic weighted fusion module performs dynamic weighted fusion on the deep and shallow features according to the channel attention weights of the deep and shallow features obtained in step S203 to obtain weighted weights. ε is a constant; S206, the adaptive hierarchical refinement module H in the dynamic weighted fusion module, according to the different feature attention weights obtained in step S204 and step S205 and the global weight W G , adjust the local weight according to the global weight to avoid local information dominating network decisions, and the calculation is as follows: S207, the adaptive hierarchical refinement module H in the dynamic weighted fusion module refines and weighted fuses the shallow and deep features according to the different feature attention weights obtained in step S206 Then, the weighted fused feature layer is enhanced with spatial correlation and enriched with spatial information and hierarchical information through spatial attention features, that is: F H represents the feature layer finally output by the adaptive hierarchical refinement module H; S208. The dynamic weighted fusion module consists of two parallel branches. Branch two consists of a convolution and a grouped enhanced attention module E, that is: F i_2 = E(conv(F i ) + up(conv(F i+1 )), (2 ≤ i ≤ 4) where E is the grouped enhanced attention module, up represents upsampling the feature, F i represents the shallow feature, and F i+1 represents the deep feature; S209. In the grouped enhancement attention module E in the dynamic weighted fusion module, the shallow and deep features are respectively used to extract features through grouped convolution, and then the shallow and deep features are fused by addition. Then, a global attention weight is calculated through a convolution and sigmoid operation. The important target information is captured by calculating the weight, and a residual connection is introduced to avoid losing detailed information. The calculation formula is: F E = sigmoid(conv(Groupconv(F i+1 ) + Groupconv(F i )))*F i + F i , where Groupconv represents grouped convolution, and F E represents the features after being fused by the grouped enhancement attention module; S210, for the dynamic weighted fusion module, the finally dynamically weighted and fused feature is the sum of the fusion results of the two branches, which is expressed by the formula: F M(i,i+1) = F H + F E ; S211, The finally fused feature layer of the dynamic weighted fusion module passes through a residual convolution module for further processing and enhancing detail information; S212, The dynamic weighted fusion module, through the adaptive hierarchical refinement module, according to steps S201 - S211, from deep layer to shallow layer, progressively adaptively fuses the deep and shallow features to achieve efficient feature fusion and enhance the target detection performance.

5. The feature fusion method according to claim 2, wherein: The specific content of step S3 is as follows: S301, For the final fusion result obtained in step S2, pass it through a simple target segmentation head to obtain a predicted mask image; S302, For the mask image obtained in step S301, perform threshold segmentation with a threshold of T to obtain the final predicted binary image.