Infrared small target detection algorithm based on multi-mode guidance

By generating multimodal features and utilizing cross-modal guided residual modules and multi-head cross-attention mechanisms, the problem of underutilization of modal complementarity in infrared small target detection is solved, achieving efficient detection of weak and small targets.

CN120976755APending Publication Date: 2025-11-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511097275.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing infrared small target detection technologies lack a mechanism to explicitly introduce cross-modal mutual guidance information in the feature extraction and residual enhancement stages, which fails to fully utilize the statistical and spatial consistency between different modalities. Furthermore, the deep feature fusion method is too coarse, resulting in insufficient response to weak targets.

Method used

A multimodal extraction module is used to generate three modal features: local energy map, contrast map, and grayscale map. The cross-modal guided residual module and multi-head cross-attention mechanism of the encoder-decoder structure are used to perform feature interaction and dynamic weighted fusion through depthwise separable convolution and pointwise convolution, and the channel and spatial mutual guidance information between modalities are explicitly integrated.

Benefits of technology

It significantly improves the detection accuracy and robustness of small targets in infrared images, effectively suppresses background interference in complex backgrounds, and enhances the salience of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976755A_ABST
    Figure CN120976755A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target detection algorithm based on multi-modal guidance, and the algorithm comprises the steps: generating three complementary modal features from an input infrared image through a multi-modal extraction module, and enabling the features to reflect the local gray fluctuation, region comparison relation and original brightness information respectively; and then inputting the multi-modal features into a neural network adopting an encoder-decoder structure, introducing a cross-modal guide residual module and explicitly integrating channel and space transconductance information between modals into each layer of an encoder and a decoder, enhancing weak small target features, and after the encoder completes multi-level feature extraction, extracting the multi-modal features by using the neural network. Fine-grained interaction is carried out among different modals through a cross-modal cross fusion module by utilizing multiple groups of multi-head cross attention mechanisms, efficient token mixing is realized by combining depth separable convolution and point-by-point convolution, and dynamic fusion is carried out through adaptive modal weight, so that the detection accuracy and robustness of a weak small target are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and image processing, and particularly relates to an infrared small target detection algorithm based on multi-modal guidance. BACKGROUND

[0002] At present, as an important branch of target detection, infrared small target detection is widely used in key scenes such as sea monitoring, air early warning, unmanned driving, and infrared search and tracking. However, due to the low resolution, few texture details, and weak contrast of the images obtained by infrared imaging devices, small targets in the background are usually weak and not obvious in shape, and are easily submerged by high-frequency noise or clutter in complex background. Therefore, how to effectively suppress background interference and highlight the saliency of weak small targets in infrared images has become a hot and difficult research topic. For this reason, the academic and industrial circles have proposed a variety of detection methods, which can be mainly divided into traditional detection algorithms based on artificial features and end-to-end detection methods based on deep learning.

[0003] In traditional methods, local contrast, local energy, gradient, and other artificially designed saliency indicators are usually used to enhance the response of infrared small targets. For example, a weighted saliency operator is constructed based on local mean and variance, or a local mean square difference is used to reflect energy distribution, which can highlight abnormal points in the image to a certain extent. However, such methods generally rely on manually selected fixed windows and thresholds, have poor adaptability, are sensitive to background intensity changes or noise, and are difficult to maintain detection stability in complex scenes. With the rapid development of deep learning technology, more and more research attempts to use convolutional neural networks (CNN) to automatically learn features, extract multi-scale features of small targets through multiple nonlinear mappings, and thus improve the robustness of detection.

[0004] Existing infrared small target detection methods based on deep learning often use an encoder-decoder structure to capture global context information through downsampling, and then gradually recover spatial details through upsampling to realize pixel-level segmentation detection. At the same time, in order to further improve the sensitivity of the network to small targets, some works introduce attention mechanisms, such as channel attention (SENet) that adjusts the importance of each channel to enhance the discriminative ability, and spatial attention (CBAM) that generates weights in the spatial dimension to highlight the region of interest. However, these attention modules usually only use the statistical information of the current branch features to generate weights, and fail to fully consider the prior complementary nature between different modalities.

[0005] On the other hand, considering the limitations of single modal, some studies have proposed multi-modal infrared detection schemes. The typical approach is to splice multi-modal information in the input layer or shallow features by extracting multiple saliency modalities from the original image, such as local energy map reflecting local gray level fluctuation, edge map highlighting spatial gradient change, and contrast map emphasizing the brightness ratio between different regions. Although this method makes use of the complementary information of multi-modal data to some extent, the fusion method is often too simple, such as directly splicing or layer-by-layer addition, without fully exploiting the internal relationship between modalities in the feature encoding and enhancement stage.

[0006] In addition, some existing multi-modal neural networks attempt to use simple feature-level splicing or averaging when fusing deep features, which can easily lead to dilution of key information in the fusion, or inability to effectively highlight the synergistic discriminative features provided by different modalities at the same spatial location, thereby making the detection response insufficient for small targets. At the same time, the single-modal attention module cannot utilize the auxiliary prior provided by other modalities in the same area, resulting in ineffective guidance of attention weights when small targets are not obvious in a certain modality, which can easily cause missed detection or false detection.

[0007] Therefore, the existing multi-modal infrared small target detection technology has the following shortcomings: first, there is a lack of mechanism to explicitly introduce cross-modal mutual guidance information in the feature extraction and residual enhancement stage, which cannot utilize the statistical and spatial consistency between different modalities to guide feature enhancement during encoding; second, current attention mechanisms mostly only adaptively adjust feature responses within a single modality, lacking cross-modal interaction guidance, making it difficult to fully leverage the advantages of multi-modal information in suppressing background clutter and enhancing weak small targets; third, the deep feature fusion method is too rough, and cannot effectively couple and dynamically weight multi-modal features through fine-grained multi-head cross-attention. Therefore, there is an urgent need for a residual enhancement module that can utilize cross-modal channel and spatial information guidance in an encoder-decoder structure, and an efficient fusion module combined with multi-head cross-attention, to fully exploit the complementarity and consistency between multi-modalities and effectively improve the detection accuracy and robustness of weak small targets in infrared images. SUMMARY

[0008] 1. A multi-modal guided infrared small target detection algorithm, comprising the following steps:

[0009] Step S1: using a multi-modal extraction module (Multi-modal Extraction Module) to simultaneously generate a local energy map, a contrast map and a gray Figure Three map from the input infrared image, wherein the local energy map is obtained by calculating the local mean square error of the image and mapping through the hyperbolic tangent function, and the energy map F e The calculation form is:

[0010]

[0011] where μ2(x, y) represents the local mean value of the window centered at pixel (x, y), μ(x, y) 2 is the square mean value of the corresponding window, to avoid the introduction of a small constant by the square root of a negative number;

[0012] to maintain the complete physical thermal radiation information, the original infrared gray image F b is directly taken as the second modal input, that is:

[0013] F b = I

[0014] where I is the model input;

[0015] In addition, in view of the fact that the infrared small target is often weak in local contrast, an adaptive transformation based on local statistical characteristics is introduced in this paper to enhance the perceived contrast, and an adaptive contrast image F c is obtained, which is calculated as:

[0016] F c (x, y) = I(x, y) γ(x,y) ,

[0017]

[0018] where σ is the local standard deviation centered at pixel (x, y), and γ is dynamically adjusted according to the local mean and variance;

[0019] Through the above multi-modal feature construction, the infrared image is explicitly decoupled into three complementary feature spaces F e , F b , and F c .

[0020] Step S2: The above three kinds of modal feature maps are respectively input into a neural network with an encoder-decoder structure, the encoder adopts a four-layer structure, and the encoder layers all use two stacked cross-modal guide residual blocks to extract and enhance the local context and multi-modal mutual guidance information; the module is mainly used to fully mine the complementary and consistent features of different modalities in small target detection, and by integrating cross-modal channel attention and spatial attention mechanisms, the collaborative representation of multi-modal information in local regions is enhanced,

[0021] Step S3: After the encoder completes all levels of downsampling, the high-level features output by each modality branch are input into a multi-modal cross fusion module. The module realizes fine-grained feature interaction between modalities through a multi-head cross attention mechanism with each modality as a query and the remaining modalities as key values, and performs token mixing using a depth separable convolution and a point-wise convolution. The different modality information is dynamically weighted and summed by combining global adaptive modality weights;

[0022] Step S4: The decoder adopts a four-layer structure. Each layer of the decoder also uses two stacked cross-modal guidance residual modules to extract and enhance local context and multi-modal mutual guidance information. The formula of the cross-modal guidance residual module is shown in step 2.

[0023] Step S5: Finally, a standard anchor-free detection head is used. The network is trained using standard stochastic gradient descent (SGD) with a momentum of 0.937, 300 batches per training period, a batch size of 4, a weight decay of 0.0005, and an output that outputs the final prediction results and the training process.

[0024] Step S6: Using the network weights obtained in step S5, input the infrared and visible light remote sensing images to be detected to obtain the prediction results.

[0025] Preferably, the specific process of step 2 is as follows:

[0026] F = BN2(W2 * δ(BN1(W1 * X)))

[0027] Where W1 and W2 are 3x3 convolution kernel weights, BN1 and BN2 are batch normalization (BatchNorm), and δ(·) is a ReLU activation function.

[0028] After obtaining the convolution feature F, a cross-modal channel attention mechanism is used for channel-by-channel weighting. First, F and the guide feature are concatenated in the channel dimension:

[0029] F c = Concat(F, G1, G2)

[0030] Where Concat(·) is channel dimension concatenation, and G1 and G2 are features of the other two modalities.

[0031] Then, global average pooling (GAP) and global maximum pooling (GMP) are used to extract their global statistical information and add them together, and a two-layer $1\times1$ convolution (bottleneck structure) is used to generate channel attention weights:

[0032] F s = GAP(F c ) + GMP(Fc )

[0033] M c = sig(W4 x delta(W3 x F s ))

[0034] where F s is the aggregated feature, and W3, W4 are 1x1 convolution kernel weights, and sig(·) represents the SiLU activation function.

[0035] The weight M c acts on the convolution output feature F channel by channel to obtain the channel enhancement result:

[0036] F′ b,c,h,w = M c [b,c,0,0] x F[b,c,h,w]

[0037] In order to further tap the consistency of multi-modal in spatial position, the channel mean and maximum feature maps of F, G1, and G2 are calculated and spliced as:

[0038]

[0039] Then, a 7x7 convolution and a Sigmoid activation are used to generate spatial attention weights, and a pixel-by-pixel broadcast multiplication is used to obtain spatial enhancement features:

[0040]

[0041] F″ b,c,h,w = M s [b,c,0,0] x F′[b,c,h,w]

[0042] where W s is a 7x7 convolution kernel; the weight M s is used to enhance the feature pixel by pixel; finally, the output is obtained through residual connection and ReLU activation:

[0043] F out = delta(F″+Proj(X))

[0044] where Proj(X) is an identity mapping to ensure that the number of channels matches F″.

[0045] Preferably, the step S3 has the following specific process: for the input energy map feature F e , the grayscale feature F b , and the contrast feature F c , a multi-head cross-modal attention mechanism is used to construct the explicit guidance relationship between modalities, which is specifically represented as

[0046] F e ′ = MHCAeb (F e ,F b )+MHCA ec (F e ,F c ),

[0047] F b ′=MHCA be (F b ,F e )+MHCA bc (F b ,F c ),

[0048] F c ′=MHCA ce (F c ,F e )+MHCA cb (F c ,F b )

[0049] where MHCA ij (F i ,F j ) denotes updating modality F j guided by modality F i , thus explicitly capturing the interaction dependency between different modalities and enhancing the saliency response of small target regions;

[0050] Subsequently, the above three features updated by guidance are mixed in space and channel, and local spatial token mixing is realized by using a channel-wise convolution (DWConv), and then the interaction between channels is realized by a point-wise convolution (PWConv) to form a mixed feature:

[0051] F mix =σ(PWConv(DWConv(F e ′+F b ′+F c ′)))

[0052] where DWConv(·) and PWConv(·) correspond to channel-wise convolution and point-wise convolution operations, respectively, which can effectively promote the spatial-channel coupling of different modalities;

[0053] In order to further adaptively integrate the discriminative ability of the three modalities, we use global pooling and Softmax normalization mechanism to calculate the dynamic weight of each modality, and accordingly adjust the three features:

[0054]

[0055] where GAP(·) denotes the global average pooling operation, W g is a 1x1 convolution for generating the modality weight vector, w e b c correspond to the adaptive coefficients of the energy, gray, and contrast modalities, respectively, and denotes the element-wise weighting.

[0056] Finally, a residual fusion layer containing convolution, batch normalization, and SiLU activation is used to further align the modality distribution and enhance feature consistency, outputting the final fused feature:

[0057] F out = δ(BN(W f * F weighted ))

[0058] where W f is the convolution kernel weight, and BN(·) denotes the batch normalization operation.

[0059] Preferably, the energy map is obtained by calculating the squared difference between the local mean and the squared mean, taking the square root, and then performing nonlinear mapping using the hyperbolic tangent function.

[0060] Preferably, the contrast map is obtained by calculating the ratio of the local standard deviation to the mean of the input image, adaptively adjusting the Gamma function exponent based on the ratio, and then performing a power transformation on the input image.

[0061] Preferably, in the calculation of the cross-modality channel attention, the current modality feature and the two guided modality features are concatenated in the channel dimension, and then the sum of the global average pooling and the maximum pooling is calculated. Then, two layers of point-by-point convolution and nonlinear activation are used to generate channel attention weights, which are then multiplied by the local convolution output feature in each channel. The simplified formula is as follows:

[0062] F′ = CAM c (F, G1, G2) ⊙ F + SAM s (F, G1, G2) ⊙ F

[0063] F out = δ(F′ + Proj(X)).

[0064] Preferably, the cross-modality spatial attention is obtained by taking the average and maximum values of the current modality feature and the two guided modality features in the channel, concatenating them in the channel dimension, and then using a 7x7 convolution and Sigmoid activation to generate spatial attention weights, which are then multiplied by the input feature pixel by pixel. The simplified formula is as follows:

[0065] F e ′, F b ′, F c ​​= MHCA (F e , b , c )

[0066] F mix = DWConv (PWConv (F e + F b + F c ))

[0067]

[0068] Preferably, the method is suitable for the detection task of small targets in infrared images, and improves the detection accuracy and robustness.

[0069] Compared with the prior art, the present application has the following beneficial effects

[0070] The present application generates a local energy map, a contrast map and a gray Figure Three Complementary modal features are extracted from the input infrared image by a multi-modal extraction module, which reflects local gray fluctuation, regional contrast relationship and original brightness information, respectively. Then the multi-modal features are input into a neural network with an encoder-decoder structure, and a cross-modal guiding residual module is introduced at each layer of the encoder and decoder to explicitly integrate the channel and spatial mutual induction information between modalities, enhance the weak small target features. After the encoder completes multi-level feature extraction, the cross-modal cross-fusion module uses multi-group multi-head cross-attention mechanism to interact between different modalities in fine granularity, and combines depth separable convolution and pointwise convolution to realize efficient token mixing, and then performs dynamic fusion through adaptive modal weight, which significantly improves the detection accuracy and robustness of weak small targets. BRIEF DESCRIPTION OF DRAWINGS

[0071] Figure 1 The flowchart of the infrared small target detection algorithm based on multi-modal neural network feature fusion described in the present application.

[0072] Figure 2 An example of the infrared small target gray image for detection in the present application;

[0073] Figure 3 An example of the strong contrast of the multi-modal extraction module of the present application; Figure One An example of the local energy of the multi-modal extraction module of the present application;

[0074] Figure 4 An example of the local energy of the multi-modal extraction module of the present application; Figure One An example of the local energy of the multi-modal extraction module of the present application;

[0075] Figure 5The whole architecture schematic diagram of the algorithm of the application shows the whole process from image input to detection output, including a backbone network, a multi-modal extraction module, and a cross-modal attention fusion module.

[0076] Figure 6 The structure diagram of the cross-modal guiding residual module;

[0077] Figure 7 The structure diagram of the multi-modal cross fusion module;

[0078] Figure 8 The result finally obtained by using the trained weight to predict the graph in the application. DETAILED DESCRIPTION

[0079] In order to make the technical means, creative features, purposes and effects achieved by the application easy to understand, the application will be further described below in combination with specific embodiments.

[0080] The multi-modal infrared small target detection method flow chart of the application is shown in Figure 1 , and specifically includes the following steps,

[0081] Step 1: using a multi-modal extraction module (Multi-modal Extraction Module) to simultaneously generate a local energy map, a contrast map and a gray Figure Three modal feature map from the input infrared image, wherein the local energy map is obtained by calculating the local mean square difference of the image and mapping through the hyperbolic tangent function, the energy map F e is calculated as follows:

[0082]

[0083] Wherein, μ2(x,y) represents the local mean value with the pixel (x,y) as the center window, μ(x,y) 2 is the square mean value of the corresponding window, and ò is a small constant introduced to avoid negative numbers in square roots. This formula can explicitly highlight the area with large gray value variance, thereby enhancing the saliency of small targets in a statistical sense, and effectively suppressing large-area uniform background.

[0084] Secondly, in order to maintain the complete physical thermal radiation information, the original infrared gray image F b is directly taken as the second modal input, that is,

[0085] F b = I

[0086] Wherein I is the model input;

[0087] This way can completely pass on the original gray contrast information, and ensure that the physical brightness difference of small targets is not lost in the early processing stage.

[0088] In addition, in view of the fact that the infrared small target is often weak in local contrast, an adaptive transformation based on local statistical characteristics is introduced to enhance the perceptual contrast, and an adaptive contrast map F is obtained c which is calculated as:

[0089] F c (x,y)=I(x,y) γ(x,y) ,

[0090]

[0091] where σ is the local standard deviation centered at pixel (x, y), and γ is dynamically adjusted according to the local mean and variance, so as to maintain smoothness in the background smooth area and significantly improve the gray difference in the small target mutation area.

[0092] Through the above multi-modal feature construction, the infrared image is explicitly decoupled into three complementary feature spaces: the energy map F e which highlights the areas with significant local fluctuations at the statistical level, helping to detect weak signal targets; the gray map F b which retains the complete thermal radiation intensity distribution, avoiding the loss of the physical contrast of small targets; and the adaptive contrast map F c which amplifies subtle local gray mutations at the perceptual level and improves the separability of weak small targets. The three together provide rich prior information for the subsequent multi-modal cross-modal fusion network, enabling the network to enhance the saliency of small targets from statistical, physical, and perceptual perspectives simultaneously, thereby significantly improving the accuracy and robustness of detection in complex backgrounds.

[0093] Step 2: The above three modal feature maps are respectively input into a neural network with an encoder-decoder structure. The encoder has a four-layer structure, and each layer of the encoder uses two stacked cross-modal guide residual blocks to extract and enhance local context and multi-modal mutual guidance information. This module is mainly used to fully exploit the complementary and consistent features of different modalities (such as infrared energy map, gray map, and contrast map) in small target detection, and to enhance the collaborative representation of multi-modal information in local areas by integrating cross-modal channel attention and spatial attention mechanisms. The specific process is as follows:

[0094] F=BN2(W2*δ(BN1(W1×X)))

[0095] where W1 and W2 are 3x3 convolution kernel weights, BN1 and BN2 are batch normalization (BatchNorm), and δ(·) is the ReLU activation function.

[0096] After obtaining the convolutional feature F, cross-modal channel attention mechanism is used for channel-by-channel weighting. First, F is concatenated with the guide feature in the channel dimension:

[0097] F c = Concat(F, G1, G2)

[0098] where Concat(·) is the concatenation in the channel dimension, and G1, G2 are features of the other two modalities.

[0099] Subsequently, global average pooling (GAP) and global maximum pooling (GMP) are used to extract their global statistical information and add them together, and two layers of 1x1 convolution (bottleneck structure) are used to generate channel attention weights:

[0100] F s = GAP(F c ) + GMP(F c )

[0101] M c = sig(W4x δ(W3x F s ))

[0102] where F s is the aggregated feature, and W3, W4 are 1x1 convolution kernel weights, and sig(·) represents the SiLU activation function.

[0103] The weight M c acts on the convolutional output feature F in a channel-by-channel manner to obtain a channel-enhanced result:

[0104] F′ b,c,h,w = M c [b,c,0,0]·F[b,c,h,w]

[0105] In order to further tap the consistency of multi-modalities in spatial positions, the channel mean and maximum feature maps of F, G1, and G2 are calculated and concatenated as:

[0106]

[0107] Subsequently, a 7x7 convolution and a Sigmoid activation are used to generate spatial attention weights, and a pixel-by-pixel broadcast multiplication is used to obtain spatially enhanced features:

[0108]

[0109] F b ′,′ c,h,w = M s [b,c,0,0]·F′[b,c,h,w]

[0110] where W sis 7x7 convolution kernel. Weight M s is used for pixel-wise feature enhancement. Finally, the output is obtained by residual connection and ReLU activation:

[0111] F out = δ (F" + Proj (X))

[0112] where Proj (X) is an identity mapping to ensure the number of channels matches F".

[0113] Step 3: After the encoder completes all levels of downsampling, the high-level features output by each modality branch are input into the multi-modal cross-fusion module. Through multiple sets of multi-head cross-attention mechanisms with each modality as the query and the remaining modalities as the key-value, fine-grained feature interaction between modalities is achieved, and token mixing is performed using depth separable convolution and point-wise convolution. The different modal information is dynamically weighted and summed by combining global adaptive modal weights. The specific calculation process is as follows.

[0114] First, for the input energy map feature F e , grayscale feature F b and contrast feature F c , we use multi-head cross-modal attention mechanism to construct explicit guidance relationship between modalities, which is specifically represented as

[0115] F e ' = MHCA eb (F e , F b ) + MHCA ec (F e , F c ),

[0116] F b ' = MHCA be (F b , F e ) + MHCA bc (F b , F c ),

[0117] F c ' = MHCA ce (F c , F e ) + MHCA cb (F c , F b )

[0118] where MHCA ij (F i , F j ) represents taking F jThe modalities are guided to update the modalities F by a multi-head query-key value attention mechanism i Thus, the interaction dependency between different modalities is explicitly captured, and the saliency response of small target regions is enhanced.

[0119] Subsequently, the above three guided updated features are mixed in space and channel. Specifically, local spatial token mixing is achieved by using a channel-wise convolution (DWConv), and then point-wise convolution (PWConv) is used to realize the interaction between channels to form a mixed feature:

[0120] F mix =σ(PWConv(DWConv(F e ′+F b ′+F c ′)))

[0121] Where DWConv(·) and PWConv(·) correspond to channel-wise convolution and point-wise convolution operations, respectively, which can effectively promote the spatial-channel coupling of different modalities.

[0122] In order to further adaptively integrate the discriminative ability of the three modalities, we use global pooling and softmax normalization mechanism to calculate the dynamic weight of each modality, and then the three features are weighted according to the dynamic weight:

[0123]

[0124] Where GAP(·) represents the global average pooling operation, W g is a 1×1 convolution used to generate a modality weight vector, w e ,w b ,w c correspond to the adaptive coefficients of energy, gray and contrast three modalities respectively, and ⊙ represents element-wise weighting.

[0125] Finally, a residual fusion layer containing convolution, batch normalization and SiLU activation is used to further align the modality distribution and enhance the feature consistency, and the final fused feature is output:

[0126] F out =δ(BN(W f *F weighted ))

[0127] Where W f is the convolution kernel weight, and BN(·) represents the batch normalization operation. This operation not only maintains the response of the salient region, but also effectively suppresses the false target and background interference in the local context.

[0128] Through the above multi-stage design, the proposed fusion module can model the physical, statistical and perceptual characteristics of the infrared image from multiple angles at the input stage, explicitly enhance the saliency expression of small target regions, and maintain high sensitivity and robustness in the subsequent detection network.

[0129] Step 4: The decoder adopts a four-layer structure, and each layer of the decoder also uses two stacked cross-modal guided residual modules to extract and enhance local context and multi-modal mutual information. The formula of the cross-modal guided residual module is shown in step 2.

[0130] Step 5: Finally, a standard anchor-free detection head is used. The network is trained using standard stochastic gradient descent (SGD) with a momentum of 0.937, 300 batches per training period, a batch size of 4, and a weight decay of 0.0005. The output part finally outputs the prediction results and the training process, as shown in the following examples. Figure 7

[0131] Step 6: Using the network weights obtained in step 5, input the infrared and visible remote sensing images to be detected to obtain the prediction results, as shown in the following examples. Figure 8

[0132] The above examples are only for illustrating the technical concept and characteristics of the present application, and the purpose is to enable those skilled in the art to understand the content of the present application and implement it, and cannot limit the protection scope of the present application. For ordinary skilled persons in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the present application.​​

Claims

1. An infrared small target detection algorithm based on multimodal guidance, characterized in that: Includes the following steps: Step S1: Using the Multi-modal Extraction Module, three modal feature maps—local energy map, contrast map, and grayscale map—are simultaneously generated from the input infrared image. The local energy map is obtained by calculating the squared difference of the local mean of the image and mapping it using the hyperbolic tangent function. This energy map F... e The calculation form is: Where μ2(x,y) represents the local mean of the window centered at pixel (x,y), μ(x,y) 2 This represents the mean square value of the corresponding window. A small constant is introduced to avoid negative square roots; To preserve complete physical thermal radiation information, the original infrared grayscale image F was directly processed. b As the second modal input, that is: F b =I Where I is the model input; Furthermore, considering the characteristic that small infrared targets often have weak local contrast, this paper introduces an adaptive transformation based on local statistical properties to enhance perceived contrast, resulting in an adaptive contrast image F. c The calculation is as follows: F c (x,y)=I(x,y) γ(x,y) , Where σ is the local standard deviation centered at pixel (x,y), and γ is dynamically adjusted based on the local mean and variance; Through the construction of the above multimodal features, the infrared image is explicitly decoupled into three complementary feature spaces F. e F b F c ; Step S2: Input the feature maps of the three modalities mentioned above into a neural network with an encoder-decoder structure. The encoder has a four-layer structure, and each layer of the encoder uses two stacked cross-guided residual blocks to extract and enhance local context and multimodal cross-correlation information. This module is mainly used to fully explore the complementary and consistent features of different modalities in small target detection. By integrating cross-modal channel attention and spatial attention mechanisms, it enhances the collaborative representation of multimodal information in local regions. Step S3: After the encoder completes downsampling at all levels, the high-level features output by each modal branch are input into the Multi-modal Cross Fusion Module. Fine-grained feature interaction between modalities is achieved through multiple sets of multi-head cross attention mechanisms with each modality as the query and the other modalities as the key. Token mixing is performed using depthwise separable convolution and pointwise convolution, and global adaptive modal weights are combined to dynamically weight and sum the information of different modalities. Step S4: The decoder adopts a four-layer structure. Each layer of the decoder also uses two stacked cross-modal guided residual modules to extract and enhance local context and multimodal intermodal information. The formula for the cross-modal guided residual module is shown in Step 2. Step S5: Finally, a standard anchorless detection head is used. This network is trained using standard stochastic gradient descent (SGD) with a momentum of 0.937, 300 batches per training epoch, a batch size of 4, and a weight decay of 0.0005. The output section shows the final prediction results and the training process. Step S6: Using the network weights obtained in step S5, input the infrared and visible light remote sensing images to be detected to obtain the prediction results.

2. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that, The specific process for step 2 is as follows: F = BN2(W2*δ(BN1(W1×X))) Where W1 and W2 are the weights of the 3×3 convolution kernel, BN1 and BN2 are batch normalization (BatchNorm), and δ(·) is the ReLU activation function; After obtaining the convolutional feature F, it is weighted channel-wise using a cross-modal channel attention mechanism; firstly, F is concatenated with the guiding feature along the channel dimension: F c =Concat(F,G1,G2) Where Concat(·) is the concatenation of channel dimensions, and G1 and G2 are the features of the other two modalities; Subsequently, global average pooling (GAP) and global max pooling (GMP) are used to extract global statistics and sum them. Then, two layers of 1×1 convolutions (bottleneck structure) are used to generate channel attention weights. F s =GAP(F c )+GMP(F c ) M c =sig(W4×δ(W3×F s )) Among them, F s W3 and W4 are the aggregated features, while W3 and W4 are the weights of the 1×1 convolution kernel, and sig(·) represents the SiLU activation function. Weight M c Channel-by-channel application of the convolutional output feature F yields the channel enhancement result: F′ b,c,h,w =M c [b,c,0,0]·F[b,c,h,w] To further explore the spatial consistency of multimodal data, the channel mean and maximum feature maps of F, G1, and G2 were calculated and concatenated as follows: Spatial attention weights are then generated using 7×7 convolution and Sigmoid activation, and spatial enhancement features are obtained through pixel-wise broadcast multiplication. F″ b,c,h,w =M s [b,c,0,0]·F′[b,c,h,w] Among them W s It has a 7×7 convolution kernel; weights M s Used for pixel-wise feature enhancement; finally, the output is obtained through residual connections and ReLU activation: F out =δ(F″+Proj(X)) Proj(X) is an identity mapping used to ensure that the number of channels matches that of F″.

3. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that: The specific process of step S3 is as follows: For the input energy map feature F e Grayscale image features F b Compared with feature F of the comparison image c We employ a multi-head cross-modal attention mechanism to construct explicit guidance relationships between modalities, specifically represented as follows: F e ′=MHCA eb (F e ,F b )+MHCA ec (F e ,F c ), F b ′=MHCA be (F b ,F e )+MHCA bc (F b ,F c ), F c ′=MHCA ce (F c ,F e )+MHCA cb (F c ,F b ) Among them MHCA ij (F i ,F j ) indicates F j Modality F is updated via a multi-head query-key-value attention mechanism, guided by the modality. i This allows for the explicit capture of interaction dependencies between different modalities, enhancing the salience response in small target regions; Subsequently, the three updated features are spatially and channel-wise blended. Specifically, channel-wise convolution (DWConv) is used to achieve local spatial token mixing, and pointwise convolution (PWConv) is used to achieve inter-channel interaction, forming blended features. F mix =σ(PWConv(DWConv(F e ′+F b ′+F c ′))) Among them, DWConv(·) and PWConv(·) correspond to channel-wise convolution and point-wise convolution operations, respectively, which can effectively promote spatial-channel coupling of different modalities; To further adaptively integrate the discriminative capabilities of the three modalities, we utilize global pooling and Softmax normalization mechanisms to calculate the dynamic weights of each modality, and then explicitly weight and fuse the three features accordingly: Where GAP(·) represents the global average pooling operation, W g A 1×1 convolution is used to generate the modality weight vector, w e ,w b ,w c These correspond to the adaptive coefficients for the three modes of energy, grayscale, and contrast, respectively, with ⊙ indicating element-wise weighting. Finally, a residual fusion layer incorporating convolution, batch normalization, and SiLU activation is used to further align modal distributions and enhance feature consistency, outputting the final fused features: F out =δ(BN(W f *F weighted )) Among them W f represents the convolution kernel weights, and BN(·) denotes the batch normalization operation.

4. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that: The energy map is obtained by calculating the local mean and the squared mean, then taking the square root of the difference between the squares, and finally applying a nonlinear mapping using the hyperbolic tangent function.

5. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that; The comparison image is obtained by calculating the ratio of the local standard deviation to the mean of the input image, adaptively adjusting the exponent of the Gamma function based on this ratio, and then performing a power transformation on the input image.

6. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that: When calculating cross-modal guided residual module, the current modal feature and the two guided modal features are concatenated along the channel dimension. After being summed by global average pooling and max pooling respectively, channel attention weights are generated through two layers of pointwise convolution and nonlinear activation. Then, the weights are multiplied by local convolutions channel by channel to output features. The simplified formula is as follows: F′=CAM c (F,G1,G2)⊙F+SAM s (F,G1,G2)⊙F F out = δ(F′+Proj(X)).

7. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that: The cross-modal spatial attention is generated by averaging and maximizing the current modal features and the two guiding modal features in each channel, concatenating them in the channel dimension, using 7×7 convolution and Sigmoid activation to generate spatial attention weights, and multiplying them by the input features pixel by pixel. The simplified formula is as follows: F e ′,F b ′,F c ′=MHCA(F e ,F b ,F c ) F mix =DWConv(PWConv(F e ′+F b ′+F c ′)) 8. The infrared small target detection algorithm based on multimodal guidance according to claim 1, characterized in that: The method is applicable to the detection of small targets in infrared images and improves detection accuracy and robustness.

Citation Information

Cited By

  • Optical and SAR multi-level cross-modal fusion flood inundation monitoring method and device

    CN121482612A

  • A fire warning method, device, equipment and medium based on image spatial domain

    CN122454464A

  • A fire warning method, device, equipment and medium based on image spatial domain

    CN122454464B