Infrared small target detection method based on conditional diffusion model

Through the multi-scale feature extraction and reverse diffusion process of the conditional diffusion model, the missed detection and missed detection problems in infrared small object detection are solved, and the detection effect with high accuracy and high robustness is achieved, reducing the cost of data acquisition.

CN120355904APending Publication Date: 2025-07-22ZHEJIANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510500705.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing infrared small object detection methods are prone to missed detection and missed detection in complex backgrounds, and the data acquisition cost is high, and lack the ability to model the target generation process.

Method used

The infrared small object detection method based on the conditional diffusion model is adopted, through the multi-scale feature extraction and reverse diffusion process, combined with the conditional perception encoding module and feature decoding module, the data enhancement strategy and mean square error loss function are used for supervision and training to generate a high-precision object mask.

Benefits of technology

Achieve high-precision and high-rootability infrared small object detection in the context of low signal-to-noise ratio and complexity, reducing dependence on large-scale annotation data and improving the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355904A_ABST
    Figure CN120355904A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target detection method based on a conditional diffusion model. Specifically, the conditional diffusion model comprises a conditional perception coding module, a conditional guidance feature fusion coding module and a feature decoding module. According to the method, multi-scale feature extraction and fusion are carried out on an infrared image through a condition-guided condition perception coding module to generate condition features, feature fusion is carried out on the condition features and a noise mask of a time step t, and a model is guided to directionally optimize a target mask. In the training process of the conditional diffusion model, a mean square error loss function based on a time step t is adopted to supervise the accuracy of mask generation, and model parameters are dynamically adjusted through an Adam optimizer to minimize noise prediction errors. In the target detection stage, denoising is carried out step by step from the initial state of Gaussian noise through Markov chain iteration in the reverse diffusion process, and finally a high-precision target mask is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an infrared small target detection method based on a conditional diffusion model. Background Art

[0002] In the field of infrared imaging technology, small target detection is a core task in applications such as complex environment monitoring, military reconnaissance, and industrial inspection. Traditional infrared small target detection methods mainly rely on manually designed features (such as local contrast enhancement, morphological filtering, etc.) and threshold segmentation techniques. Such methods have significant limitations in scenarios where the target size is extremely small, the background is complex, or the signal-to-noise ratio is extremely low.

[0003] In recent years, discriminative models based on deep learning have gradually become the mainstream methods in the field of infrared small target detection due to their powerful feature expression capabilities. Such models can distinguish small targets from relatively complex backgrounds through end-to-end learning, but still face significant challenges in practical applications. First, the performance of the model highly depends on large-scale labeled data, and infrared small targets require fine annotation by professionals, resulting in high data acquisition costs and scarce samples. Second, existing methods are mostly based on static discriminative strategies and lack the ability to model the target generation process, prone to missed detection and misdetection problems in complex background scenarios. Summary of the Invention

[0004] Aiming at the defects existing in the above background art, the present invention provides an infrared small target detection method based on a conditional diffusion model. This method is implemented based on a conditional diffusion model. During the model training process, various data augmentation strategies can be adopted to expand the training data, significantly improving the generalization ability of the model in different scenarios. The model uses a residual nested dense block to extract multi-scale features of infrared images and dynamically adjusts the feature weights through a time step embedding vector to accurately capture the spatial details of small targets. In the process of generating a target mask based on the noise prediction result, based on the Markov chain characteristics of the reverse diffusion process, the target mask is gradually restored through multi-step denoising. The conditional diffusion model uses the multi-scale features of the conditional perception encoding module to guide the network to dynamically suppress background noise. At the same time, the shallow detail features of the encoder are fused with the deep semantic features of the decoder through skip connections, avoiding the loss of target information caused by multiple downsamplings in traditional methods and ensuring the integrity of the small target structure. The method realizes high-precision and high-robustness infrared small target detection in low signal-to-noise ratio, complex background, and small sample scenarios through the collaborative optimization of the generative framework and discriminative features.

[0005] The present invention is implemented by adopting the following technical solutions:

[0006] An infrared small target detection method based on a conditional diffusion model. The method is implemented based on a conditional diffusion model, and the conditional diffusion model is implemented using an end-to-end convolutional neural network. At each time step t, a supervised training method is used to guide the update of network parameters, giving full play to the unique advantages of the conditional diffusion model in generating fine features and improving the performance of small target detection.

[0007] The construction method of the conditional diffusion model includes the following steps:

[0008] (1) Construct a dataset: Use a variety of data augmentation methods to expand the dataset and improve the generalization performance of the conditional diffusion model in different scenarios;

[0009] (2) Build a conditional diffusion model: The network of the conditional diffusion model uses a fully convolutional neural network, which can adapt to inputs of different sizes without retraining. Specifically, it can be divided into three modules, namely a conditional perception encoding module, a conditional-guided feature fusion encoding module, and a feature decoding module;

[0010] (3) Use the dataset obtained in step (1) to train the model built in step (2): Use the mean squared error loss function to perform supervised training on the state at each time step t, guiding the network to adaptively learn the corresponding noise distribution based on the information at time step t; Update the network parameters of the model through the Adam optimization algorithm, combined with a stepped learning rate decay strategy, to improve the convergence speed and detection performance of the model (minimize the noise prediction error).

[0011] The infrared small target detection method based on the conditional diffusion model specifically includes the following steps:

[0012] (1) Process the infrared image using a variety of data augmentation methods, and use the processed infrared image as the conditional input of the conditional diffusion model;

[0013] (2) Input the processed infrared image into the conditional diffusion model; First, use the conditional perception encoding module to perform multi-scale feature extraction and fusion on the processed infrared image to obtain conditional features; Then, use the conditional-guided feature fusion encoding module to fuse the conditional features and the noise mask obtained based on the forward diffusion process to obtain conditional fusion features of different scales; Finally, use the feature decoding module to fuse the conditional fusion features of different scales, and then obtain the noise prediction result after processing by a group normalization layer, a SiLU activation function, and a 3×3 convolutional layer;

[0014] (3) Based on the obtained noise prediction result, gradually recover from Gaussian noise to a clear image by repeatedly performing the reverse diffusion process through multiple iterations, thereby generating an accurate target mask.

[0015] In the above technical solution, further, the data augmentation method includes: scale transformation, image translation, random cropping, and image flipping. Perform a normalization operation on the infrared image after data augmentation (which can improve the stability and convergence speed of model training), specifically expressed as:

[0016]

[0017] Among them, represents the normalized output of the i-th channel, and X (i) is the original input of the i-th channel, μ (i) is the statistical mean of the i-th channel, and σ (i) is the statistical standard deviation of the i-th channel.

[0018] Further, the conditional perception encoding module is constructed based on a composite residual dense architecture, and is used to perform multi-scale feature extraction and fusion on the input infrared image to obtain conditional features. The conditional perception encoding module contains two cascaded Residual-in-Residual Dense Blocks (RRDBs), where each RRDB is composed of four 3×3 convolutional layers connected in sequence, and a Leaky ReLU activation function is embedded between layers to enhance the non-linear expression ability of features; the output of the convolutional layer is concatenated with the features of all previous levels along the channel dimension through a dense connection mechanism to achieve cross-level feature reuse and information transfer; the RRDB further performs element-wise superposition of its initial input and the final output of the four 3×3 convolutional layers according to certain weights (to obtain conditional features) through a global residual path, forming a residual nested structure to optimize the gradient propagation path and suppress the deep network degradation effect, thereby improving the stability and efficiency of feature extraction.

[0019] Further, the conditional-guided feature fusion encoding module fuses the conditional features and the noise mask obtained based on the forward diffusion process to obtain conditional fusion features of different scales, specifically as follows:

[0020] Based on the forward diffusion process, the noise mask x is obtained after t steps of diffusion from the initial state x0 t is defined as:

[0021]

[0022] α t = 1 - β t , where β t represents the preset variance decay coefficient of the time step t, and ∈ represents the noise obeying the standard normal distribution.

[0023] The conditional-guided feature fusion encoding module first passes the noise mask x through a 3×3 convolutional layert Map to the conditional features output by the conditional perception encoding module with the same channel dimension to generate aligned features where H, W, and D are the height, width, and dimension of the feature map respectively; then, perform an element-wise addition operation on the conditional feature C and the aligned feature to achieve preliminary fusion and obtain a fused feature The fused feature and the embedding vector at time step t undergo deep feature interaction through an n-level encoder module to obtain conditional fusion features at different scales; each level of the encoder module includes a residual block.

[0024] The operation of using the fused feature and the embedding vector at time step t to perform deep feature interaction through an n-level encoder module to obtain conditional fusion features at different scales specifically includes the following steps:

[0025] S1: Map the time step t to an initial embedding vector e using sine encoding t ′ , and then generate a final embedding vector e through two linear layers (with a SiLU function in the middle) t , which can be specifically expressed as:

[0026] e t = Linear(SiLU(Linear(e t ′ )))

[0027] S2: In the first-level encoder module, the fused feature sequentially passes through a non-downsampling 3x3 convolution and a 7x7 convolution to output the conditional fusion feature corresponding to the first-level encoder module Take as the input of the second-level encoder module;

[0028] S3: In the second-level to the n-level encoder modules, the method for obtaining conditional fusion features at different scales is as follows: For the k-th level encoder module, the embedding vector e t is linearly projected to the same feature dimension as in each residual block and then element-wise added to the input of the k-th level encoder module to generate a spatio-temporal fusion feature, which is then integrated through a 3x3 convolutional layer and output as an intermediate feature through a residual connection. Subsequently, the intermediate feature is spatially downsampled and channel-expanded through a 3x3 convolutional layer to generate the conditional fusion feature corresponding to the k-th level encoder module The conditional fusion feature corresponding to the k-th level encoder module As the input of the (k + 1)-th level encoder module, when k = n, the conditional fusion feature of the last scale is output Among them, The acquisition method of can be specifically expressed as:

[0029]

[0030] Among them, the conditional fusion features output by each level of encoder module have different scales.

[0031] Furthermore, the feature decoding module fuses the conditional fusion features of different scales, and then obtains the noise prediction result after being processed by a group normalization layer, a SiLU activation function, and a 3×3 convolutional layer. The specific method is as follows:

[0032] The conditional fusion feature of the last scale output by the conditional guidance feature fusion and encoding module After being processed by a 3×3 depthwise separable convolution, a Leaky ReLU activation, and a 1×1 pointwise convolution in sequence to achieve cross-channel information fusion, a high-level semantic feature is obtained; then it is input into the feature decoding module.

[0033] The feature decoding module adopts a reverse architecture symmetric to the n-level encoder module, including n-level cascaded decoder modules. The feature decoding module specifically performs the following operations:

[0034] S1: First, the high-level semantic feature is concatenated with the conditional fusion feature output by the symmetric-level encoder module along the channel dimension to generate a cross-layer fusion feature; after the spatial resolution of the cross-layer fusion feature is expanded by bicubic interpolation upsampling of the first-level decoder module, feature interaction and enhancement are realized through a residual block;

[0035] S2: In the second-level to the (n - 1)-th level decoder modules, the output of the residual block of the previous-level decoder module is first concatenated with the conditional fusion feature of the corresponding scale output by the symmetric-level encoder module along the channel dimension to generate a cross-layer fusion feature; after the spatial resolution of the cross-layer fusion feature is expanded by bicubic interpolation upsampling of the current-level decoder module, feature interaction and enhancement are realized through a residual block;

[0036] S3: In the n-th level decoder module, the upsampling module is cancelled, and the cross-layer fusion feature of the (n - 1)-th level decoder module after skip connection is directly input into a residual block and a 1x1 convolutional layer for feature interaction, and its output passes through a group normalization layer and a SiLU activation function in sequence, and finally the noise prediction result ∈ is output by a 3×3 convolutional layer pred .

[0037] Furthermore, during the training process of the conditional diffusion model, the mean squared error loss function is used to supervise the training of the state at each time step t, guiding the model to adaptively learn the corresponding noise distribution based on the information at time step t; and the network parameters are updated through the Adam optimization algorithm, combined with the stepped learning rate decay strategy, to improve the convergence speed and detection performance of the model.

[0038] Furthermore, the derivation process of the mean squared error loss function is as follows:

[0039] Based on the Markov chain assumption of the reverse diffusion process, given the noise mask x at time step t t , x t-1 's probability distribution p θ (x t-1 |x t ) can be expressed as:

[0040] p θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , C, t), Σ θ (x t , C, t))

[0041] where μ θ (·, ·, ·), Σ θ (·, ·, ·) are the mean and variance parameters predicted by the network respectively; C represents the conditional input; based on Bayes' theorem and the forward diffusion process assumption, given the initial state x0 and the noise mask x at time step t t , x t-1 's probability distribution can be expressed as:

[0042]

[0043] where I represents the identity matrix; represents the posterior mean and variance given x0 and x t , x t-1 , and the specific definitions are:

[0044]

[0045] and α t = 1 - β t , where ∈ represents noise following the standard normal distribution; β t represents the preset variance decay coefficient at time step t.

[0046] The training objective of the conditional diffusion model is to minimize p θ (xt-1 |x t ) and q(x t-1 |x t , x0), for simplicity of calculation, set the variance and the mean μ θ (x t , C, t) is reconstructed as:

[0047]

[0048] where ∈ θ (x t , C, t) is the noise prediction value output by the model, related to the time step t, the noise mask x t and the conditional input C.

[0049] During the forward diffusion process, the noise mask x t can be expressed as a linear combination of the initial state x0:

[0050]

[0051] Based on this, the following mean squared error loss function is defined to supervise the model training:

[0052]

[0053] Furthermore, based on the obtained noise prediction results in step (3), by repeatedly performing the reverse diffusion process through multiple iterations, gradually recovering from Gaussian noise to a clear image, thereby generating an accurate target mask, which specifically includes the following steps:

[0054] (1) Randomly sample the initial noise from the Gaussian distribution where the total number of time steps is T; according to the reverse diffusion process, the noise mask x at time step t - 1 t-1 is generated from the noise mask x at time step t t and its calculation formula is:

[0055]

[0056] where, ∈ θ (x t , C, t) is the noise prediction result output by the model at time step t; represents the standard Gaussian noise;

[0057] (2) Starting from time step t = T, repeatedly execute step (1) to gradually generate x t-1 until t = 1, and finally output the target mask.

[0058] The advantages of the present invention are:

[0059] The infrared small target detection method based on the conditional diffusion model proposed by the present invention gradually refines the target mask through the Markov chain characteristics of the reverse diffusion process, and combines the multi-scale feature guidance of the conditional perception coding module to significantly improve the positioning accuracy and detail recovery ability of small targets, effectively solving the problem of information loss caused by downsampling in traditional methods; it learns the target distribution through the noise prediction task, reduces the dependence on large-scale labeled data, and realizes cross-scene generalization by using the conditional guidance mechanism and various data augmentations. Through the deep integration of the generative framework and discriminative guidance, this method achieves high-precision and high-generalization infrared small target detection in complex backgrounds and small-sample scenarios, combining theoretical rigor and industrial practicality. Description of the Drawings

[0060] Figure 1 It is the training framework for infrared small target detection based on the conditional diffusion model.

[0061] Figure 2 It is the RRDB structure of the conditional perception coding module.

[0062] Figure 3 It is the encoder and decoder structures in the training framework.

[0063] Figure 4 It is the comparison of experimental results of different detection methods. Detailed Implementation Manner

[0064] The following further describes the detailed implementation manner of the present invention with reference to the attached drawings.

[0065] The present invention provides an infrared small target detection method based on the conditional diffusion model. As Figure 1 shown in the training method of the conditional diffusion model, the specific steps are as follows:

[0066] (1) First, perform data augmentation operations such as random scaling, image translation, random cropping, and image flipping on the input infrared image to simulate the target morphology and background changes in different scenarios; the enhanced image is processed by channel normalization to stabilize the input distribution and generate a standardized input suitable for learning.

[0067] (2) Input the infrared image processed in step (1) into the conditional diffusion model. First, use the conditional perception coding module composed of two-stage cascaded RRDBs to perform multi-scale feature extraction and fusion on the processed infrared image to obtain conditional features. The composition of RRDB is as Figure 2As shown. Each RRDB consists of four sequentially connected 3×3 convolutional layers, with a Leaky ReLU activation function embedded between the layers. The output of each convolutional layer is concatenated with the outputs of all previous convolutional layers along the channel dimension through a dense connection mechanism; the initial input of the RRDB and the final output of the four 3×3 convolutional layers are element-wise superimposed according to certain weights to obtain conditional features, and this operation can preserve the distribution of the original features. The standardized infrared image input finally generates conditional feature C through the conditional perception encoding module, which is used to guide the denoising direction of the diffusion process;

[0068] (3) Use the conditional-guided feature fusion and encoding module to fuse the conditional features and the noise mask obtained based on the forward diffusion process to obtain conditional fusion features of different scales. Based on the forward diffusion process, the initial state x0 is diffused for t steps to obtain the noise mask x t ; The conditional-guided feature fusion and encoding module first passes through a 3x3 convolutional layer to map the noise mask x t to the same channel dimension as the conditional features to generate the aligned feature where H, W, and D are the height, width, and dimension of the feature map respectively; then the conditional feature C and the aligned feature are initially feature-fused by element-wise superposition to generate the fused feature For it is input into an encoder-decoder based on the U-Net architecture for deeper feature fusion to obtain conditional fusion features of different scales. The specific encoder-decoder structure is as Figure 3 shown, and specifically includes 4 levels of encoder modules. First, preliminary feature extraction is performed through 3x3 convolution and 7x7 convolution; then, through multiple downsampling convolutions and residual block processing, the resolution of the feature map is gradually reduced and multi-scale features are extracted. Each of the second to nth level encoder modules includes a residual block, and the embedding vector e t is input into the residual block. The method for obtaining the embedding vector e t is: using sine encoding to map the time step t to the initial embedding vector e t ′ , and then generating the final embedding vector e t through two linear layers with SiLU functions embedded in the middle.。The output of the last encoder module sequentially passes through a 3x3 depthwise separable convolution, a Leaky ReLU activation, and a 1x1 convolution layer to generate high-level semantic features, which serve as the initial input to the decoder. The decoder adopts a reverse architecture symmetric to the encoder. The input features are first concatenated along the channel dimension with the conditional fusion features output by the symmetric-level encoder module to generate cross-layer fusion features. After the cross-layer fusion features are upsampled by bicubic interpolation to expand the spatial resolution, feature interaction and enhancement are achieved through residual blocks. The last decoder module directly inputs the cross-layer fusion features after skip connection into the residual block and the 1x1 convolution layer for feature interaction. Its output sequentially passes through a group normalization layer and a SiLU activation function, and finally the noise prediction result ∈ is output by a 3×3 convolution layer pred 。

[0069] (4) The mean squared error loss function is used to supervise the training of the noise mask at each time step t, guiding the network to adaptively learn the corresponding noise distribution based on the information at time step t. The network parameters are updated through the Adam optimization algorithm, combined with a stepped learning rate decay strategy, to improve the convergence speed and detection performance of the model

[0070] The trained conditional diffusion model described above is used for noise prediction. Based on the obtained noise prediction results, by repeatedly executing the reverse diffusion process, the Gaussian noise is gradually restored to a clear image, thereby generating an accurate target mask

[0071] The experimental comparison results of the present invention with the traditional method RIPT and the advanced algorithm ISNet are as Figure 4 shown. Under the condition of complex background interference, both the traditional RIPT method and the ISNet algorithm have target misdetection phenomena, while the present invention can still maintain accurate target detection performance, verifying the robustness and technical superiority of the present invention in complex environments

[0072] The above is only the specific implementation manner of the present invention, and the scope of the present invention cannot be limited thereby. Equivalent changes made by those of ordinary skill in the art according to this creation, as well as changes well-known to those skilled in the art, should still fall within the scope covered by the present invention

Claims

1. An infrared small target detection method based on a conditional diffusion model, characterized in that, The conditional diffusion model includes a conditional perception encoding module, a condition-guided feature fusion encoding module, and a feature decoding module; the prediction method specifically includes the following steps: (1) Process the infrared image using a variety of data augmentation methods, and use the processed infrared image as the conditional input of the conditional diffusion model; (2) Input the processed infrared image into the conditional diffusion model; First, use the conditional perception encoding module to perform multi-scale feature extraction and fusion on the processed infrared image to obtain conditional features; Then, use the condition-guided feature fusion encoding module to fuse the conditional features and the noise mask obtained based on the forward diffusion process to obtain conditional fusion features of different scales; Finally, use the feature decoding module to fuse the conditional fusion features of different scales, and then obtain the noise prediction result after being processed by a group normalization layer, a SiLU activation function, and a 3×3 convolutional layer; (3) Based on the obtained noise prediction result, gradually restore from Gaussian noise to a clear image by repeatedly executing the reverse diffusion process multiple times, so as to generate an accurate target mask.

2. The infrared small target detection method based on a conditional diffusion model according to claim 1, characterized in that, In the step (1): The data augmentation methods used include: scale transformation, image translation, random cropping, and image flipping; the normalization operation on the processed infrared image is specifically expressed as: Among them, represents the normalized output of the i-th channel, X (i) is the original input of the i-th channel, μ (i) is the statistical mean of the i-th channel, σ (i) is the statistical standard deviation of the i-th channel.

3. The infrared small target detection method based on a conditional diffusion model according to claim 1, characterized in that, Using the conditional perception encoding module to perform multi-scale feature extraction and fusion on the processed infrared image to obtain conditional features, specifically: The conditional perception encoding module adopts a composite residual dense architecture, which includes two cascaded residual nested dense blocks RRDB, where each RRDB is composed of four 3×3 convolutional layers connected in sequence, and a Leaky ReLU activation function is embedded between layers; the output of the convolutional layer is concatenated with the outputs of all previous convolutional layers along the channel dimension through a dense connection mechanism; the initial input of the RRDB and the final output of the four 3×3 convolutional layers are element-wise superimposed according to weights to obtain conditional features.

4. The infrared small target detection method based on a conditional diffusion model according to claim 1, characterized in that, The method of using the condition-guided feature fusion encoding module to fuse the conditional features and the noise mask obtained based on the forward diffusion process to obtain conditional fusion features of different scales is specifically: Based on the forward diffusion process, the initial state x0 is diffused through t steps to obtain the noise mask x t Defined as: Among them, β t represents the variance decay coefficient preset for time step t, and ∈ represents noise following a standard normal distribution; The condition-guided feature fusion and encoding module first maps the noise mask x through a 3×3 convolutional layer t to the same channel dimension as the conditional features output by the condition-aware encoding module to generate aligned features where H, W, and D are the height, width, and dimension of the feature map respectively; then, the conditional features C and the aligned features are preliminarily fused by element-wise addition to obtain fused features The fused features and the embedding vector at time step t undergo deep feature interaction through an n-level encoder module to obtain conditional fusion features at different scales; each encoder module includes a residual block; The fused feature and the embedding vector at time step t are subjected to deep feature interaction through an n-level encoder module to obtain conditional fused features of different scales, which specifically include the following steps: S1: Map the time step t to the initial embedding vector e′ using sine encoding t , and then generate the final embedding vector e through two linear layers with SiLU functions embedded in the middle t , which is specifically expressed as: e t = Linear(SiLU(Linear(e′ t ))) S2: In the first-level encoder module, fuse the features successively pass through a non-downsampling 3x3 convolution and a 7x7 convolution to output the conditional fusion features corresponding to the first-level encoder module Take as the input of the second-level encoder module; S3: In the encoder modules from the second level to the nth level, the method for obtaining the conditional fusion features of different scales is as follows: for the kth encoder module, the embedding vector e t after each residual block is linearly projected to the same feature dimension and then added element-wise to the input of the kth encoder module to generate spatio-temporal fusion features, which are then integrated by a 3x3 convolutional layer and output intermediate features through a residual connection. Subsequently, the intermediate features are spatially downsampled and channel-expanded by a 3x3 convolutional layer to generate the conditional fusion features corresponding to the kth encoder module The conditional fusion features corresponding to the kth encoder module serve as the input of the (k + 1)th encoder module. When k = n, the conditional fusion features of the last scale are output wherein, the acquisition method is specifically expressed as: Among them, the conditional fusion features output by each level of encoder module have different scales.

5. The infrared small target detection method based on the conditional diffusion model according to claim 4, wherein, The method of using the feature decoding module to fuse the conditional fusion features of different scales, and then obtaining the noise prediction result after being processed by a group normalization layer, a SiLU activation function, and a 3×3 convolutional layer is specifically: The conditional fusion feature of the last scale output by the conditional-guided feature fusion and encoding module After being processed by a 3×3 depthwise separable convolution, a Leaky ReLU activation, and a 1×1 pointwise convolution in sequence to achieve cross-channel information fusion, high-level semantic features are obtained; Then input the high-level semantic features into the feature decoding module; The feature decoding module adopts a reverse architecture symmetric to the n-level encoder module, which includes n-level cascaded decoder modules. The feature decoding module specifically performs the following operations: S1: First, concatenate the high-level semantic features with the conditional fusion features output by the symmetric-level encoder module along the channel dimension to generate a cross-layer fusion feature; after the cross-layer fusion feature is upsampled by bicubic interpolation of the first-level decoder module to expand the spatial resolution, feature interaction and enhancement are realized through a residual block; S2: In the decoder modules from the second level to the (n - 1)th level, the output of the residual block in the previous decoder module is first concatenated with the conditional fusion features of the corresponding scale output by the symmetric hierarchical encoder module along the channel dimension to generate cross-layer fusion features; after the spatial resolution of the cross-layer fusion features is expanded by bicubic interpolation upsampling of the decoder module at the current level, feature interaction and enhancement are achieved through the residual block. S3: In the n-th level decoder module, the upsampling module is removed. The cross-layer fusion features of the (n - 1)-th level decoder module after skip connection are directly input into the residual block and the 1x1 convolutional layer for feature interaction. Its output passes through the group normalization layer and the SiLU activation function in sequence, and finally the noise prediction result ∈ is output by the 3×3 convolutional layer pred .

6. The infrared small target detection method based on the conditional diffusion model according to claim 5, characterized in that, During the training process of the conditional diffusion model, the mean squared error loss function is used to supervise the training of the state at each time step t, guiding the model to adaptively learn the corresponding noise distribution based on the information at time step t; and the network parameters are updated through the Adam optimization algorithm, combined with the stepped learning rate decay strategy, to improve the convergence speed and detection performance of the model.

7. The infrared small target detection method based on a conditional diffusion model according to claim 6, wherein The derivation process of the mean squared error loss function is as follows: Based on the Markov chain assumption of the reverse diffusion process, given the noise mask x at time step t t , x t-1 's probability distribution p θ (x t-1 |x t ) is expressed as: p θ (x t-1 |x t ) = N(x t-1 ; μ θ (x t , C, t), Σ θ (x t , C, t)) where μ θ (·, ·, ·), Σ θ (·, ·, ·) are the mean and variance predicted by the model respectively; C represents the conditional input; Based on Bayes' theorem and the forward diffusion process assumption, given the initial state x0 and the noise mask x at time step t t , x t-1 's probability distribution q(x t-1 |x t , x0) is expressed as: where I represents the identity matrix; represent the known x0 and x respectively t , x t-1 's posterior mean and variance, specifically defined as: and α t = 1 - β t , where ∈ represents noise following a standard normal distribution; β t represents the preset variance decay coefficient at time step t; The training objective of the conditional diffusion model is to minimize the difference between p θ (x t-1 |x t ) and q(x t-1 |x t , x0). For simplicity in calculation, set the variance and reconstruct the mean μ θ (x t , C, t) as: where, ∈ θ (x t , C, t) is the predicted noise value output by the model, which is related to the time step t, the noise mask x t and the conditional input C; Noise mask x during the forward diffusion process t Expressed as a linear combination of the initial state x0: Accordingly, the following mean squared error loss function is defined to supervise the model training:

8. The infrared small target detection method based on the conditional diffusion model according to claim 7, wherein, The specific steps of step (3) include the following steps: (1) Randomly sample the initial noise from a Gaussian distribution where the total number of time steps is T; according to the reverse diffusion process, the noise mask x at time step t-1 t-1 is generated from the noise mask x at time step t t and its calculation formula is: where, ∈ θ (x t , C, t) is the noise prediction result output by the model at time step t; represents standard Gaussian noise; (2) Starting from time step t = T, iteratively execute step (1) to gradually generate x t-1 until t = 1, and finally output the target mask.

Citation Information

Cited By

  • Moire pattern removing method and system based on generative AI model

    CN120997068A

  • Video target detection method and system based on multi-scale perception diffusion

    CN121190746A

  • Crane rotating part fault diagnosis method, device and equipment and storage medium

    CN121256576A