Infrared small target detection method and system fusing model driving and deep learning
Through the internal and external variational collaborative deep neural network and the global feature fusion module, the problem of infrared small target detection in complex backgrounds is solved, and high-precision target and background distinction is achieved.
Patent Information
- Application Number
- CN202510647294.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-23
AI Technical Summary
In complex backgrounds, traditional infrared small target detection methods have difficulty effectively distinguishing targets from backgrounds, resulting in poor detection results. Deep learning methods still face challenges in complex scenarios.
The internal and external variational collaborative deep neural network is adopted, combined with the global feature fusion module and the variational module, and the network is trained through cross entropy loss to achieve accurate distinction between the target and the background.
The accuracy and robustness of infrared small target detection are improved, and it can effectively distinguish the target from the background under complex backgrounds and output binary images of the target area and background.
Smart Images

Figure CN120689659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent object recognition, and in particular to a method and system for detecting small infrared targets in complex backgrounds. Background Art
[0002] Small target detection is a key research area in infrared imaging, widely used in fields such as traffic monitoring, maritime target search, and fire alarm. Small infrared targets are typically characterized by small size, low brightness, high similarity to the background, and dynamic changes, making detection in complex environments extremely challenging. In particular, at sea, at night, or in adverse weather conditions, traditional detection methods often struggle to effectively distinguish between the target and the background due to the extremely low contrast between background noise and the target, resulting in poor target detection. Traditional image processing methods typically rely on hand-crafted features and physical models. These methods, including image filtering, low-rank approximation, and edge detection, extract local or global features from the image for target detection. However, these methods rely on extensive experience and hyperparameter settings, such as filter size and weight selection. These hyperparameters are often highly sensitive and difficult to maintain consistency and efficiency in complex environments. Furthermore, traditional methods often fail to fully capture the high-dimensional semantic information in the image, resulting in unsatisfactory performance in dynamic scenes. In recent years, deep learning techniques, particularly convolutional neural networks and other neural network architectures, have made significant progress in the task of infrared small target detection. Deep learning methods can automatically learn more complex and abstract features from large amounts of data, avoiding the reliance of traditional methods on handcrafted features. Therefore, the application of deep learning in target detection can significantly improve detection accuracy and robustness. However, while deep learning methods have demonstrated superior performance on some standard datasets, they still face many challenges in complex scenarios. The present invention relates to the field of intelligent object recognition technology, and more particularly to a method and system for detecting small infrared targets in complex backgrounds. Summary of the Invention
[0003] The present invention aims to propose an infrared small target detection method and system that integrates model driving and deep learning. It uses the designed internal and external variational collaborative deep neural network to realize infrared target prediction in complex scenes and output a predicted image that can effectively distinguish the target from the background.
[0004] The present invention is achieved by utilizing the following technical solutions:
[0005] In a first aspect, a method for detecting infrared small targets by integrating model-driven and deep learning is provided, the method comprising the following steps:
[0006] Step S1, capturing a single frame of infrared image data from an infrared video to construct an infrared small target image dataset containing a complex background;
[0007] Step S2, preprocessing the single-frame infrared image data to obtain infrared small target images under complex backgrounds, and dividing the images into a training set, a validation set, and a test set;
[0008] Step S3: constructing an internal and external variation collaborative deep neural network, taking the training set as input, using the internal and external variation collaborative deep neural network to extract the infrared small target image in the complex background based on the infrared small target image data, and outputting the predicted small target image; and using the internal and external variation collaborative deep neural network to obtain a template image through the test set;
[0009] Step S4, calculating the cross entropy loss between the predicted small target image and the template image;
[0010] Step S5, using the cross entropy loss to guide the neural network training process, using the evidence set to verify the training results respectively, to achieve infrared small target prediction under complex background, and output the result as a binary image including the target area and background.
[0011] In some embodiments, step S3 further includes configuring the variational module to reconstruct the error The error of the variational module reconstruction is The sum is used to update the features generated by the variational module As shown below:
[0012]
[0013] in, represents the input features of the variational module, represents the output features of the variational module, i, j represent the two-dimensional spatial coordinates, k represents the time series representing the input features, and λ is the control parameter. Indicates the pixel value corresponding to the feature of sequence k at the spatial coordinate (i, j+1), represents the original input image, Indicates the image oscillation information of the feature of sequence k in the j+1 direction, Indicates the image oscillation information of the feature with sequence k in the j-1 direction, Indicates the image oscillation information of the feature in the i+1 direction, Indicates the image oscillation information of the feature in the i-1 direction, Represents Represents the deviation information from the original input image.
[0014] In some embodiments, step S3 further includes that the internal and external variational collaborative deep neural network includes a variational module configured with an outer encoding layer and an outer decoding layer, the outer encoding layer further includes multiple inner feature extraction layers, and the outer decoding layer further includes multiple pairs of global feature fusion modules and inner feature extraction layers stacked together.
[0015] In some embodiments, step S3 further includes configuring an inner feature extraction layer, wherein the inner feature extraction layer includes multiple inner encoding layers and multiple inner decoding layers. After three layers of inner encoding layers are connected in series at the input end, they are correspondingly connected in series with three layers of inner decoding layers. The structure of the inner encoding layer includes a 3×3 convolutional layer, a BatchNorm layer and a first ReLU activation function connected in sequence; the inner decoding layer includes 4 layers connected in sequence, namely a Concat convolutional layer, a 3×3 convolutional layer, a BatchNorm layer and a second ReLU activation function layer; the various features processed by the inner encoding layer and the inner decoding layer are fused through the Cancat convolutional layer to form a U-shaped feature extraction structure.
[0016] In some embodiments, step S3 further includes configuring a global feature fusion module, which includes a channel attention layer, a Concat convolution layer, a spatial attention layer, a 3×3 convolution layer, a BatchNorm layer, and a third ReLU activation function layer connected in sequence.
[0017] In some embodiments, the global feature fusion module is combined with the inner feature extraction layer to achieve the fusion of the outer encoder output and the decoder output. The fusion process is shown in the following formula:
[0018]
[0019] u′ e =Z c ⊙u e
[0020] Among them, the average pooling layer of the channel attention layer is defined as u e represents the features output by the inner feature extraction layer in the outer encoding layer, W c represents the first linear transformation layer in the channel attention layer, δ represents the ReLU activation function, ⊙ represents the dot product, u′ e represents the infrared small target image features after spatial attention layer enhancement, c, H, W represent the number of channels, height and width of the features respectively, Z c represents channel attention, u eRepresents the image features of infrared small targets, and F(i, j, c) represents the value of feature F corresponding to the spatial coordinate (i, j, c);
[0021] u concat =[u′ e ;u d ]
[0022]
[0023] u f =(1-Z s )⊙u e ′+Z s ⊙u d
[0024] Among them, u d Represents the infrared small target image features output by the inner feature extraction layer in the outer decoding layer, and the average pooling layer of spatial attention is defined as The maximum pooling layer is defined as W s Indicates that the second linear transformation layer in the spatial attention mechanism is used to enhance the ability of spatial feature extraction, u f Represents the fused features of the encoder and decoder.
[0025] In some implementations, the cross entropy loss between the predicted small object image and the template image in step 4 is as follows:
[0026]
[0027] Among them, N represents the number of samples, which is used to average the loss value, y n and Represent the label template image of the nth sample and the predicted sample of the nth network, The logarithmic value of the probability that the predicted result is the target is used to measure the confidence of the model in the target. The logarithm of the probability that the prediction result is background is used to measure the model's confidence in the background.
[0028] Secondly, a fusion model-driven and deep learning infrared small target detection system is provided for performing an infrared small target detection method. The system includes an interception module, a preprocessing module, an extraction module, a loss design module, and an infrared small target detection module.
[0029] The interception module is used to intercept single-frame infrared image data from the infrared video and construct an infrared small target image dataset containing a complex background;
[0030] A preprocessing module is used to preprocess the single-frame infrared image data to obtain an infrared small target image under a complex background, and divide the image into a training set, a verification set, and a test set;
[0031] An extraction module is configured to construct an internal and external variation collaborative deep neural network, take a training set as input, use the internal and external variation collaborative deep neural network to extract the infrared small target image in a complex background based on the infrared small target image data in the complex background, and output a predicted small target image; and use the internal and external variation collaborative deep neural network to obtain a template image through a test set;
[0032] A loss design module, used to design a cross entropy loss between the predicted small target image and the template image;
[0033] The infrared small target detection module uses the evidence set to verify the training results respectively, realizes the prediction of infrared small targets under complex backgrounds, and outputs a binary image including the target area and the background.
[0034] Compared with the prior art, the present invention can achieve the following beneficial technical effects:
[0035] 1) Build an internal and external variational collaborative deep neural network to extract infrared small target images in complex backgrounds based on the infrared small target image data. The input is the infrared small target image in complex backgrounds. The network output is a predicted binary image in which the target area is white and the background is black. The inner feature extraction layer is designed to include multiple inner encoding layers and multiple inner decoding layers. The features processed by the inner encoding layer and the inner decoding layer are fused through the Cancat convolution layer to form a U-shaped structure to achieve effective feature extraction.
[0036] 2) Through the configuration of the global feature fusion module, the features with rich spatial information output by the encoder and the features with rich semantic information output by the decoder are fused to improve the model's ability to express fine-grained structures and achieve complementary advantages;
[0037] 3) The variational module can effectively utilize the local information of infrared small targets provided by the variational module, thereby achieving accurate prediction results; the reconstruction error of the variational module is used to update the output features of the variational module. This feature can provide the inner encoder and decoder with spatial information of small targets at different scales, which is conducive to the detection of infrared small targets;
[0038] 4) The output result includes a binary image of the target area and the background, which can effectively distinguish the target from the background;
[0039] 5) Combined with deep learning methods, it can automatically learn more complex and abstract features from large amounts of data, avoiding the traditional method's reliance on manual features. Its application in target detection can significantly improve the accuracy and robustness of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is the overall flow chart of the infrared small target detection method based on the fusion model drive and deep learning of the present invention;
[0041] Figure 2 Schematic diagram of the implementation process of the infrared small target detection method based on the fusion model drive and deep learning of the present invention;
[0042] Figure 3 Schematic diagram of the deep neural network structure of an embodiment of the present invention; (3A) Schematic diagram of the internal and external variational collaborative deep neural network structure; (3B) Schematic diagram of the inner feature extraction layer structure;
[0043] Figure 4 Schematic diagram of the structure of the global feature fusion module according to an embodiment of the present invention;
[0044] Figure 5 This is a module diagram of the infrared small target detection system based on the fusion model drive and deep learning of the present invention.
[0045] Figure 6 Schematic diagram of the effect of infrared small target detection under complex background according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0047] Example 1
[0048] Figure 1 The flowchart of the infrared small target detection method provided by the present invention is shown. The flowchart specifically includes the following steps:
[0049] Step 1: Construct an infrared small target image dataset with a complex background. Specifically, by adjusting the exposure time of the infrared camera, under different background noise and environmental conditions, multiple infrared videos containing the target at different angles are captured by recording. Single-frame infrared images are captured from each infrared video and saved as raw single-frame infrared image data. All single-frame images captured from the video constitute the raw infrared image data and the raw infrared image dataset.
[0050] Step 2: Process the original single-frame infrared image data:
[0051] In step 2.1, the original single-frame infrared image is cropped, rotated, and spliced to enrich the diversity of infrared targets in the dataset. Template images of small targets are manually calibrated and labeled to ensure that the input and output image sizes are consistent for unified training and testing.
[0052] Step 2.2: Finally, randomly split the dataset into training set, validation set, and test set according to a certain ratio (e.g., 8:1:1);
[0053] Step 3: Build an internal and external variation collaborative deep neural network to extract infrared small targets under complex backgrounds; Figure 3 A is a schematic diagram of the structure of the internal and external variational collaborative deep neural network according to an embodiment of the present invention. In the internal and external variational collaborative deep neural network, multiple outer coding layers and outer decoding layers are included. Each outer coding layer and outer decoding layer is configured with a variational module. The outer coding layer further includes multiple inner feature extraction layers. The outer decoding layer further includes multiple pairs of global feature fusion modules and inner feature extraction layers for stacking. The global feature fusion module is used to fuse the features extracted by the outer coding layer and provide them to the inner feature extraction layer as input; the input is an infrared small target image under a complex background, and the network output is a predicted binary image, in which the target area is white and the background is black;
[0054] Step 3.1, configure the variational module 100, which is used to extract the local features of the infrared small target image, including local structure and edge features. Inspired by the total variation model, the infrared small target image is iterated in the form of a partial differential equation to perform variational module reconstruction error Configuration,, the error of reconstruction of variational module and The sum is used to update the features generated by the variational module As shown in formula (1):
[0055]
[0056] in, represents the input features of the variational module, represents the output features of the variational module, i, j represent the two-dimensional spatial coordinates, k represents the time series representing the input features, and λ is the control parameter. Indicates the pixel value corresponding to the feature of sequence k at the spatial coordinate (i, j+1), represents the original input image, Indicates the image oscillation information of the feature of sequence k in the j+1 direction, Indicates the image oscillation information of the feature with sequence k in the j-1 direction, Indicates the image oscillation information of the feature in the i+1 direction, Indicates the image oscillation information of the feature in the i-1 direction, Represents Represents the deviation information from the original input image.
[0057] Respectively represent a semi-discrete coefficient, as shown in formula (2):
[0058]
[0059] Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i, j±1), Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i+1,j±1), Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i-1, j±1), Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i±1,j), Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i±1,j+1), Indicates the pixel value corresponding to the feature of time series k at the spatial coordinate (i±1,j-1);
[0060] Step 3.2, configure the inner feature extraction layer 200. As shown in Figure 3B, it is a schematic diagram of the structure of the inner feature extraction layer of an embodiment of the present invention. The inner feature extraction layer 200 is used to combine the local features of the infrared small target image extracted by the variational module 100 to achieve good prediction to obtain feature extraction results. The inner feature extraction layer includes multiple inner encoding layers and multiple inner decoding layers. After the multiple inner encoding layers are connected in series, they are connected to the input features; the structure of each inner encoding layer includes three layers connected in sequence, namely a 3×3 convolution layer, which is used for feature extraction such as edges, textures, and shapes; a BatchNorm layer, which speeds up network training and prevents gradient disappearance or explosion by normalizing the output of each layer; and a first ReLU activation function layer, which enables the network to fit complex mapping relationships. The structure of each inner decoding layer includes four layers connected in sequence, namely a Concat convolution layer for connecting features of different scales; a 3×3 convolution layer for extracting features such as edges, textures, and shapes; a BatchNorm layer that speeds up network training and prevents gradient vanishing or exploding by normalizing the output of each layer; and a second ReLU activation function layer that enables the network to fit complex mapping relationships. At the same time, the features processed by the inner encoding layer and the inner decoding layer are fused through the Cancat convolution layer to form a U-shaped structure for effective feature extraction.
[0061] After setting up three layers of inner encoding layers in series on the input side, it is then connected to the corresponding three layers of inner decoding layers in series. The input features are encoded and decoded according to the inner encoding layers and inner decoding layers set up by the network layer defined above.
[0062] Step 3.3, configure the global feature fusion module 300, such as Figure 4As shown in FIG, the module includes 6 layers connected in sequence, namely a channel attention layer, a Concat convolution layer, a spatial attention layer, a 3×3 convolution layer, a BatchNorm layer and a third ReLU activation function layer. Among them, the channel attention layer is used to focus on important channel information, and the spatial attention layer is used to focus on important spatial positions to strengthen meaningful areas. The small target features processed by the inner encoder and decoder are fused through the Concat convolution layer, the 3×3 convolution layer, the BatchNorm layer and the ReLU activation function layer to obtain features that integrate semantic information and spatial information at different scales, enhance the detection and reconstruction capabilities of small targets, and improve the model's ability to express fine-grained structures. That is, the global feature fusion module is combined with the inner feature extraction layer to achieve targeted fusion of small target image features at different levels. Targeted fusion at different levels refers to the fusion of features with rich spatial information output by the encoder and features with rich semantic information output by the decoder to complement each other. The fusion process is shown in formula (3):
[0063]
[0064] The average pooling layer of the channel attention layer is defined as u e represents the features output by the inner feature extraction layer in the outer encoding layer, W c represents the first linear transformation layer in the channel attention layer, δ represents the ReLU activation function, which is used to enhance the nonlinear expression ability, ⊙ represents the dot product, u′ e represents the infrared small target image features after spatial attention layer enhancement, c, H, W represent the number of channels, height and width of the features respectively, Z c Represents channel attention, used to strengthen feature u e Channel features, F(i,j,c) represents the value of feature F corresponding to the spatial coordinate (i,j,c).
[0065]
[0066] Among them, u d Represents the infrared small target image features output by the inner feature extraction layer in the outer decoding layer, and the average pooling layer of spatial attention is defined as The maximum pooling layer is defined as W s The second linear transformation layer in the spatial attention mechanism is used to enhance the ability of spatial feature extraction, and the fused feature u f It is obtained by a linear combination method, which combines the information of the encoder and decoder. srepresents the weight parameter generated by the second formula in formula (4), which is used for the linear combination of the third formula in formula (4), u concat represents u′ generated by the first formula in formula (4) e and u d After the features are connected in the channel dimension, F(c,i,j) represents the value corresponding to the feature F at the spatial coordinate (c,i,j);
[0067] Step 3.4: Configure an inner-outer variational collaborative deep neural network, which consists of multiple outer encoding layers and multiple outer decoding layers. Each outer encoding layer and outer decoding module is configured with a variational module. This allows the network to effectively utilize the local information of small infrared targets provided by the variational module during the encoding and decoding process, thereby achieving accurate prediction results.
[0068] In step 4, the last outer decoding layer maps the output value of the inner and outer variational collaborative deep neural network to the range of (0, 1) through the Sigmoid function layer, and calculates the cross entropy loss between the predicted small target image and the label template image through the binary cross entropy loss function. The cross entropy loss is used to guide the training process of the neural network to achieve the prediction of infrared small targets in complex backgrounds.
[0069] The step 4 also includes:
[0070] Using the binary cross entropy loss function (BCELoss), the cross entropy loss between the predicted small target image and the label template image is calculated, and this loss is used to guide the training process of the neural network. The loss function is shown in (4):
[0071]
[0072] Among them, N represents the number of samples, which is used to average the loss value, y n and Represent the label template image of the nth sample and the predicted sample of the nth network, and Used to measure the difference between the predicted value and the true label template, The logarithmic value of the probability that the predicted result is the target is used to measure the confidence of the model in the target. The logarithm of the probability that the prediction result is background is used to measure the model's confidence in the background. The two together constitute the cross-entropy loss, which is used to supervise the model's training performance in the binary classification task.
[0073] Step 5, according to the requirements of recognition, adjust the parameters of the internal and external variation collaborative deep neural network, including the initial learning rate, batch size and inner U-shaped feature extraction structure. The U-shaped feature extraction structure is composed of the outer coding layer and the outer decoding layer. The above process mentioned that the global feature fusion module connects the inner feature extraction layer of the outer coding layer and the inner feature extraction layer of the outer decoding layer to form a U-shaped structure. The number of inner feature extraction layers of the outer coding layer and the number of inner feature extraction layers of the outer decoding layer (the number of outer coding layers and decoding layers is the same and both are four, and the outer coding layer and decoding layer each contain an inner feature extraction layer). The deep neural network based on the internal and external full variation collaborative is trained, and the input image is flipped and rotated during the training process to achieve data expansion; specifically, in this embodiment, the minimum batch size is 16, the learning rate is initialized to 0.001, the training cycle is set to 1000, the significance index is 0.5, and the Adam algorithm is used to optimize the loss function.
[0074] Example 2
[0075] like Figure 5 As shown, the module diagram of the infrared small target detection system of the fusion model driven and deep learning provided by the present invention for executing the above-mentioned infrared small target detection method of the fusion model driven and deep learning is shown; the system includes an interception module 510, a preprocessing module 520, an extraction module 530, a loss design module 540 and an infrared small target detection module 550;
[0076] The interception module 510 is used to intercept single-frame infrared image data from the infrared video and construct an infrared small target image dataset containing a complex background;
[0077] The preprocessing module 520 is used to preprocess the single-frame infrared image data to obtain infrared small target images under complex backgrounds, and divide them into a training set, a validation set, and a test set;
[0078] The extraction module 530 builds an internal and external variation collaborative deep neural network, takes the training set as input, uses the internal and external variation collaborative deep neural network to extract the infrared small target image in the complex background based on the infrared small target image data, and outputs a predicted small target image; and uses the internal and external variation collaborative deep neural network to obtain a template image through the test set;
[0079] The loss design module 540 is used to design the cross entropy loss between the predicted small target image and the template image;
[0080] The infrared small target detection module 550 uses the cross entropy loss to guide the neural network training process, and uses the test set and the validation set to test and verify the training results respectively, so as to achieve infrared small target prediction in complex backgrounds.
[0081] In the embodiment of the present invention, the infrared small target detection network based on the collaboration of internal and external total variation is used to obtain the detected infrared small target, such as Figure 6 The figure shows the effect of infrared small target detection under complex background according to an embodiment of the present invention. (6A) is the observed image of the infrared small target, (6B) is the infrared small target detected using the embodiment of the present invention, and (6C) is the template image of the calibration target corresponding to the observed image. It can be seen that the underwater polarization image infrared small target detection method proposed by the present invention using internal and external total variation and deep learning can effectively identify infrared small targets under complex backgrounds.
[0082] In summary, this paper proposes a method for infrared small target detection that integrates physical models and deep learning, namely, the Internal-External Variation Collaborative Network. This method fully utilizes the target's spatial information by enhancing the target's structural features at different network layers. Specifically, this method introduces a module inspired by the total variation model, which aims to capture the key structural information of the target and is implemented in the form of partial differential equations. In addition, this paper proposes an Internal-External Variation Collaborative Architecture that enhances the target's spatial information at different depth levels by combining the total variation-inspired module with the encoder and decoder. At the same time, this method also introduces a dual spatial attention mechanism that uses channel information to enhance the semantic distinction between the target and the background, and simultaneously fuses spatial features from different layers to achieve more accurate target detection.
[0083] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, various modifications and variations may be made to the present application. Those skilled in the art may make any modifications, equivalent substitutions, or variations without departing from the spirit and scope of the present invention. All modifications, equivalent substitutions, or variations made within the spirit and principles of the present application fall within the scope of protection of the present invention as defined by the appended claims.
[0084] It should be noted that although the present invention has been shown and described with reference to specific exemplary embodiments of the present invention, those skilled in the art should understand that the present invention is not limited to the above-mentioned embodiments and all kinds of changes to the present invention fall within the scope of protection of the present invention.
Claims
1. A method for infrared small target detection that integrates model driving and deep learning, characterized in that: The method comprises the following steps: Step S1, capturing a single frame of infrared image data from an infrared video to construct an infrared small target image dataset containing a complex background; Step S2, preprocessing the single-frame infrared image data to obtain infrared small target images under complex backgrounds, and dividing the images into a training set, a validation set, and a test set; Step S3: constructing an internal and external variation collaborative deep neural network, taking the training set as input, using the internal and external variation collaborative deep neural network to extract the infrared small target image in the complex background based on the infrared small target image data, and outputting the predicted small target image; and using the internal and external variation collaborative deep neural network to obtain a template image through the test set; Step S4, designing a cross entropy loss between the predicted small target image and the template image; Step S5, using the cross entropy loss to guide the neural network training process, using the evidence set to verify the training results respectively, to achieve infrared small target prediction under complex background, and output the result as a binary image including the target area and background.
2. The infrared small target detection method based on fusion model driving and deep learning according to claim 1 is characterized in that: Step S3 further includes configuring the variational module to reconstruct the error The error of the variational module reconstruction is The sum is used to update the features generated by the variational module As shown below: in, represents the input features of the variational module, represents the output features of the variational module, i, j represent the two-dimensional spatial coordinates, k represents the time series representing the input features, and λ is the control parameter. Indicates the pixel value corresponding to the feature of sequence k at the spatial coordinate (i, j+1), represents the original input image, Indicates the image oscillation information of the feature of sequence k in the j+1 direction, Indicates the image oscillation information of the feature with sequence k in the j-1 direction, Indicates the image oscillation information of the feature in the i+1 direction, Indicates the image oscillation information of the feature in the i-1 direction, Represents Represents the deviation information from the original input image.
3. The infrared small target detection method based on fusion model driving and deep learning according to claim 1 is characterized in that: Step S3 further includes that the internal and external variational collaborative deep neural network includes a variational module configured with an outer encoding layer and an outer decoding layer, the outer encoding layer further includes multiple inner feature extraction layers, and the outer decoding layer further includes multiple pairs of global feature fusion modules and inner feature extraction layers stacked together.
4. The infrared small target detection method based on fusion model driving and deep learning according to claim 1 is characterized in that: Step S3 further includes configuring an inner feature extraction layer, wherein the inner feature extraction layer includes multiple inner encoding layers and multiple inner decoding layers. After three layers of inner encoding layers are connected in series at the input end, they are correspondingly connected in series with three layers of inner decoding layers. The structure of the inner encoding layer includes a 3×3 convolution layer, a BatchNorm layer and a first ReLU activation function connected in sequence; the inner decoding layer includes 4 layers connected in sequence, namely a Concat convolution layer, a 3×3 convolution layer, a BatchNorm layer and a second Relu activation function layer; the various features processed by the inner encoding layer and the inner decoding layer are fused through the Cancat convolution layer to form a U-shaped feature extraction structure.
5. The infrared small target detection method integrating model driving and deep learning according to claim 1 is characterized in that: Step S3 further includes configuring a global feature fusion module, which includes a channel attention layer, a Concat convolution layer, a spatial attention layer, a 3×3 convolution layer, a BatchNorm layer, and a third ReLU activation function layer connected in sequence.
6. The infrared small target detection method integrating model driving and deep learning according to claim 1 is characterized in that: The global feature fusion module is combined with the inner feature extraction layer to achieve the fusion of the outer encoder output and the decoder output. The fusion process is shown in the following formula: at' e =Z c ⊙u e Among them, the average pooling layer of the channel attention layer is defined as u e represents the features output by the inner feature extraction layer in the outer encoding layer, W c represents the first linear transformation layer in the channel attention layer, δ represents the ReLU activation function, ⊙ represents the dot product, u′ e represents the infrared small target image features after spatial attention layer enhancement, c, H, W represent the number of channels, height and width of the features respectively, Z c represents channel attention, u e Represents the image features of infrared small targets, and F(i, j, c) represents the value of feature F corresponding to the spatial coordinate (i, j, c); in concat =[u′ e ;in d ] at f =(1-Z s )⊙u e ′+Z s ⊙u d Among them, u d Represents the infrared small target image features output by the inner feature extraction layer in the outer decoding layer, and the average pooling layer of spatial attention is defined as represents the maximum pooling layer, W s Indicates that the second linear transformation layer in the spatial attention mechanism is used to enhance the ability of spatial feature extraction, u f Represents the fused features of the encoder and decoder.
7. The infrared small target detection method integrating model driving and deep learning according to claim 1 is characterized in that: The cross entropy loss between the predicted small target image and the template image in step 4 is as follows: Among them, N represents the number of samples, which is used to average the loss value, y n and Represent the label template image of the nth sample and the predicted sample of the nth network, The logarithmic value of the probability that the predicted result is the target is used to measure the confidence of the model in the target. The logarithm of the probability that the prediction result is background is used to measure the model's confidence in the background.
8. A fusion model-driven and deep learning infrared small target detection system that implements the fusion model-driven and deep learning infrared small target detection method according to any one of claims 1 to 7, characterized in that: The system includes an interception module, a preprocessing module, an extraction module loss design module and an infrared small target detection module; The interception module is used to intercept single-frame infrared image data from the infrared video and construct an infrared small target image dataset containing a complex background; The preprocessing module is used to preprocess the single-frame infrared image data to obtain infrared small target images under complex backgrounds, and divide the images into a training set, a verification set, and a test set; The extraction module is used to build an internal and external variation collaborative deep neural network, take the training set as input, use the internal and external variation collaborative deep neural network to extract the infrared small target image with a complex background based on the infrared small target image data with a complex background, and output the predicted small target image; and use the internal and external variation collaborative deep neural network to obtain a template image through the test set; The loss design module is used to design the cross entropy loss between the predicted small target image and the template image; The infrared small target detection module uses the cross entropy loss to guide the neural network training process, uses the test set and the validation set to test and verify the training results respectively, and realizes the prediction of infrared small targets in complex backgrounds.