Infrared Small Target Detection Method Based on Runge-Kutta Residual Block

By constructing an infrared small object detection model based on Longge-Kuta residual block, the combination of codec network and edge extraction network is used to solve the problem of low accuracy of infrared small object detection in complex backgrounds, and higher detection accuracy is achieved.

CN116580276BActive Publication Date: 2025-07-29XIDIAN UNIV HANGZHOU RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310235707.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2025-07-29
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

The existing infrared small object detection method has low detection accuracy in complex backgrounds, making it difficult to effectively suppress high-frequency noise and retain detailed characteristics.

Method used

A codec and edge extraction network based on Longge-Kuta residual block is adopted, combining attention mechanisms and convolutional neural networks, and the clarity of target feature extraction and edge information is improved through iterative training.

Benefits of technology

It significantly improves the accuracy of infrared small target detection, can accurately locate targets in complex backgrounds, and enhances the effectiveness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580276B_ABST
    Figure CN116580276B_ABST
Patent Text Reader

Abstract

The present invention discloses an infrared small target detection method based on a Runge-Kutta residual block, which mainly solves the problem of low detection accuracy of existing methods. The implementation steps are as follows: obtaining a training sample set and a test sample set; constructing an infrared small target detection model based on a Runge-Kutta residual block, including an encoding-decoding network and a Head module connected in sequence, and an edge extraction network; performing iterative training on the infrared small target detection model; and obtaining the infrared image small target detection result. The Runge-Kutta residual block in the encoding-decoding network of the present invention utilizes the advantages of the attention mechanism in capturing long-range dependence relationships and the convolutional neural network in extracting local features to extract semantics and retain details, which can effectively enhance target features and suppress high-frequency noise. The edge extraction network can give clear edge information to the target from multiple levels, improving the accuracy of infrared small target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an infrared small target detection method based on a Runge-Kutta residual block, which can be used for maritime monitoring and traffic management. Background Art

[0002] Due to its unique imaging mechanism, an infrared image can capture more information about a target even under extremely challenging visible light imaging conditions. Detecting targets, especially small targets, in infrared images has attracted a great deal of attention and is of great significance in practical applications such as maritime monitoring and traffic management. However, due to the low contrast between small targets and complex backgrounds in infrared images, it is challenging to detect infrared small targets.

[0003] Traditional infrared small target detection methods usually model small target detection as a filtering and target enhancement problem, including methods based on filtering, spectral residuals, local contrast, and low-rank representation, etc. They can provide good performance in some simple backgrounds, but cannot suppress complex background noise and heavily rely on hyperparameter tuning and handcrafted features.

[0004] Many researchers have introduced deep learning into the field of infrared small target detection. For example, the University of Electronic Science and Technology proposed an infrared small target detection method based on global mean contrast spatial attention in its patent document "Infrared Small Target Detection Method Based on Global Mean Contrast Spatial Attention" (Patent Application No.: CN202211398750.X, Publication No.: CN115527098A). The detection network of global contrast learning mainly consists of a feature extraction module, a feature fusion module, and a detection module. The feature extraction module includes 1 Focus-Conv sub-module and a dual-channel feature extraction sub-module. The dual-channel feature extraction sub-module includes a Conv-CSP component and a SAG-Conv component respectively, for extracting features of infrared small targets. The feature fusion module includes 3 feature fusion channels. The first channel is the output feature of the SAGG component at the end of the feature extraction module, the second channel is the output feature of the CSPG component at the end of the feature extraction module, and the third channel is the feature result after upsampling the output of the CSPG component at the end of this fusion module, for multi-level feature fusion. The detection module includes two groups of Conv-SKConv sub-modules and a small target detection head designed based on the YOLOV technology framework. In this invention, low-level detailed feature information extracted by the network is easily lost, resulting in a low detection accuracy. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention provides an infrared small target detection method based on a Runge-Kutta residual block, which is used to solve the problem of low detection accuracy of existing infrared small target detection methods.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] (1) Obtain a training sample set and a test sample set:

[0008] Obtain K infrared images, perform edge detection on each infrared image to obtain K infrared edge images, then label the targets in each infrared image and its corresponding infrared edge image to obtain the labels of K infrared images and the labels of K infrared edge images. Then, form a training sample set R1 with M infrared images and their corresponding labels, and the labels of M infrared edge images, and form a test sample set E1 with the remaining K - M infrared images and their corresponding labels, and the labels of K - M infrared edge images, where K ≥ 500.

[0009] (2) Construct an infrared small target detection model based on a Runge-Kutta residual block:

[0010] Construct an infrared small target detection model O including an encoding-decoding network and a Head module connected in sequence, and an edge extraction network, where:

[0011] The encoding-decoding network includes a Stem module, an encoding module containing H encoding sub-modules, and a decoding module containing H - 1 transposed convolutional layers connected in sequence; the Stem module includes a plurality of convolutional layers and a pooling layer connected in sequence; each encoding sub-module includes an embedding layer and a Runge-Kutta residual block connected in sequence; the output end of the h-th encoding sub-module is also connected to the output end of the (H - h)-th transposed convolutional layer.

[0012] The edge extraction network is loaded between the output end of the encoding module and the input end of the Head module, and includes an edge detection module including H edge detection sub-modules arranged in parallel and an edge feature fusion module cascaded therewith; the input end of the h-th edge detection sub-module is connected to the output end of the h-th encoding sub-module; the h-th edge detection sub-module includes a Sobel operator, a convolutional layer, and h - 1 transposed convolutional layers connected in sequence; the edge feature fusion module includes two convolutional layers arranged in parallel, and the output end of one of the convolutional layers is also connected to a normalization layer and a non-linear activation layer; the Head module includes three branches arranged in parallel, and each branch includes two convolutional layers connected in sequence.

[0013] (3) Iteratively train the infrared small target detection model O:

[0014] (3a) Initialize the iteration count as t, the maximum iteration count as T, where T ≥ 1000. The weight and bias parameters of the infrared small target detection model O in the t-th iteration are w t and b t respectively. Let t = 0 and O t = O; t

[0015] (3b) Randomly and with replacement, select L infrared images from the training sample set R1 as the input for the forward propagation of the infrared small target detection model O, where 1 ≤ L ≤ M:

[0016] (3b1) The Stem module in the encoding-decoding network downsamples each infrared image, and the H encoding sub-modules in the encoding module perform feature encoding on the result of each downsampling. The H - 1 transposed convolutional layers in the decoding module upsample the result of each feature encoding to obtain L target feature information;

[0017] (3b2) The h-th edge detection sub-module in the edge extraction network performs edge detection on the output features of the h-th encoding sub-module, and the edge feature fusion module fuses the outputs of the H edge detection sub-modules to obtain L target edge feature information;

[0018] (3b3) The Head module predicts the L multi-layer feature information obtained by merging the L target feature information and their corresponding L target edge feature information to obtain L infrared small target detection results;

[0019] (3c) Using the cross-entropy loss function, calculate the loss value of the edge extraction network in O t through the generated L target edge feature information and the labels of their corresponding L infrared edge images and use the Dice loss function to calculate the loss value L t of O Dice through the generated L infrared small target detection results and the labels of their corresponding L infrared images. Then, use the chain rule to calculate and L Dice to obtain the partial derivatives of the total loss Lt composed of t with respect to the weight parameter ω t and the bias parameter b and Finally, update ω and b t according to t to obtain the network model O t for this iteration;

[0020] (3d) Determine whether t ≥ T holds. If so, obtain the trained infrared small target detection model O*. Otherwise, set t = t + 1 and execute step (3b);​

[0021] (4) Obtain the detection results of small infrared targets:

[0022] Use the test sample set E1 as the input of the trained small infrared target detection model O* for forward propagation to obtain the detection results of small infrared targets corresponding to all infrared images.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] The small infrared target detection model based on the Runge-Kutta residual block constructed by the present invention includes an encoding-decoding network and a Head module connected in sequence, as well as an edge extraction network. During the training of the model and the process of obtaining the detection results of small infrared targets, the Runge-Kutta residual block in the encoding-decoding network utilizes the advantages of the attention mechanism in capturing long-range dependence relationships and the convolutional neural network in extracting local features to extract semantics and retain details, which can effectively enhance the target features and suppress high-frequency noise. The edge extraction network can give clear edge information of the target from multiple levels, which helps to accurately locate the target. Experimental results show that the present invention can effectively improve the accuracy of small infrared target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the implementation flowchart of the present invention;

[0026] Figure 2 is the structural schematic diagram of the small infrared target detection model based on the Runge-Kutta residual block adopted in the embodiment of the present invention;

[0027] Figure 3 (a) is the structural schematic diagram of the Runge-Kutta residual block adopted in the embodiment of the present invention;

[0028] Figure 3 (b) is the structural schematic diagram of the feature extraction block adopted in the embodiment of the present invention;

[0029] Figure 3 (c) is the structural schematic diagram of the attention transformation block adopted in the embodiment of the present invention;

[0030] Figure 3 (d) is the structural schematic diagram of the convolutional residual block adopted in the embodiment of the present invention;

[0031] Figure 3 (e) is the structural schematic diagram of the channel attention module adopted in the embodiment of the present invention;

[0032] Figure 4 is the structural schematic diagram of the edge detection sub-module adopted in the embodiment of the present invention;

[0033] Figure 5Schematic diagram of the edge feature fusion block adopted in the embodiment of the present invention. Detailed implementation manners

[0034] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Referring to Figure 1 , the present invention includes the following steps:

[0036] (1) Obtain a training sample set and a test sample set:

[0037] Obtain 1000 infrared images included in the IRSTD-1k dataset, perform edge detection on each infrared image to obtain 1000 infrared edge images, then label the targets in each infrared image and its corresponding infrared edge image to obtain 1000 labels for infrared images and 1000 labels for infrared edge images. Then, form a training sample set R1 with 600 infrared images and their corresponding 600 labels, as well as 600 labels for infrared edge images, and form a test sample set E1 with the remaining 400 infrared images and their corresponding 600 labels, as well as 600 labels for infrared edge images;

[0038] (2) Construct an infrared small target detection model based on the Runge-Kutta residual block:

[0039] Construct an infrared small target detection model O including an encoding-decoding network, a Head module, and an edge extraction network connected in sequence, where:

[0040] The encoding and decoding network includes a Stem module, an encoding module containing three encoding sub-modules, and a decoding module containing two transposed convolutional layers that are connected in sequence. The output end of the first encoding sub-module is also connected to the output end of the first transposed convolutional layer, and the output end of the second encoding sub-module is also connected to the output end of the second transposed convolutional layer; the Stem module includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a first pooling layer that are connected in sequence; the encoding sub-module includes an embedding layer and a Runge-Kutta residual block that are connected in sequence, where the embedding layer includes a Patch Embedding and a Position Embedding, and the Runge-Kutta residual block includes a first feature extraction block and a second feature extraction block that are connected in sequence. The feature extraction block includes an attention transformation block and a convolutional residual block arranged in parallel. After channel merging, a channel attention module and a fourth convolutional layer are connected in sequence. Among them, the attention transformation block includes a randomly connected attention block, a first normalization layer, a first fully connected layer, a first non-linear activation layer, a first dropout layer, a second fully connected layer, a second dropout layer, and a second normalization layer. The input end of the randomly connected attention block is also added to the output end of the randomly connected attention block, and the output end of the first normalization layer is also added to the output end of the second dropout layer. The convolutional residual block includes a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a third normalization layer, and a second non-linear activation layer. The input end of the fifth convolutional layer is also added to the output end of the third normalization layer. The channel attention module includes a second pooling layer arranged in parallel and a third fully connected layer and a fourth fully connected layer connected thereto, and a third pooling layer and a fifth fully connected layer and a sixth fully connected layer connected thereto. After channel merging, a third non-linear activation layer is connected in sequence. The input end of the channel attention module is also multiplied by the input end of the third non-linear activation layer. The specific parameters are as follows: the convolutional kernel size of the first convolutional layer is 3*3, the sliding step is 2, and the padding is 1. The convolutional kernel sizes of the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are all 3*3, the sliding steps are all 1, and the padding is 1. The convolutional kernel sizes of the fifth convolutional layer and the seventh convolutional layer are 1*1, and the sliding step is 1. The convolutional kernel size of the sixth convolutional layer is 3*3; both the first pooling layer and the second pooling layer use max pooling, and the third pooling layer uses average pooling; both the first normalization layer and the second normalization layer use layer normalization, and the third normalization layer uses batch normalization; the first non-linear activation layer uses the GELU activation function, the second non-linear activation layer uses the ReLU activation function, and the third non-linear activation layer uses the Sigmoid activation function;

[0041] Among them, the structure of the infrared small target detection model based on the Runge-Kutta residual block adopted in the embodiment of the present invention is as shown in Figure 2 shown; the structure of the Runge-Kutta residual block adopted in the embodiment of the present invention is as shown in Figure 3As shown in (a); the structure of the feature extraction block adopted in the embodiment of the present invention is as Figure 3 shown in (b); the structure of the attention transformation block adopted in the embodiment of the present invention is as Figure 3 shown in (c); the structure of the convolutional residual block adopted in the embodiment of the present invention is as Figure 3 shown in (d); the structure of the channel attention module adopted in the embodiment of the present invention is as Figure 3 shown in (e).

[0042] The connection method between the two feature extraction blocks in the Runge-Kutta residual block is specifically as follows:

[0043] The Runge-Kutta method is a single-step method for solving ordinary differential equations with high precision. Using the second-order Runge-Kutta method to solve f(x j-1 , y j-1 ) = y j-1 -x j-1 , we get:

[0044]

[0045] Regarding x j-1 and y j-1 as the input and output of a feature extraction block, and x j-1 can also represent the output y j-2 of the previous feature extraction block, that is, x j-1 = y j-2 , so equation (1) can be rewritten as:

[0046]

[0047] where Δy j-1 = y j-1 -y j-2 .

[0048] Stack two feature extraction blocks in each Runge-Kutta residual block and add specific connections between them according to equation (2). The Runge-Kutta residual block can be expressed as:

[0049]

[0050] where y j-4 is the input of the Runge-Kutta residual block, y j-3 is the output of the first feature extraction block, y j-1 is the output of the second feature extraction block, and y j is the output of the Runge-Kutta residual block. By calculating the residual between the input and output features of the feature extraction block and then compensating in the output features, the Runge-Kutta residual block can act as an information bottleneck to suppress high-frequency noise and at the same time strengthen the target features through backpropagation gradients.

[0051] Specifically, the Random Connection Attention Block (RCA) consists of N intermediate neurons. Let the number of input and output neurons be D. Denote the input value, intermediate value, and output value by u(t), s(t), and f(t) respectively. The intermediate value update rule can be expressed as:

[0052] s(t) = tanh(W s s(t - 1)+W in u(t)) (4)

[0053] where t refers to the currently input image patch, and t - 1 refers to the previous input image patch.

[0054] The output f(t) can be calculated as:

[0055] f(t) = W out s(t) (5)

[0056] where W in , W s and W out are the weights corresponding to the input value, intermediate value, and output value respectively, and tanh(·) is the activation function.

[0057] The weight W in of the input value is initialized from a uniform distribution in the range [-W inmax , W inmax , where W inmax represents the maximum weight. The weight W s of the intermediate value is initialized by randomly generated W random and is transformed according to the following method:

[0058] W s = W random (ρ desired / {ρ(W random )}) (6)

[0059] where ρ(·) represents the radius of the value range of W random , and ρ desired is a hyperparameter set to 0.9.

[0060] The edge extraction network is loaded between the output end of the encoding module and the input end of the Head module, and includes a first edge detection sub-module, a second edge detection sub-module, and a third edge detection sub-module arranged in parallel, as well as an edge feature fusion module cascaded therewith. The input end of the first edge detection sub-module is also connected to the output end of the first encoding sub-module, the input end of the second edge detection sub-module is also connected to the output end of the second encoding sub-module, and the input end of the third edge detection sub-module is also connected to the output end of the third encoding sub-module; the first edge detection sub-module includes a Sobel operator and an eighth convolutional layer connected in sequence, the second edge detection sub-module includes a Sobel operator, a ninth convolutional layer, and a third transposed convolutional layer connected in sequence, and the third edge detection sub-module includes a Sobel operator, a tenth convolutional layer, a fourth transposed convolutional layer, and a fifth transposed convolutional layer connected in sequence; the edge feature fusion module includes an eleventh convolutional layer, a twelfth convolutional layer arranged in parallel, and a fourth normalization layer and a fourth non-linear activation layer connected thereto. The specific parameters are as follows: the kernel sizes of the eighth convolutional layer, the ninth convolutional layer, the tenth convolutional layer, and the eleventh convolutional layer are all 1*1, the sliding strides are all 1, the kernel size of the twelfth convolutional layer is 3*3, and the padding is 1; the kernel sizes of the third transposed convolutional layer, the fourth transposed convolutional layer, and the fifth transposed convolutional layer are all 3*3, the sliding strides are all 2, and the padding is 1; the fourth normalization layer uses batch normalization, and the fourth non-linear activation layer uses the ReLU function.

[0061] The Head module includes a thirteenth convolutional layer arranged in parallel and a fourteenth convolutional layer connected thereto, a fifteenth convolutional layer arranged in parallel and a sixteenth convolutional layer connected thereto, and a seventeenth convolutional layer arranged in parallel and an eighteenth convolutional layer connected thereto. The specific parameters are as follows: the kernel size of the thirteenth convolutional layer is 1*1, the sliding stride is 1, the fourteenth convolutional layer, the sixteenth convolutional layer, and the eighteenth convolutional layer are all dilated convolutions, the kernel sizes are all 3*3, and the dilation rates are 1, 3, and 5 respectively. The kernel size of the fifteenth convolutional layer is 3*3, the sliding stride is 1, the padding is 1, and the kernel size of the seventeenth convolutional layer is 5*5, the sliding stride is 1.

[0062] (3) Iteratively train the infrared small target detection model O:

[0063] (3a) Initialize the iteration number as t, the maximum iteration number as T = 1000, and the weight and bias parameters of the infrared small target detection model O in the t-th iteration are w t and b t respectively, and let t = 0, O t = O; t

[0064] (3b) Randomly select 8 infrared images from the training sample set R1 with replacement as the input of the infrared small target detection model O for forward propagation:

[0065] (3b1) The Stem block downsamples each infrared image to obtain the feature map x1. The first encoding sub-module further extracts features from the feature map x1 to obtain the feature map x2. The second encoding sub-module performs feature encoding on the feature map x2 to obtain the feature map x3. The second encoding sub-module performs feature encoding on the feature map x3 to obtain the feature map x4. The first transposed convolutional layer upsamples the feature map x4 to obtain the feature map x5. The feature map x2 and the feature map x5 are added element-wise to obtain the feature map x6. The second transposed convolutional layer upsamples the feature map x6 to obtain the feature map x7. The feature map x3 and the feature map x7 are added element-wise to obtain the feature map x8;

[0066] (3b2) The first edge detection sub-module performs edge detection on the feature map x2 to obtain the feature map x9. The second edge detection sub-module performs edge detection on the feature map x3 to obtain the feature map x 10 , and the third edge detection sub-module performs edge detection on the feature map x 10 to obtain the feature map x 11 . The feature map x9, the feature map x 10 and the feature map x 11 are merged in channels to obtain the feature map x 12 . The edge feature fusion module further extracts features from the feature map x 12 to obtain the feature map x 13 , that is, the target edge feature information is obtained;

[0067] (3b3) The feature map x8 and the feature map x 13 are merged in channels to obtain the feature map x 14 . The feature map x 14 passes through the Head module to obtain 8 infrared small target detection results.

[0068] (3c) Using the cross-entropy loss function, calculate the loss value of the edge extraction network in O t through the generated 8 target edge feature information and the labels of the corresponding 8 infrared edge images and use the Dice loss function to calculate the loss value L t of O Dice through the generated 8 infrared small target detection results and the labels of the corresponding 8 infrared images. Then, use the chain rule to calculate and the partial derivative of the total loss Lt composed of L Dice with respect to the weight parameter ω t and the bias parameter b t and Finally, according to for ω t , b tUpdate to obtain the network model O for this iteration t ;

[0069] (3d) Determine whether t≥1000 holds. If so, obtain the trained infrared small target detection model O*. Otherwise, set t = t + 1 and execute step (3b);

[0070] (4) Obtain the infrared image small target detection result:

[0071] Use the test sample set E1 as the input of the trained infrared small target detection model O* for forward propagation to obtain the infrared small target detection results corresponding to all test samples.

[0072] The technical effects of the present invention will be described below in combination with simulation experiments

[0073] Analysis of simulation conditions, content, and results:

[0074] The hardware platform for the simulation experiment is as follows: The processor is an Intel(R) Core i9-9900K CPU with a main frequency of 3.5 GHz, the memory is 32 GB, and the graphics card is an NVIDIA GeForce RTX 2080Ti. The software platform for the simulation experiment is: Ubuntu 16.04 operating system, python version is 3.7, and Pytorch version is 1.7.1.

[0075] Use the intersection over union (IoU) evaluation metric to compare the detection accuracy of the patent document "Infrared Small Target Detection Method Based on Global Mean Contrast Spatial Attention" (Patent Application No.: CN202211398750.X, Publication No. CN115527098A) and the present invention on test samples. The intersection over union (IoU) of the infrared small target detection results of the prior art is 62.37%, and the intersection over union of the infrared small target detection results of the present invention is 64.29%. Compared with the prior art, the detection accuracy of the present invention has been significantly improved.

Claims

1. An infrared small target detection method based on a Runge-Kutta residual block, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Obtain K infrared images, perform edge detection on each infrared image to obtain K infrared edge images, then label the targets in each infrared image and its corresponding infrared edge image to obtain the labels of K infrared images and the labels of K infrared edge images. Then, form a training sample set R1 with M infrared images and their corresponding labels, as well as the labels of M infrared edge images, and form a test sample set E1 with the remaining K - M infrared images and their corresponding labels, as well as the labels of K - M infrared edge images, where K ≥ 500. (2) Construct an infrared small target detection model based on the Runge-Kutta residual block: Construct an infrared small target detection model O including an encoder-decoder network and a Head module connected in sequence, and an edge extraction network, where: The encoder-decoder network includes a Stem module, an encoding module containing H encoding sub-modules, and a decoding module containing H-1 transposed convolutional layers connected in sequence; the Stem module includes a plurality of convolutional layers and a pooling layer connected in sequence; the encoding sub-module includes an embedding layer and a Runge-Kutta residual block connected in sequence; the output end of the h-th encoding sub-module is also connected to the output end of the (H-h)-th transposed convolutional layer; The edge extraction network is loaded between the output end of the encoding module and the input end of the Head module, and includes an edge detection module including H edge detection sub-modules arranged in parallel and an edge feature fusion module cascaded therewith; the input end of the h-th edge detection sub-module is connected to the output feature of the h-th encoding sub-module; the h-th edge detection sub-module includes a Sobel operator, a convolutional layer, and h-1 transposed convolutional layers connected in sequence; the edge feature fusion module includes two convolutional layers arranged in parallel, and the output end of one of the convolutional layers is also connected to a normalization layer and a non-linear activation layer; the Head module includes three branches arranged in parallel, and each branch includes two convolutional layers connected in sequence; (3) Iteratively train the infrared small target detection model O: (3a) Initialize the number of iterations to t, the maximum number of iterations to T, T ≥ 1000, the infrared small target detection model O of the tth iteration t The weight and bias parameters in are w t , b t , and let t = 0, O t =O; (3b) Randomly and with replacement select L infrared images from the training sample set R1 as the input of the infrared small target detection model O for forward propagation, where 1 ≤ L ≤ M: (3b1) The Stem module in the encoder-decoder network downsamples each infrared image, the H encoding sub-modules in the encoding module perform feature encoding on the result of each downsampling, and the H-1 transposed convolutional layers in the decoding module upsample the result of each feature encoding to obtain L target feature information; (3b2) The h-th edge detection sub-module in the edge extraction network performs edge detection on the output feature of the h-th encoding sub-module, and the edge feature fusion module performs feature fusion on the outputs of the H edge detection sub-modules to obtain L target edge feature information; (3b3) The Head module predicts the L multi-layer feature information obtained by combining the L target feature information and their corresponding L target edge feature information to obtain L infrared small target detection results; (3c) Calculate \(O\) by using the cross - entropy loss function, based on the \(L\) generated target edge feature information and the labels of the corresponding \(L\) infrared edge images t The loss value of the edge extraction network And calculate \(O\) by using the Dice loss function, based on the \(L\) generated infrared small target detection results and the labels of the corresponding \(L\) infrared images t The loss value \(L\) of Dice , and then calculate by using the chain rule The partial derivative of the total loss \(L_t\) composed of \(O\) and \(L\) Dice with respect to the weight parameter \(\omega\) t and the bias parameter \(b\) t ; and Finally, according to Update \(\omega\) t , \(b\) t to obtain the network model \(O\) for this iteration t ; (3d) Determine whether t ≥ T holds. If so, obtain the trained infrared small target detection model O*, otherwise, set t = t + 1 and execute step (3b); (4) Obtain the infrared image small target detection result: Use the test sample set E1 as the input of the trained infrared small target detection model O* for forward propagation to obtain the infrared small target detection results corresponding to all infrared images.

2. The infrared small target detection method based on the Runge-Kutta residual block according to claim 1, characterized in that For the infrared small target detection model O based on the Runge-Kutta residual block described in step (2), where: The number of convolutional layers contained in the Stem module is 3, the number of encoding sub-modules H contained in the encoding module is 3, and the encoding sub-module includes an embedding layer and a Runge-Kutta residual block connected in sequence. Among them, the embedding layer includes a segmentation encoding vector and a position encoding vector, and the Runge-Kutta residual block includes two feature extraction blocks connected in sequence. The feature extraction block includes an attention transformation block and a convolutional residual block arranged in parallel, and a channel attention module and a convolutional layer connected in sequence. Among them, the attention transformation block includes a randomly connected attention block, a normalization layer, a fully connected layer, a non-linear activation layer, a random dropout layer, a fully connected layer, a random dropout layer, and a normalization layer connected in sequence. The convolutional residual block includes three convolutional layers, a normalization layer, and a non-linear activation layer connected in sequence. The channel attention module includes a pooling layer arranged in parallel and two fully connected layers connected thereto, a pooling layer and two fully connected layers connected thereto, and a non-linear activation layer connected in sequence.

3. The infrared small target detection method based on the Runge-Kutta residual block according to claim 1, characterized in that For the infrared small target detection results of the L infrared images described in step (3b), the specific implementation steps of the acquisition process are as follows: (3b1) The implementation steps for the encoding and decoding network to obtain L target feature information are as follows: (3b11) The Stem module downsamples each infrared image to obtain a feature map x1; (3b12) The h-th encoding sub-module in the encoding module performs feature encoding on the feature map x h to obtain the feature map x h+1 , and the encoding module outputs the feature map x H+1 ; (3b13) The H-1 transposed convolutional layers in the decoding module perform upsampling on the feature map x H+1 layer by layer. The output feature map of the (H-h)th transposed convolutional layer is also element-wise added to the output feature map of the hth encoding sub-module, and the decoding module outputs the feature map x 3H-1 ; (3b2) The implementation steps for the edge extraction network to obtain L target edge feature information are as follows: (3b21) The h-th edge detection sub-module performs edge detection on the output feature map x of the h-th encoding sub-module h+1 to obtain the feature map x 3H+h-1 ; (3b22) Merge the output feature maps of H edge detection sub-modules to obtain the feature map x 4H , and the edge feature fusion module performs feature fusion on the feature map x 4H to obtain the feature map x 4H+1 ; (3b3)The Head module makes predictions on the feature map x 3H-1 and the feature map x 4H+1 to obtain the feature map x obtained by channel merging 4H+2 and gets L infrared small target detection results.

4. The infrared small target detection method based on the Runge-Kutta residual block according to claim 1, characterized in that The expression for calculating the loss value Lt described in step (3c), and according to for ω t , b t The update formulas for updating are respectively: Among them, Y represents the output image of the network, and Y GT represents the label of the infrared image, X represents the output image of the edge extraction network, and X GT represents the label of the infrared edge image, β and λ are weight coefficients, and ω t , b t represents the weights and biases of all learnable parameters, and w t represents the weights and biases of all learnable parameters, and w t ', b t ' represents the updated learnable parameters, and α r represents the learning rate.

Citation Information

Patent Citations

  • Infrared small target detection method based on global mean contrast space attention

    CN115527098A

  • Infrared small target detection method based on bidirectional attention aggregation mechanism

    CN114882322A