Highly Reliable Small Target Detection Method in Complex Environments

By constructing a weak object detection network inspired by higher-order differential equations, using differential hierarchical generation and feature fusion modules, the fusion and robustness problems of weak object detection in complex backgrounds are solved, and efficient object detection is achieved.

CN117253117BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202310826170.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-07-29
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect weak targets in complex contexts, the filters and convolutional neural networks are poorly integrated, the network is poorly regular, the robustness is insufficient, and the data sets are strict, resulting in a high false alarm rate.

Method used

Build a weak object detection network inspired by higher-order differential equations, including a differential hierarchical generation module, a feature fusion module, a fourth-order Adams boot module and a Predict module. Through multiple differential operations and feature fusion, we use the prior knowledge of the scale and translation invariance of the image to reduce dependence on large-scale data sets and avoid overfitting.

Benefits of technology

It improves the weak target detection rate, suppresses background aliasing, reduces false alarm rate, and enhances the accuracy and robustness of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253117B_ABST
    Figure CN117253117B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting highly reliable small targets in a complex environment. The implementation steps are as follows: construct and train a small target detection network inspired by a higher-order differential equation, use the difference layer generation module in the network to perform multiple difference operations on the image to obtain small target features at different scales and different levels. The fourth-order Adams guidance module in the small target detection network inspired by the higher-order differential equation establishes connections between multiple feature terms, constructing a more powerful interpretable network. The present invention mainly solves the detection of small targets in images with low signal-to-noise ratio in a complex background, and has the advantages of suppressing noise, enhancing the visibility of targets, accurately extracting different-level features by the network to obtain the features of small targets at all levels, and having a low false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and further relates to a high-trust small target detection method in a complex environment in the field of image detection technology. The present invention uses the high-trust small target detection method in a complex environment to detect small and weak targets in an image, and can be used to detect small and weak targets in images with low signal-to-noise ratio under complex backgrounds. Background Art

[0002] Due to the interference of noise in the image receiver and the relatively long distance between the target to be measured and the detector, the available features of the target to be measured are few, the scale is small, there is no texture, and the detailed information is easily lost. It has a very small proportion of the area in the image, resulting in an increased difficulty in detecting small and weak targets. However, it is difficult for the existing technology to extract clear detailed features under complex backgrounds, and it is easy to lose small target features, resulting in missed detections and false alarms. The performance of small target detection still needs to be improved.

[0003] Liaoning Technical University disclosed an infrared small target detection method based on background suppression and feature fusion in its patent document "An Infrared Small Target Detection Method Based on Background Suppression and Feature Fusion" (Patent Application No.: 202211540004.X, Publication No. CN 115830502A). The specific steps of this method are as follows: First, the original infrared image is input into a three-layer window local contrast mechanism module, which divides the window into three layers: the central layer, the middle layer, and the outermost layer. Second, Gaussian filtering is performed on the central layer, the outermost layer of the window is divided into 8 directions, and the value closest to the center is selected as the background value for contrast calculation. Third, the difference joint contrast calculation is performed between the matched filtering result and the closest filtering result to highlight the target and suppress the background. Fourth, the background clutter is further suppressed through a convolutional neural network. Fifth, an adaptive threshold operation is used to perform a binarization operation on the result of the non-negative constraint calculation to extract the region of interest. The disadvantages of this method are that, since only a Gaussian filter is used to filter and denoise the image in the second step of this method, the method is single and the filter cannot be well integrated with the convolutional neural network, and the network regularity is poor. At the same time, the binarization operation is extremely likely to cause the loss of key features containing small and weak targets, resulting in poor feature fusion effect and low detection rate of small and weak targets in the image.

[0004] Chongqing University of Technology and the Academy of Ordnance Science of China disclosed an infrared dim and small target detection method based on Swin-Transformer and multi-scale feature fusion in their jointly applied patent document "An Infrared Dim and Small Target Detection Method Based on Swin-Transformer and Multi-Scale Feature Fusion" (Patent Application No. 202310205449.0, Publication No. CN 116188944 A). The specific steps of this method are as follows: In the first step, a Swin-Transformer module is introduced into the Unet network to replace the original convolutional layer for feature extraction to form a target detection model. In the second step, the infrared image to be detected is input into the trained target detection model, and the feature information of the infrared image is extracted layer by layer through multiple Swin-Transformer modules to generate feature maps of multiple scales. In the third step, starting from the feature map of the highest scale, multiple cross-layer feature fusion modules are used to fuse the feature maps of each scale in turn to generate corresponding multi-layer fusion feature maps. In the fourth step, the multi-layer fusion feature maps are input into a classifier for normalization processing and the corresponding target prediction results are output. The deficiencies of this method are as follows: In the first step of this method, when introducing the Swin-Transformer module, it is impossible to utilize prior knowledge such as the scale, translation invariance, and feature locality inherent in the image. Moreover, training this neural network must use a large-scale dataset, and the requirements for the dataset are extremely strict, which easily leads to the network's difficulty in capturing the feature information of small targets in the image, poor robustness, and an increased false alarm rate. Summary of the Invention

[0005] The object of the present invention is to address the deficiencies existing in the above-mentioned prior art, and propose a high-trust small target detection method for complex environments to solve the problems that the filter cannot be well integrated with the convolutional neural network and the network has poor regularity; the problem of being unable to utilize prior knowledge such as the scale, translation invariance, and feature locality inherent in the image; the problem that it is necessary to use a large-scale dataset, and the requirements for the dataset are extremely strict, resulting in difficulty in capturing the feature information of medium and small targets in the image, poor robustness, and a high false alarm rate.

[0006] The technical idea for achieving the object of the present invention is that the differential layer generation module constructed by the present invention uses a convolutional neural network to perform multiple differential operations on the input image to obtain small target features at different scales and levels, overcoming the defect that the filters in the prior art cannot be well integrated with the convolutional neural network, and improving the detection rate of small and weak targets in the image. The feature fusion module constructed by the present invention maps all features of different scales to the low layer through upsampling operations for three levels of features with different detail information and semantic features, and fuses features at different levels using feature weight factors, effectively solving the defect that the prior art cannot utilize prior knowledge such as scale, translation invariance, and feature locality inherent in the image itself. The fourth-order Adams guidance module constructed by the present invention applies the method of solving ordinary differential equations through the fourth-order Adams implicit equation to three residuals, establishing connections between more feature terms and constructing a more powerful interpretable network. It effectively solves the defect of poor network regularity in the prior art. The Predict module designed by the present invention includes a random dropout layer, which reduces the interaction between neurons and avoids the phenomenon of overfitting when using a small-scale dataset, effectively solving the defect that the prior art must use a large-scale dataset.

[0007] To achieve the above object, the specific implementation steps of the present invention are as follows:

[0008] Step 1, construct a feature fusion module:

[0009] Build a feature fusion module including a high-level feature branch, a middle-level feature branch, a low-level feature branch, a feature weight fusion unit, and an output layer; the outputs of the three branches of the high-level feature branch, the middle-level feature branch, and the low-level feature branch are fused through the feature weight fusion unit and then output through the output layer;

[0010] Step 2, construct a fourth-order Adams guidance module:

[0011] Build a fourth-order Adams guidance module including a first differential input group, a second differential input group, a negative unit, an addition unit, an inner product unit, a merging unit, and an output layer; the first differential input group, the negative unit, the addition unit, the inner product unit, the merging unit, and the output layer are connected in series in sequence; the second differential input group is connected to the addition unit;

[0012] Step 3, construct a differential layer generation module:

[0013] Build a differential layer generation module in which a negative third-order differential input layer, a first differential transformation unit, a negative second-order differential output layer, a second differential transformation unit, a negative first-order differential output layer, a third differential transformation unit, and a zero-order differential output layer are connected in series in sequence, and the structures and parameters of the first to third differential transformation units are the same;

[0014] Step 4, construct the Predict module:

[0015] Build a Predict module composed of a first convolutional layer, a normalization layer, an activation layer, a random dropout layer, and a second convolutional layer connected in series in sequence;

[0016] Set the convolutional kernel sizes of the first convolutional layer and the second convolutional layer to 3×3, the sliding strides to 1, the output channels to 16, and use the ReLU function as the activation function for the activation layer;

[0017] Step 5, construct the U-Net sub-module:

[0018] Build a U-Net sub-module including an input layer, a first residual convolutional group, a second residual convolutional group, a third residual convolutional group, a first merging unit, a first residual deconvolutional group, a second merging unit, a second residual deconvolutional group, a first output layer, a second output layer, and a third output layer. Among them, the input layer, the first residual convolutional group, the second residual convolutional group, the third residual convolutional group, the first merging unit, the second merging unit, and the second residual deconvolutional group are connected in series in sequence; the first residual convolutional group is also connected to the second merging unit, the second residual convolutional group is also connected to the first merging unit, the third residual convolutional group is also connected to the first output layer, the first residual deconvolutional group is also connected to the second output layer, and the second residual deconvolutional group is also connected to the third output layer; the first to third residual convolutional groups are composed of three identical residual convolutional units connected in series in sequence; the first and second residual deconvolutional groups are composed of three identical residual deconvolutional units connected in series in sequence;

[0019] Step 6, construct a small and weak target detection network inspired by high-order differential equations:

[0020] Build a small and weak target detection network inspired by high-order differential equations composed of an input layer, a convolutional layer, a target feature enhancement module, a U-Net sub-module, a feature fusion module, a difference layer generation module, a fourth-order Adams guidance module, a Predict module, and an output layer connected in series in sequence; set the convolutional kernel size of the convolutional layer to 3×3, the sliding stride to 1, and the output channels to 16;

[0021] Step 7, generate a training set;

[0022] Select at least 400 images to form a sample set; after performing random mirroring, random scaling, extended cropping, and Gaussian blurring operations on each image in the sample set in sequence, obtain processed images with a specification of 480×480 for each image; input each processed image into the torchvision module for normalization processing, and form a training set with all the normalized samples;

[0023] Step 8, training a small and weak target detection network inspired by high-order differential equations:

[0024] Input the training set into the small and weak target detection network inspired by high-order differential equations. Use the dice hybrid loss function and AdaGrad as the optimizer. Optimize the network parameters through the gradient optimization algorithm of stochastic gradient descent, and iteratively update the weight values of the network until the dice hybrid loss function of the network converges, obtaining a trained small and weak target detection network inspired by high-order differential equations.

[0025] Step 9, detecting small and weak targets:

[0026] Adopt the same method as in Step 7 to process the small and weak target images to be detected; input each processed image into the trained small and weak target detection network inspired by high-order differential equations for detection; regard the region composed of multiple pixels with pixel value 1 in the image output by the network as the detected small target region, and the pixels with pixel value 0 represent the region where small targets are not detected.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] First, the differential layer generation module in the small and weak target detection network constructed by the present invention performs multiple convolutional difference operations on the obtained multi-scale features in the convolutional neural network, overcoming the defect that the filters in the prior art cannot be well integrated with the convolutional neural network, so that the present invention improves the detection rate of small and weak targets in images.

[0029] Second, the feature fusion module in the small and weak target detection network designed by the present invention adopts the exponential autocorrelation fusion method when fusing three different levels and different semantic features, overcoming the defect that the prior art cannot utilize the prior knowledge such as scale, translation invariance and feature locality inherent in the image, so that the present invention can better adaptively learn the spatial weights of feature fusion at each scale and suppress the background aliasing.

[0030] Third, the fourth-order Adams guidance module in the small and weak target detection network designed by the present invention constructs a more powerful interpretable network. The high-order differential equation relates multiple feature terms, overcoming the defect of poor network regularity in the prior art, so that the present invention avoids the loss of small target information during the network propagation process and reduces the false alarm rate.

[0031] Fourth, the Predict module in the small and weak target detection network designed by the present invention includes a random dropout layer, which overcomes the defect that the prior art must use a large-scale data set and avoids the phenomenon of overfitting when using a small-scale data set. Therefore, the present invention has the advantages of improving the training effect, making the network easily capture the small target feature information in the image, and improving the robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flowchart of the implementation of the present invention;

[0033] Figure 2 is a schematic structural diagram of the feature fusion module of the present invention;

[0034] Figure 3 is a schematic structural diagram of the fourth-order Adams guidance module of the present invention;

[0035] Figure 4 is a schematic structural diagram of the difference layer generation module of the present invention;

[0036] Figure 5 is a schematic structural diagram of the Predict module of the present invention;

[0037] Figure 6 is a schematic structural diagram of the U-Net sub-module of the present invention;

[0038] Figure 7 is a schematic structural diagram of the small and weak target detection network inspired by the high-order differential equation of the present invention;

[0039] Figure 8 is a schematic diagram of the image to be measured in the embodiment of the present invention;

[0040] Figure 9 is a schematic diagram of the true position and shape of the small and weak target in the image to be measured in the embodiment of the present invention;

[0041] Figure 10 is a schematic diagram of the position and shape of the small and weak target detected in the image to be measured in the embodiment of the present invention;

[0042] Figure 11 is a simulation result diagram of the simulation experiment of the present invention;

[0043] Figure 12 is a simulation result diagram of the simulation experiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0044] The following describes the present invention in further detail with reference to the drawings and embodiments.

[0045] Refer to Figure 1 , and further describe the implementation steps of the embodiment of the present invention.

[0046] Step 1, construct a feature fusion module.

[0047] Refer to Figure 2 for a further description of the structure of the feature fusion module constructed by the present invention.

[0048] Step 1.1, build a feature fusion module including a high-level feature branch, a middle-level feature branch, a low-level feature branch, a feature weight fusion unit, and an output layer. Among them:

[0049] The high-level feature branch is composed of a first input layer, a first deconvolution layer, and a first convolution layer connected in series in sequence. The middle-level feature branch is composed of a second input layer, a second deconvolution layer, and a second convolution layer connected in series in sequence. The low-level feature branch is composed of a third input layer and a third convolution layer connected in series in sequence. The high-level feature branch, the middle-level feature branch, and the low-level feature branch are fused through the feature weight fusion unit and then output through the output layer.

[0050] Step 1.2, set the channel parameters of the first to third input layers to 64, 32, and 16 respectively; set the convolution kernel sizes in the first and second deconvolution layers to 3×3 and the sliding strides to 1; set the convolution kernel sizes in the first to third convolution layers to 3×3 and the sliding strides to 1; set the channel parameter of the output layer to 16.

[0051] Step 2, construct a fourth-order Adams guidance module.

[0052] Refer to Figure 3 for a further description of the structure of the fourth-order Adams guidance module constructed by the present invention.

[0053] Step 2.1, build a fourth-order Adams guidance module including a first difference input group, a second difference input group, a negative unit, an addition unit, an inner product unit, a merging unit, and an output layer. The first difference input group, the negative unit, the addition unit, the inner product unit, the merging unit, and the output layer are connected in series in sequence; the second difference input group is connected to the addition unit. Among them, the first difference input group is composed of a negative third-order difference input layer, a negative second-order difference layer, and a negative first-order difference input layer connected in parallel; the second difference input group is composed of a negative second-order difference input layer, a negative first-order difference input layer, and a zero-order difference input layer connected in parallel. The zero-order difference input layer in the second difference input group is connected to the merging unit. Set the product factor of the negative unit to -\\(1\\).

[0054] Step 3, construct a difference layer generation module.

[0055] Step 3.1, build a differential layer generation module composed of a negative third-order difference input layer, a first differential transformation unit, a negative second-order difference output layer, a second differential transformation unit, a negative first-order difference output layer, a third differential transformation unit, and a zero-order difference output layer connected in series in sequence. The structures and parameters of the first to third differential transformation units are the same, and its structure is composed of a first convolutional layer, a first activation layer, a second convolutional layer, and a second activation layer connected in series in sequence.

[0056] Step 3.2, set the convolutional kernel sizes of the first and second convolutional layers to 3×3, the sliding strides to 1, and the output channels to 16. The activation functions of the first activation layer and the second activation layer both use the ReLU function.

[0057] Step 4, build a Predict module.

[0058] Refer to Figure 5 , and further describe the structure of the Predict module constructed in the present invention.

[0059] Step 4.1, build a Predict module composed of a first convolutional layer, a normalization layer, an activation layer, a random dropout layer, and a second convolutional layer connected in series in sequence.

[0060] Step 4.2, set the convolutional kernel sizes of the first convolutional layer and the second convolutional layer to 3×3, the sliding strides to 1, and the output channels to 16. The activation function of the activation layer uses the ReLU function;

[0061] Step 5, build a U-Net sub-module.

[0062] Refer to Figure 6 , and further describe the structure of the U-Net sub-module constructed in the present invention.

[0063] Build a U-Net sub-module including an input layer, a first residual convolutional group, a second residual convolutional group, a third residual convolutional group, a first merging unit, a first residual deconvolutional group, a second merging unit, a second residual deconvolutional group, a first output layer, a second output layer, and a third output layer. Among them, the input layer, the first residual convolutional group, the second residual convolutional group, the third residual convolutional group, the first merging unit, the second merging unit, and the second residual deconvolutional group are connected in series in sequence; the first residual convolutional group is also connected to the second merging unit, the second residual convolutional group is also connected to the first merging unit, the third residual convolutional group is also connected to the first output layer, the first residual deconvolutional group is also connected to the second output layer, and the second residual deconvolutional group is also connected to the third output layer.

[0064] The first to third residual convolutional groups are composed of three residual convolutional units with the same structure connected in series in sequence.

[0065] The residual convolution unit is composed of a residual input layer, a first convolution layer, a first pooling layer, a second convolution layer, a second pooling layer, an addition unit, and a residual output layer connected in series in sequence. The residual input layer is also connected to the addition unit. The activation functions of the first and second pooling layers adopt the ReLU function. The convolution kernel sizes of the first and second convolution layers are both set to 3×3, and the sliding strides are both set to 1. The output channels of the first to third residual convolution groups are set to 16, 32, and 64 respectively.

[0066] The first and second residual deconvolution groups are composed of three residual deconvolution units with the same structure connected in series in sequence.

[0067] The residual deconvolution unit is composed of a residual input layer, a first deconvolution layer, a first pooling layer, a second deconvolution layer, a second pooling layer, an addition unit, and a residual output layer connected in series in sequence. The residual input layer is also connected to the addition unit. The activation functions of the first and second pooling layers adopt the ReLU function. The convolution kernel sizes of the first and second deconvolution layers are both set to 3×3, and the sliding strides are both set to 1. The output channels of the first and second residual deconvolution groups are set to 32 and 16 respectively.

[0068] Step 6, construct a small and weak target detection network inspired by high-order differential equations.

[0069] Refer to Figure 7 , and further describe the structure of the small and weak target detection network inspired by high-order differential equations constructed by the present invention.

[0070] Step 6.1, build a small and weak target detection network inspired by high-order differential equations, which is composed of an input layer, a convolution layer, a target feature enhancement module, a U-Net sub-module, a feature fusion module, a difference layer generation module, a fourth-order Adams guidance module, a Predict module, and an output layer connected in series in sequence. The structure and parameters of the target feature enhancement module are detailed in the residual feature compensation module in the publicly disclosed patent document "Multi-scale forward feature gain infrared small and weak target detection method in complex environments" (Patent Application No. 202211167891.0, Publication No. CN 115661443A). Set the convolution kernel size of the convolution layer to 3×3, the sliding stride to 1, and the output channel to 16.

[0071] Step 7, generate a training set and a test set.

[0072] Step 7.1, in the embodiment of the present invention, 427 images are selected from the publicly available SIRST dataset to form a sample set. 80% of the images in the sample set are used to form a training sample set, and 20% of the images are used to form a test sample set.

[0073] Step 7.2, to ensure the training effect of the network, perform data augmentation on the images in the training sample set. That is, for each image in the training sample set, perform random mirroring, random scaling, extended cropping, and Gaussian blurring operations in sequence to obtain processed images with a specification of 480×480 for each image.

[0074] Step 7.3, input each processed image into the torchvision module for normalization processing, and form a training set with all the normalized training samples. Obtain a test set for the test sample set using the same method.

[0075] Step 8, train a small and weak target detection network inspired by high-order differential equations.

[0076] Input the training set into the small and weak target detection network inspired by high-order differential equations. For the training set, use the dice hybrid loss function, use AdaGrad as the optimizer, with a learning rate of 0.05, and adopt the MSRA initialization method for the weight initialization strategy. The training process consists of 3000 epochs in total, with a weight decay of 10 -4 , and the batch size is 32. Calculate the loss value of the small and weak target detection network inspired by high-order differential equations after the selected images are input into it. Use the gradient optimization algorithm of the stochastic gradient descent method to optimize the network parameters, and iteratively update the weight values of the small and weak target detection network inspired by high-order differential equations until the dice hybrid loss function of the network converges, obtaining a trained small and weak target detection network inspired by high-order differential equations.

[0077] The described dice hybrid loss function is as follows:

[0078]

[0079] where loss represents the network loss value output after one round of iteration for the images input into the small and weak target detection network inspired by high-order differential equations, ∑ represents the summation operation, N represents the total number of image samples in the training set, n represents the index of the image samples in the training set, p n represents the probability that each pixel in the predicted input image belongs to the true target, r n represents the category of each pixel in the input image, ∈ represents a correction factor, and the value of ∈ is any real number selected within the range of (0, 0.1). Its function is to prevent the denominator of the fraction from being zero.

[0080] Step 9, detect small and weak targets.

[0081] The input in the test set is fed into the trained small target detection network inspired by high-order differential equations to detect small targets in the image. The region composed of multiple pixels with pixel value 1 in the image output by the network is used as the detected small target region, and the pixels with pixel value 0 represent the regions where small targets are not detected.

[0082] The following combines the embodiments of the present invention and Figure 8 、 9 、10 to further describe the present invention.

[0083] Figure 8 In the embodiment of the present invention, an original image without any preprocessing is randomly selected from the test sample set. The small target to be detected in this image is located next to a strong light source and occupies very few pixels, making it difficult to observe and identify and prone to false alarms. After preprocessing this image, it is input into the small target detection network inspired by high-order differential equations constructed and trained by the present invention. The small target detection result output by this network is as Figure 9 shown. Figure 9 The white small pixels on the left in Figure 8 show the position of the small target detected in Figure 9 The shape of the detected small target is shown in the gray square in the lower right corner of

[0084] Figure 10 This is the true position and true shape diagram of the small target corresponding to the image to be tested in the test sample set of the present invention. Figure 8 The white small pixels on the left in Figure 10 are the true position information of the small target in the image to be tested, and the true shape of the small target is shown in the white square in the lower right corner.

[0085] Comparing Figure 9 with Figure 10 it is obvious that the similarity between the position and shape information of the small target detected by the method of the present invention and the true information is extremely high, indicating that the small target detection network inspired by high-order differential equations constructed and trained by the present invention has obvious effects in suppressing the background and enhancing the target, the differential layer generation module extracts small target features accurately, and the fourth-order Adams guidance module fuses the features of each layer well, greatly improving the accuracy in small target detection.

[0086] The effect of the present invention can be further demonstrated by the following simulation.

[0087] 1. Simulation experiment conditions.

[0088] The software platform for the simulation experiment of the present invention uses the Linux5 operating system and the professional version of Pycharm2021.1, and the hardware platform uses NVIDIA RTX3090.

[0089] 2. Simulation content and result analysis.

[0090] In the first simulation experiment of the present invention, five pictures randomly selected from the test sample set are preprocessed and then input into the network trained by the present invention and four publicly available trained networks of the prior art (GAU network, ACM network, ALC network, FC3 network) respectively for small target detection, and five groups of 25 small target detection result pictures are obtained. Comparing the five pictures randomly selected from the test set, the 25 predicted result pictures with the 5 real small target pictures, the 7 groups of 35 pictures in total are as Figure 11 shown. The position and shape of the small target are shown in the light gray box in each picture.

[0091] The detection results are as Figure 11 shown. It can be seen from the figure that for real infrared images, the prior art can detect small targets, but there are still large differences between the detected small targets and the true label values, and false alarm phenomena occur. The network proposed by the present invention can detect contours similar to the true label values, retain the angle and detail information of the small targets, and the detection results are significantly better.

[0092] The prior art adopted in the simulation experiment of the present invention refers to:

[0093] The GAU network refers to the GAU network proposed by Weizhe Hua et al. in their paper "Transformer Quality in Linear Time" (10.48550 / arXiv.2202.10447) in a method for small target detection.

[0094] The ACM network refers to the ACM network proposed by Yimian Dai et al. in their paper "Asymmetric Contextual Modulation for Infrared Small Target Detection" (Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 950-959) in a method for small target detection.

[0095] The ALC network refers to the ALC network proposed by Yimian Dai et al. in their paper "Attentional Local Contrast Networks for Infrared Small Target Detection" (IEEE Transactions on Geoscience and Remote Sensing, 05 January 2021, 9813-9824) in a method for detecting small and weak targets.

[0096] The FC3 network refers to the FC3 network proposed by Mingjin Zhang et al. in their paper "Exploring Feature Compensation and Cross-level Correlation for Infrared Small Target Detection" (MM'22: Proceedings of the 30th ACM International Conference on Multimedia October 2022 Pages 1857–1865) in a method for detecting small and weak targets.

[0097] In Simulation Experiment 2 of the present invention, the test set is input into the network trained by the present invention and four trained networks disclosed in the prior art (IPI, PSTNN, ALC network, FC3 network) respectively for small and weak target detection.

[0098] The prior art adopted in Simulation Experiment 2 of the present invention refers to:

[0099] The IPI method refers to the IPI method proposed by Landan Zhang et al. in their paper "Infrared Patch-Image Model for Small Target Detection in a Single Image" (IEEE Transactions on Image Processing Volume:22, Issue:12, December 2013 4996-5009) in a method for detecting small and weak targets.

[0100] PSTNN refers to the PSTNN method proposed by Hong Zhang et al. in their published paper "Infrared Small Target Detection Based on Partial Sum of the Tensor Nuclear Norm" (Remote Sensing for Target Object Detection and Identification, 13 February 2019) in a small and weak target detection method.

[0101] In the simulation experiment of the present invention, evaluation indicators such as intersection over union (IoU), normalized intersection over union (nIoU), receiver operating characteristic curve (ROC), detection rate (Pd), and false alarm rate (Fa) are used to evaluate existing small and weak target detection methods. The definitions of IoU and nIoU are as follows:

[0102]

[0103] Among them, T, P, and TP respectively represent the number of true pixels, pixels predicted correctly, and pixels predicted correctly and being true. N represents the total number of image samples in the training set, and i represents the index of the image samples in the training set. The larger the values of IoU and nIoU, the better the detection performance of the network.

[0104] The true positive rate (P d ) represents the proportion of true values predicted correctly in the total true values, and the false positive rate (F a ) represents the proportion of false values predicted correctly in the total false values:

[0105]

[0106] Among them, FP, TN, and FN respectively represent the number of pixels predicted correctly and being false, pixels predicted wrongly and being false, and pixels predicted wrongly and being true. The larger the true positive rate obtained from the test of a network, the better the detection effect, and the smaller the false positive rate, the better the detection effect. The ROC curve describes the dynamic relationship between P d and F a . For the ROC indicator, as the false positive rate increases in the ROC function graph line of a network, the larger the true positive rate, the better the detection result of the network.

[0107] The test set is respectively input into the network trained in the present invention and four networks trained and publicly disclosed in the prior art. The obtained output pictures of small and weak target detection results are processed to obtain evaluation indicators such as IoU, nIoU, P d and F a . The evaluation indicators are shown in Table 1. Regarding P d , F aFive ROC curves are drawn as Figure 12 shown below.

[0108] Table 1: Comparison Table of Evaluation Indicators

[0109]

[0110] As can be seen from Table 1, compared with the four existing technologies, the IoU, nIoU, and P obtained from the network test proposed by the present invention d are all the largest, and F a is the smallest. All four evaluation indicators are better than the existing technologies. As Figure 12 can be seen, the ROC function graph line corresponding to the network of the present invention is higher than that of the existing technology, which proves that the network proposed by the present invention has the best effect in detecting small and weak targets. In summary, the present invention is superior to the existing technology models.

Claims

1. A high-trust small target detection method for complex environments, characterized in that, A small and weak target detection network inspired by high-order differential equations is composed of a feature fusion module, a fourth-order Adams guidance module, a difference layer generation module, and a Predict module constructed; the steps of this detection method are as follows: Step 1, construct a feature fusion module: Build a feature fusion module including a high-level feature branch, a middle-level feature branch, a low-level feature branch, a feature weight fusion unit, and an output layer; The outputs of the three branches of the high-level feature branch, the middle-level feature branch, and the low-level feature branch are fused by the feature weight fusion unit and then output through the output layer; Step 2, construct a fourth-order Adams guidance module: Build a fourth-order Adams guidance module including a first difference input group, a second difference input group, a negative unit, an addition unit, an inner product unit, a merging unit, and an output layer; the first difference input group, the negative unit, the addition unit, the inner product unit, the merging unit, and the output layer are connected in series in turn; the second difference input group is connected to the addition unit; Step 3, construct a difference layer generation module: Build a difference layer generation module composed of a negative third-order difference input layer, a first difference transformation unit, a negative second-order difference output layer, a second difference transformation unit, a negative first-order difference output layer, a third difference transformation unit, and a zero-order difference output layer connected in series in turn, and the structures and parameters of the first to third difference transformation units are the same; Step 4, construct a Predict module: Build a Predict module composed of a first convolutional layer, a normalization layer, an activation layer, a random dropout layer, and a second convolutional layer connected in series in turn; Set the convolutional kernel sizes of the first convolutional layer and the second convolutional layer to 3×3, the sliding strides to 1, and the output channels to 16, and use the ReLU function as the activation function of the activation layer; Step 5, construct a U-Net sub-module: Build a U-Net sub-module including an input layer, a first residual convolutional group, a second residual convolutional group, a third residual convolutional group, a first merging unit, a first residual deconvolutional group, a second merging unit, a second residual deconvolutional group, a first output layer, a second output layer, and a third output layer. Among them, the input layer, the first residual convolutional group, the second residual convolutional group, the third residual convolutional group, the first merging unit, the first residual deconvolutional group, the second merging unit, and the second residual deconvolutional group are connected in series in turn; the first residual convolutional group is also connected to the second merging unit, the second residual convolutional group is also connected to the first merging unit, the third residual convolutional group is also connected to the first output layer, the first residual deconvolutional group is also connected to the second output layer, and the second residual deconvolutional group is also connected to the third output layer; the first to third residual convolutional groups are composed of three residual convolutional units with the same structure connected in series in turn; the first and second residual deconvolutional groups are composed of three residual deconvolutional units with the same structure connected in series in turn; Step 6, construct a small and weak target detection network inspired by high-order differential equations: Build a weak small target detection network inspired by high-order differential equations, which is composed of an input layer, a convolutional layer, a target feature enhancement module, a U-Net sub-module, a feature fusion module, a difference layer generation module, a fourth-order Adams guidance module, a Predict module, and an output layer connected in series in turn; set the convolutional kernel size of the convolutional layer to 3×3, the sliding step size to 1, and the output channels to 16; Step 7, generate a training set; Select at least 400 images to form a sample set; after performing random mirroring, random scaling, extended cropping, and Gaussian blur operations on each image in the sample set in turn, obtain processed images with a specification of 480×480 for each image; input each processed image into the torchvision module for normalization processing, and all the samples after normalization processing form a training set; Step 8, train the weak small target detection network inspired by high-order differential equations: Input the training set into the weak small target detection network inspired by high-order differential equations, use the dice hybrid loss function, use AdaGrad as the optimizer, and optimize the network parameters through the gradient optimization algorithm of stochastic gradient descent method, and iteratively update the weight values of the network until the dice hybrid loss function of the network converges, and obtain the trained weak small target detection network inspired by high-order differential equations; Step 9, detect weak small targets: Adopt the same method as in Step 7 to process the weak small target image to be detected; input each processed image into the trained weak small target detection network inspired by high-order differential equations to detect the image; regard the area composed of multiple pixels with pixel value 1 in the image output by the network as the detected small target area, and the pixels with pixel value 0 represent the area where small targets are not detected.

2. The high-trust small target detection method for complex environments according to claim 1, wherein The high-level feature branch described in Step 1 is composed of an input layer, a deconvolution layer, and a convolutional layer connected in series in turn; set the channel parameter of the input layer to 64, set the convolutional kernel size in the deconvolution layer to 3×3, the sliding step size to 1, and set the convolutional kernel size in the convolutional layer to 3×3, the sliding step size to 1.

3. The high-trust small target detection method for complex environments according to claim 1, wherein, The middle-level feature branch described in Step 1 is composed of an input layer, a deconvolution layer, and a convolutional layer connected in series in turn; set the channel parameter of the input layer to 32, set the convolutional kernel size in the deconvolution layer to 3×3, the sliding step size to 1, and set the convolutional kernel size in the convolutional layer to 3×3, the sliding step size to 1.

4. The high-confidence small target detection method for complex environments according to claim 1, characterized in that The low-level feature branch described in Step 1 is composed of an input layer and a convolutional layer connected in series in turn; set the channel parameter of the input layer to 16, set the convolutional kernel size in the convolutional layer to 3×3, the sliding step size to 1.

5. The high-confidence small target detection method for complex environments according to claim 1, wherein The first difference input group described in Step 2 is composed of a negative third-order input layer, a negative second-order difference input layer, and a negative first-order difference input layer connected in parallel.

6. The high-confidence small target detection method for complex environments according to claim 1, wherein The second difference input group described in Step 2 is composed of a negative second-order difference input layer, a negative first-order difference input layer, and a zero-order difference input layer connected in parallel. The zero-order difference input layer in the second difference input group is connected to the merging unit; set the product factor of the negative unit to -1.

7. The high-trust small target detection method for complex environments according to claim 1, wherein The structures of the differential transformation units described in step 3 are connected in series in turn as follows: the first convolutional layer, the first activation layer, the second convolutional layer, and the second activation layer; the convolutional kernel sizes of the first and second convolutional layers are both set to 3×3, the sliding strides are both set to 1, the output channels are both set to 16, and the activation functions of the first activation layer and the second activation layer both use the ReLU function.

8. The high-trust small target detection method for complex environments according to claim 1, wherein The residual convolutional unit described in step 5 is composed of a residual input layer, a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, an addition unit, and a residual output layer connected in series in turn, and the residual input layer is also connected to the addition unit; the activation functions of the first and second pooling layers use the ReLU function, the convolutional kernel sizes of the first and second convolutional layers are both set to 3×3, the sliding strides are both set to 1, and the output channels of the first to third residual convolutional groups are set to 16, 32, and 64 respectively.

9. The high-trust small target detection method for complex environments according to claim 1, wherein The residual deconvolutional unit described in step 5 is composed of a residual input layer, a first deconvolutional layer, a first pooling layer, a second deconvolutional layer, a second pooling layer, an addition unit, and a residual output layer connected in series in turn, and the residual input layer is also connected to the addition unit; the activation functions of the first and second pooling layers use the ReLU function, the convolutional kernel sizes of the first and second deconvolutional layers are both set to 3×3, the sliding strides are both set to 1, and the output channels of the first and second residual deconvolutional groups are set to 32 and 16 respectively.

10. The high-confidence small target detection method for complex environments according to claim 1, characterized in that, The dice hybrid loss function described in step 8 is as follows: Among them, loss represents the network loss value output after one round of iteration for the image input into the multi-scale feature enhancement aggregation network. ∑ represents the summation operation, N represents the total number of image samples in the training set, n represents the index of the image sample in the training set, and p n represents the probability that each pixel in the predicted input image belongs to the true target, and r n represents the category of each pixel in the input image. ∈ represents the correction factor, and the value of ∈ is any real number selected within the range of (0, 0.1).

Citation Information

Patent Citations

  • Multi-scale forward feature gain infrared weak and small target detection method for complex environment

    CN115661443A

  • Infrared small target detection method based on background suppression and feature fusion

    CN115830502A

  • Infrared weak and small target detection method based on Swinin-Transform and multi-scale feature fusion

    CN116188944A

  • Multiband fusion detection method based on unmanned platform

    CN106096604A

  • Remote sensing image target detection method based on multi-scale feature fusion

    CN110110599A