Infrared small target detection method for sparse perception global channel pruning

By introducing sparsely perceived global channel pruning technology in infrared small object detection, the existing problems in the comprehensive consideration of detection accuracy and inference speed are solved, and compact model deployment and real-time operation on resource-constrained devices are realized.

CN120014408APending Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510038159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing infrared small object detection method has not yet achieved satisfactory results under the comprehensive consideration of two major indicators: detection accuracy and inference speed, especially on resource-constrained edge devices, which are difficult to deploy and achieve real-time operation.

Method used

The sparse perception global channel pruning method is adopted, and the sparse constraint interpretable layer is introduced to quickly achieve sparseness of feature maps and learn robust features, thereby reducing the computing resources and storage resources requirements of the model.

Benefits of technology

It effectively maintains detection performance, while significantly reducing the number of parameters and calculations of the model, achieving compact model deployment and real-time operation on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014408A_ABST
    Figure CN120014408A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small target detection method based on sparse perception global channel pruning, and the method comprises the steps: rapidly achieving the sparsification of a feature map through introducing a sparse constraint interpretable layer, learning robust features, and effectively maintaining the detection performance of a model; according to the method, an infrared small target detection network pruning problem is formalized into a sparse representation problem, extra standards do not need to be designed to identify redundant channels, the sparsity of the feature map is directly utilized, the method is more in line with the characteristics of the target, and the detection performance and the reasoning efficiency can be effectively balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a sparse-perceived global channel pruning infrared small target detection method. Background Art

[0002] In most practical infrared imaging systems, the distance between the target to be detected and the detector is relatively far, which makes the infrared target occupy a very small area of ​​the entire infrared image, generally less than 100 pixels. In addition, the complex and changeable background makes detection difficult. The traditional model-driven method faces the following difficulties:

[0003] (1) There are few available features of the target. Due to the small size of the target, the total radiation energy is less than the radiation energy of the background, and the grayscale distribution in the image is variable, which makes it difficult to describe it with a unified mathematical model; and there is no fine texture, shape and other structural information, so the target detection method of traditional visible light images cannot be directly used for infrared small target detection. (2) The signal-to-noise ratio of the image is low. Due to the long imaging distance, small targets have similar characteristics to clutter and noise such as clouds and waves, and are easily submerged and interfered, resulting in small intensity of small targets and low signal-to-noise ratio of the image. The target signal is almost submerged in the unpredictable background, making it more difficult to detect. (3) The detection accuracy is not high. In actual situations and applications, the movement direction and speed of the target are highly maneuverable, which also makes improving the detection accuracy of maneuverable targets a problem that scholars are committed to solving. (4) Poor real-time performance. The detection effect is usually approximately inversely proportional to the amount of calculation. Algorithms with good detection effects often have large amounts of calculation. Because of the large amount of modeling and the inability of computing hardware conditions to keep up, the real-time performance is poor.

[0004] Compared with the traditional model-driven method, the deep learning-based method extracts the features of infrared small targets under the data-driven paradigm to solve the problem of single-frame infrared small target detection, showing superior performance. Although the existing single-frame infrared small target detection methods based on deep learning have achieved good performance, they still need to be further improved under the comprehensive consideration of the two major indicators of detection accuracy and reasoning speed. Most of the existing algorithms need to use densely nested interactive feature fusion and various forms of attention mechanisms to maintain or enhance the features of infrared small targets. These fragmented operations or very deep structures result in very large forward reasoning and video memory overhead. Although these large-size CNNs can be run on powerful GPU clusters, if they are run in real time (>30 frames / s) in real scenes, such as mobile / embedded devices, these platforms have few memory resources, low processor performance, and limited power consumption, which makes it difficult for the current highest-precision models to be deployed on these platforms and reach real-time running speed. Therefore, how to design a model with fewer parameters, less calculation, and faster speed while ensuring accuracy is a hot issue at present and a challenging direction.

[0005] The compression and acceleration of deep neural network models refers to the use of the redundancy of neural network parameters and network structure to streamline the model, without affecting the task completion, to obtain a model with fewer parameters and a more streamlined structure. Compared with the pursuit of accuracy in traditional networks, lightweight networks pay more attention to improving detection accuracy and model efficiency. The compressed model has smaller computing resource model and memory requirements, and can meet a wider range of application needs than the original model.

[0006] Network pruning is a network lightweight solution that is currently being studied more frequently. Its main idea is to remove redundant parameters, convolution kernels, channels, etc. from a trained deep network model, thereby reducing the model storage and accelerating the network test and inference time. Depending on whether the entire convolution kernel or channel is deleted at one time, that is, the pruning granularity, network pruning can be further divided into unstructured weight connection pruning and structured convolution kernel / channel pruning, such as Figure 1 shown.

[0007] Unstructured pruning has a finer granularity and can remove any "redundant" parameters of the desired proportion in the network without restriction. However, this will lead to irregular network structure after pruning and difficulty in effective acceleration. Srinivas et al. recently proposed a pruning scheme that does not rely on training data and backpropagation. It directly constructs and sorts the significance matrix of weights and then deletes redundant weight connections. Subsequently, Han et al. proposed a threshold-based pruning strategy. Once a weight connection is deleted, the connection is always blocked in subsequent training, which may cause the mis-pruning to not be effectively restored later. In order to solve this problem, Guo et al. proposed a dynamic network pruning scheme, including pruning and recovery, which can effectively avoid the problem that important weight connections pruned in the early stage of training cannot be effectively restored later. Although unstructured pruning methods can significantly remove non-important weight connections in the network, thereby obtaining a larger model compression ratio, this type of method ultimately obtains an irregular and sparse network structure. In actual operations, although some weight positions have been set to 0, the entire convolution kernel still needs to participate in matrix operations. Therefore, even with the support of the sparse matrix acceleration library, the actual acceleration effect is still very limited.

[0008] The granularity of structured pruning is relatively coarse, and the smallest unit of pruning is the combination of parameters within the filter. By setting evaluation factors for filters or feature maps, even the entire filter or certain channels can be deleted to make the network "narrower", so that effective acceleration can be directly obtained on existing software / hardware. In 2016, Lebedev et al. proposed adding structured sparse constraints to the loss function of deep networks. After gradient descent optimization, the network gradually tends to learn this structured sparsity and sets the convolution kernels less than a given threshold to 0. In the test reasoning stage, the convolution kernels with a value of 0 can be directly discarded. Wen et al. added regularization constraints on the convolution kernels, channels, convolution kernel shapes, and number of network layers of deep networks to the loss function. Based on this structured sparse learning process, an ideal acceleration ratio was obtained. Zhou et al. proposed a forward and backward term splitting algorithm to solve the optimization problem of objective functions with structured sparse constraints. In addition, Li et al. directly judged the importance of convolution kernels based on their norm values. Convolution kernels below a certain threshold were considered unimportant. At the same time, the output feature map corresponding to the convolution kernel was removed, and the number of input channels of the next layer of convolution kernels was correspondingly reduced. Finally, the performance degradation caused by pruning was restored by fine-tuning. Since the ReLU activation layer itself can obtain sparse output feature maps, Hu et al. calculated the non-zero ratio of the feature map corresponding to each convolution kernel based on this feature as a criterion for judging the importance of the convolution kernel. Luo et al. proposed a channel pruning scheme called ThiNet. In the convolution operation, there is a one-to-one correspondence between the convolution kernel of the current layer and the input channel of the convolution kernel of the next layer. Based on this feature, ThiNet explored the importance of the input channel of the convolution kernel of the next layer instead of directly considering the convolution kernel of the current layer, and established an effective channel selection optimization target. Liu et al. cleverly used the scaling factor in the batch normalization layer as a measure of channel importance, which can directly complete the channel pruning task in the basic network training process without introducing additional storage and computational overhead.

[0009] Structured pruning directly deletes the entire convolution kernel or channel in the convolution layer without introducing other additional data types for storage. It can be directly deployed on the existing deep learning framework, thereby compressing the network parameters while accelerating the actual operation speed of the entire network. However, there are still some shortcomings, especially the importance of judging each layer and channel, which greatly increases the training complexity and cost of the network pruning process, the process is cumbersome, and the adaptability is low. In addition, the pruning strategy needs to manually judge the sensitivity of each layer and determine the appropriate pruning ratio, which also requires a lot of analysis and layer-by-layer fine-tuning. The miniaturized network automatically pruned during training is also prone to the situation that the pruning ratio of high-computation convolutional layers is small, while the pruning ratio of low-computation convolutional layers is larger, which may seriously restrict the acceleration effect of the entire network.

[0010] CNN-based methods have shown promising performance in the field of infrared small target detection. However, due to the extremely small size of infrared small targets, they are easy to disappear in deep networks. Existing infrared small target detection methods usually need to design complex structures to maintain the characteristics of small targets, which makes the model bulky and difficult to deploy on resource-constrained edge devices. Channel pruning methods are hardware-friendly and easy to implement, and are widely popular in the field of model compression. Existing pruning algorithms focus on optimizing classification networks and pay more attention to semantic information related to image classification. In fact, the spatial detail information required for infrared small target detection can easily be mistakenly pruned in the shallow layer of the network, resulting in a sharp drop in detection performance. Summary of the invention

[0011] The technical problem to be solved by the present invention is to provide a sparse-aware global channel pruning infrared small target detection method in view of the shortcomings of the existing technology, introduce a sparse constrained interpretable layer to quickly realize the sparsification of feature maps, and learn robust features, thereby improving the detection performance and reducing the computing resources and storage resources required by the model.

[0012] In order to solve the above technical problems, the technical solution adopted by the present invention is: a sparse sensing global channel pruning infrared small target detection method, comprising the following steps:

[0013] Get infrared image x∈R Cin×H×W , maps the infrared image x to a high-level sparse representation z * ∈R Cout ×H×W ; Cin, H, W are the number of input channels, spatial height and width of the infrared image respectively; the infrared image is reconstructed using the sparse representation to obtain the features after sparse reconstruction;

[0014] The sparsely reconstructed features are used as inputs of a neural network, and the neural network is trained to obtain a target detection network.

[0015] Sparse representation z * It is expressed as: is the output of the convolutional sparse coding model, z represents the feature map, λ is the set parameter, and λ>0.

[0016] λ=0.1

[0017] The sparse representation z is obtained by iteratively performing the following steps: * The optimal solution is:

[0018]

[0019] in, z [0] = 0 for sparse representation z* The initial value of yes The adjoint operator of represents the shrinkage threshold operator that operates element-wise on the input variable, p [1] =z [0] , m1=1, t is the step length. The shrinkage threshold operator is expressed as: m∈R. t<1 / λ K ; K is the number of iterations, ε [K] represents the measured value of the Kth iteration, yes The adjoint matrix of , <> represents vector multiplication, represents the measurement value of the k+1th iteration, k=1, 2,..., K-1.

[0020] t=0.9 / λ K .

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] 1. The present invention quickly realizes the sparsification of feature maps by introducing a sparse constrained interpretable layer and learns robust features, which can effectively maintain the detection performance of the model;

[0023] 2. The present invention formalizes the infrared small target detection network pruning problem into a sparse representation problem. It does not need to design additional criteria to identify redundant channels. It directly utilizes the sparsity of the feature map itself, which is more in line with the characteristics of the target and can effectively balance the detection performance and reasoning efficiency.

[0024] 3. Since the infrared input image is represented at a high level in an unsupervised manner and incorporated into the network training through the above method, sparse learning of the redundancy between feature map channels can be induced to obtain a large number of features with zero response values. Finally, pruning them can obtain a compact infrared small target detection model, that is, the present invention directly removes channels with no response (0 response value) to reduce the operations of each convolution kernel with the channel, which can greatly reduce the computing and storage resources required for the model, and further promote the deployment and application of infrared small target detection networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Schematic diagram of pruning methods; (a) convolution kernel, (b) unstructured-single weight pruning, (c) structured-intra-kernel weight pruning, (d) structured-convolution kernel (channel) pruning;

[0026] Figure 2Comparison of different pruning methods; (a) traditional three-stage pruning paradigm, (b) pruning paradigm of the embodiment of the present invention;

[0027] Figure 3 Schematic diagram of sparse-aware global channel pruning according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] The embodiment of the present invention provides an infrared small target detection method, which reconstructs the target area by using sparse coding, performs sparse training on redundant feature maps in the network based on the sparse coding, and finally prunes the sparse network to obtain a compact infrared small target detection model. The pruning method of this embodiment does not need to design additional network structures or sophisticated standards to select redundant channels, and fully utilizes the characteristics of infrared small targets to learn more robust features and obtain a more compact model.

[0030] Based on the sparsity of infrared images, we introduce sparse modeling to reconstruct high-level features to achieve rapid sparsification of feature maps. where x i represents the infrared input image, y i is the corresponding true value, then the infrared small target detection problem based on deep learning can be expressed as:

[0031]

[0032] Where L represents the loss function, Θ(·,w) is a trainable deep neural network, and w represents the model parameters.

[0033] Given an infrared image x∈R Cin×H×W , where Cin and H, W are the input channels and spatial dimensions of the image, respectively. Our goal is to reconstruct the input x using the sparse features z. Specifically, we represent z as where ζ C ∈R H×W , as sparse as possible; Cout Indicates the number of output channels. In addition, the filter F can be expressed as:

[0034]

[0035] Among them, fij ∈R k×k represents a convolution kernel of size k×k. Then the input x can be generated by the operator defined as follows:

[0036]

[0037] The above formula defines the sparse modeling layer Map x to a high-level sparse representation z * ∈R Cout×H×W .f 1C represents the element in the first row and the Cth column of the matrix F, f CinC represents the element at the Cth row and Cth column. Inspired by the convolutional sparse coding model (CSC-layer), we find the optimal sparse solution z by solving the following Lasso type optimization problem * :

[0038]

[0039] In this embodiment, λ is set to 0.1. The feature map z represents the position and size of the convolution filter to be linearly combined in F. * It is a high-level feature representation based on sparsity constraints. norm to force z to be sparse, the residual value of norm to control the reconstruction error as small as possible. The goal of CSC-layer is to Reconstruct the input x, that is, find the smallest z such that x and The reconstruction error between them is as small as possible, and the parameters of z are as few as possible.

[0040] By solving the optimization problem above, the forward propagation of the sparse coding layer can be performed. Here, the fast iterative shrinkage threshold algorithm FISTA is used. The FISTA algorithm starts with an arbitrary initialization z (here z [0] =0), introduce a [1] =z [0] The auxiliary variable p, and a scalar m initialized to m1 = 1, for The following steps are performed iteratively:

[0041]

[0042] in, yes The FISTA iteration will automatically generate a nonlinear operator It represents a shrinkage threshold operator that operates element-wise on the input variable. For a scalar m∈R, we define the following shrinkage thresholding operation:

[0043]

[0044] The parameter t is the step size of FISTA. As long as the step size t is less than the operator The above formula converges to the inverse of the main eigenvalue. We estimate this eigenvalue by the power iteration method, which iteratively performs the following calculations to estimate the main eigenvector:

[0045]

[0046] in, When updating ε [k+1] -ε [k] element-wise norm is less than a predefined threshold (we use ) or reaches the maximum number of iterations (we use 50), the iteration terminates. Assuming K iterations are performed, the estimated main eigenvalue can be given by the following formula:

[0047]

[0048] Then, we can set t to a value less than 1 / λ K To ensure the convergence of the FISTA algorithm, we set t = 0.9 / λ. K . The dominant eigenvalue of changes after each update of the layer parameters F. Therefore, during training, λ is performed after each parameter update. K Once the training is completed, λ K Can be fixed during testing, so it does not add extra inference time.

[0049] The above iterative process will automatically generate a network architecture constructed by the unfolding optimization algorithm during forward propagation, which can be back-propagated through automatic differentiation. Therefore, the features sparsely reconstructed by the CSC-layer are incorporated into the network without introducing additional loss functions, and can directly participate in training optimization until convergence.

[0050] Since the infrared input image is represented at a high level in an unsupervised manner and incorporated into the network training through the above method, sparse learning of the redundancy between feature map channels can be induced to obtain a large number of features with zero response values. Finally, pruning them can obtain a compact infrared small target detection model.

[0051] This embodiment formalizes the pruning problem as a sparse reconstruction problem, which can maximize the identification of the source of redundancy in the infrared small target detection network; in addition, based on the sparse reconstruction-guided sparse training method, the redundant feature maps between channels can be sparsely trained to achieve lossless global channel pruning.

[0052] The following experiments are conducted on two single-frame infrared small target detection datasets containing real targets.

[0053] NUAA-SIRST contains 427 infrared images from different scenes, including 480 small target instances. The diversity of these instances enables the dataset to more realistically simulate complex scenes in practical applications. For the convenience of training and testing, we adopt a training-test ratio of 4:1, that is, 341 images are used for training and 86 images are used for testing. IRSTD-1k data is larger in scale, consisting of 1,000 real images of different shapes, sizes and scenes. Here we use 80% of the images for training and 20% for testing.

[0054] All the following experiments are performed on Python 3.8. The computer used for the experiments has an Nvidia GeForce GTX3090 GPU.

[0055] Experimental setup:

[0056] First, the input images of different initial sizes were resized to a resolution of 256×256 and normalized for preprocessing; AdaGrad was used as the optimizer with a learning rate of 0.05 and the Xavier strategy for weight initialization. A total of 500 epochs were trained with a weight decay of 1×10 -4 , the batch size is set to 16.

[0057] Evaluation Methodology:

[0058] Intersection over Union (IoU): is a pixel-level evaluation metric used to compare the similarity between the prediction and the true value

[0059] It is defined as the ratio of the intersection area and the union area between the predicted result and the true value.

[0060] ●Detection rate (P d ): Measures the ratio of the number of correctly predicted targets to the total number of all targets.

[0061] False alarm rate (F a ): Measures the ratio of incorrectly predicted pixels to the total number of pixels in the image.

[0062] Experimental results:

[0063] Table 1 Comparison of detection performance and efficiency before and after pruning

[0064]

[0065] On the above two data sets, the pruning method proposed in the embodiment of the present invention can achieve lossless or even better results before and after pruning, while the number of parameters (Params) and the amount of calculation (FLOPs) are significantly reduced.

[0066] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0067] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A sparse sensing global channel pruning infrared small target detection method, characterized in that: The following steps are involved: Get infrared image x∈R Cin×H×W , maps the infrared image x to a high-level sparse representation z * ∈R Cout×H×W ; Cin, H, and W are the number of input channels, spatial height, and width of the infrared image, respectively; Reconstructing the infrared image using the sparse representation to obtain sparsely reconstructed features; The sparsely reconstructed features are used as inputs of a neural network, and the neural network is trained to obtain a target detection network.

2. The infrared small target detection method based on sparse perception global channel pruning according to claim 1 is characterized in that: Sparse representation z * It is expressed as: is the output of the convolutional sparse coding model, z represents the feature map, λ is the set parameter, and λ>0; the feature map is the output of the infrared image after the neural network convolution operation.

3. The infrared small target detection method according to claim 2, characterized in that: λ=0.

1.

4. The infrared small target detection method according to claim 2, characterized in that: The sparse representation z is obtained by iteratively performing the following steps: * The optimal solution is: Among them, l ≥ 1, z [0] = 0 for sparse representation z * The initial value of yes The adjoint operator of represents the shrinkage threshold operator that operates element-wise on the input variable, p [1] =z [0] , m1=1, t is the step size, and K is the number of iterations.

5. The infrared small target detection method based on sparse perception global channel pruning according to claim 4 is characterized in that: The shrinkage threshold operator is expressed as: μ∈R, Indicates approximately equal to.

6. The infrared small target detection method with sparse perception global channel pruning according to claim 4 is characterized in that: t<1 / λ K ; ε [K] represents the measured value of the Kth iteration, yes The adjoint matrix of , < > represents vector multiplication, represents the measurement value of the k+1th iteration, k=1, 2,..., K-1.

7. The infrared small target detection method with sparse perception global channel pruning according to claim 6 is characterized in that: t=0.9 / l K 。 8. The infrared small target detection method with sparse perception global channel pruning according to claim 1 is characterized in that: The neural network adopts a convolutional neural network.

Citation Information

Cited By

  • Structure grouping de-correlation pruning method and infrared small target lightweight detection method

    CN120806025A

  • Structure packet disassociation pruning method and infrared small target lightweight detection method

    CN120806025B