Infrared weak and small target detection method, system and device and storage medium
By encoding and feature extraction of infrared images, and using multiple operations of global feature maps to generate binary maps, the problem of inaccurate detection of weak infrared objects in the prior art is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202510203068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
Existing infrared weak target detection technology is difficult to accurately capture the semantic information and background separation capabilities of small targets, resulting in inaccurate detection results.
Binary maps are generated to improve the accuracy of object detection by encoding infrared images, local feature extraction, feature representation and multiple flattening, transposing and upsampling operations of global feature maps.
The long-distance dependence ability of neural network models to capture features is improved, accurate detection and recognition of weak infrared targets is achieved, and the accuracy of detection results is improved.
Smart Images

Figure CN120219909A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared target detection, and particularly to an infrared dim and small target detection method, system, device and storage medium.
Background Art
[0002] The purpose of infrared dim and small target detection is to accurately detect and identify target objects with low thermal signals or small sizes in infrared images. Through this technology, potential targets such as humans, vehicles, and drones can be effectively discovered, even if they are not very conspicuous or relatively small in infrared images. The application fields of this technology include reconnaissance, security monitoring, unmanned driving, border patrol, etc., providing important support for improving safety and real-time response capabilities. Accurate infrared dim and small target detection helps enhance the ability to identify specific targets and helps people better understand and respond to potential risks or threats in complex environments.
[0003] In practical applications, a convolutional neural network model can be used to detect infrared dim and small targets. The convolutional neural network model can adopt some advanced models, which have been proven to have certain effects in target detection. Although the performance of the neural network model is constantly improving, the long-distance dependence ability of the convolutional kernel of the neural network model to capture features is insufficient, and it is unable to completely extract the semantic information of small targets. The ability of the neural network model to extract context is not good, resulting in a decline in the ability of the neural network model to separate targets from the background, and the detection results of infrared dim and small targets are inaccurate.
Summary of the Invention
[0004] In view of this, the present invention provides an infrared dim and small target detection method, system, device and storage medium.
[0005] The specific technical solution of the first embodiment of the present invention is: an infrared dim and small target detection method, the method includes: encoding a to-be-detected infrared image to obtain a first global feature map of the to-be-detected infrared image; performing local feature extraction on the first global feature map to obtain a first local feature map; performing feature representation on the first local feature map to obtain multiple feature variables of the first local feature map; the feature variables include a query variable, a key variable, and a value variable; performing flattening, transposing, and downsampling operations on the multiple feature variables to obtain a second global feature map of the to-be-detected infrared image; performing local feature extraction on the second global feature map to obtain a second local feature map; performing feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map; using a preset activation function to map the third global feature map to obtain a binary map of the to-be-detected infrared image; obtaining a target detection result of the to-be-detected infrared image in the binary map; the target detection result includes the presence or absence of a target object.
[0006] Preferably, after performing feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map, the method further includes: obtaining the number of times the upsampling operation has been completed; determining whether the number of times the upsampling operation has been completed reaches a preset number; if the number of times the upsampling operation has been completed does not reach the preset number, using the third global feature map as the second global feature map, and returning to the step of performing local feature extraction on the second global feature map to obtain a second local feature map, until the number of times the upsampling operation has been completed reaches the preset number, then performing the step of mapping the third global feature map using a preset activation function to obtain a binary map of the infrared image to be detected.
[0007] Preferably, the third global feature map is obtained by the following formula:
[0008] f RCBR (·) = DoubleDCBR(·) + CBR(·)
[0009] F out = f RCBR (f bilinear (F in ))
[0010] where DoubleDCBR is a double-layer depthwise separable convolution, BatchNorm, and Relu activation function, CBR is a single-layer convolution, BatchNorm, and Relu activation function, f bilinear is a bilinear interpolation function, F out is the third global feature map, and F in is the second local feature map.
[0011] Preferably, performing flattening, transposing, and downsampling operations on the multiple feature variables to obtain a second global feature map of the infrared image to be detected includes: performing low-dimensional encoding on the key variable in the variable to obtain an encoded key variable; performing low-dimensional encoding on the value variable in the variable to obtain an encoded value variable; performing flattening and transposing on the encoded key variable, the encoded value variable, and the query variable to obtain a similarity matrix; and performing channel information expansion on the similarity matrix to obtain the second global feature map.
[0012] Preferably, the similarity matrix is obtained by the following formula:
[0013]
[0014] where P is the similarity matrix, Q is the query variable, is the encoded key variable, For the encoded value variable, T is the transpose, and d is the image depth.
[0015] Preferably, after performing local feature extraction on the second global feature map to obtain a second local feature map, it further includes: expanding the channel information of the second local feature map to obtain an optimized second local feature map; then, performing feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map includes: performing feature representation, flattening, transposing, and upsampling operations on the optimized second local feature map to obtain a third global feature map.
[0016] Preferably, the query variable, key variable, and value variable are obtained using the following formula:
[0017] Q = W q ·f BN (F i )
[0018] K = W k ·f BN (F i )
[0019] V = W v ·f BN (F i )
[0020] where Q is the query variable, K is the key variable, V is the value variable, W is the width of the feature image, q is the weight matrix corresponding to Q, k is the weight matrix corresponding to K, v is the weight matrix corresponding to V, and f BN is batch normalization processing, and F i is the first local feature map.
[0021] The specific technical solution of the second embodiment of the present invention is as follows: An infrared small and weak target detection system, the system includes: an encoding module, a first feature extraction module, a feature representation module, a downsampling operation module, a second feature extraction module, an upsampling operation module, a mapping module, and a detection module; the encoding module is used to encode the infrared image to be detected to obtain a first global feature map of the infrared image to be detected; the first feature extraction module is used to perform local feature extraction on the first global feature map to obtain a first local feature map; the feature representation module is used to perform feature representation on the first local feature map to obtain multiple feature variables of the first local feature map; the feature variables include a query variable, a key variable, and a value variable; the downsampling operation module is used to flatten, transpose, and downsample the multiple feature variables to obtain a second global feature map of the infrared image to be detected; the second feature extraction module is used to perform local feature extraction on the second global feature map to obtain a second local feature map; the upsampling operation module is used to perform feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map; the mapping module is used to map the third global feature map by using a preset activation function to obtain a binary map of the infrared image to be detected; the detection module is used to obtain the target detection result of the infrared image to be detected in the binary map; the target detection result includes the presence or absence of a target object.
[0022] The specific technical solution of the third embodiment of the present invention is as follows: An infrared small and weak target detection device, including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0023] The specific technical solution of the fourth embodiment of the present invention is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0024] Implementing the embodiments of the present invention will have the following beneficial effects:
[0025] The present invention encodes the infrared image to be detected to obtain a first global feature map, extracts local features from the first global feature map to obtain a first local feature map, and represents the features of the first local feature map to obtain query variables, key variables, and value variables. The query variables, key variables, and value variables can effectively represent the context information of the feature image, which is beneficial to separating extremely weak targets from the background of the feature image. Therefore, the second global feature map obtained using the query variables, key variables, and value variables can enhance the ability of the neural network model to capture long-distance dependencies of features, achieving the purpose of accurately extracting the semantic information of small targets in the global feature map. Performing local feature extraction, feature representation, flattening, transposing, and upsampling operations on the second global feature map to obtain a third global feature map, and mapping on the third global feature map to obtain a binary map of the infrared image to be detected can make the binary map contain richer semantic information, thereby improving the accuracy of the detection result of infrared small and weak targets based on the binary map.
Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 It is a flowchart of the steps of the infrared small and weak target detection method;
[0028] Figure 2 It is a diagram of the infrared small and weak target detection model based on the linear self-attention mechanism and deep label supervision;
[0029] Figure 3 It is a schematic diagram of the linear self-attention mechanism;
[0030] Figure 4 It is a schematic diagram of the RCBR module;
[0031] Figure 5 It is a schematic diagram of the structure of the infrared small and weak target detection system;
[0032] Among them, 201, encoding module; 202, first feature extraction module; 203, feature representation module; 204, downsampling operation module; 205, second feature extraction module; 206, upsampling operation module; 207, mapping module; 208, detection module.
Detailed Embodiments
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0034] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products or devices.
[0035] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0036] Please refer to Figure 1 , which is a flowchart of the steps of a method for detecting a small and weak infrared target in the first embodiment of the present application, to improve the accuracy of obtaining the detection result of the small and weak infrared target. The method includes:
[0037] Step 101: Encode the infrared image to be detected to obtain a first global feature map of the infrared image to be detected;
[0038] Step 102: Extract local features from the first global feature map to obtain a first local feature map;
[0039] Step 103: Perform feature representation on the first local feature map to obtain multiple feature variables of the first local feature map; the feature variables include query variables, key variables, and value variables;
[0040] Step 104: Perform flattening, transposing, and downsampling operations on the multiple feature variables to obtain a second global feature map of the infrared image to be detected;
[0041] Step 105: Extract local features from the second global feature map to obtain a second local feature map;
[0042] Step 106: Perform feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map;
[0043] Step 107: Use a preset activation function to map the third global feature map to obtain a binary map of the infrared image to be detected;
[0044] Step 108: Obtain the target detection result of the infrared image to be detected in the binary map; the target detection result includes the presence or absence of a target object.
[0045] Specifically, please refer to Figure 2 , and use a neural network based on a linear self-attention mechanism and deep label supervision for infrared small target recognition. Figure 2 In, LSA is the global feature map, RCBR is the local feature map, max pooling is used for channel information expansion, and Conv3*3 is the convolution operation. For any input infrared image to be detected First, use a convolution with a kernel size of 3×3 to encode the infrared image to be detected twice to obtain a first global feature map The formula for obtaining the first global feature map is: F0 = DoubleConv3(I), where DoubleConv3 represents a double-layer convolution with a kernel size of 3. Through preliminary encoding, the number of channels of the infrared image to be detected can be widened, effectively enriching the channel information of the feature image. Extract local features from the first global feature map to obtain a first local feature map, and use a linear-based self-attention module (LSA) to perform feature representation on the first local feature map to obtain query variables, key variables, and value variables. Among them, the structure of the linear-based self-attention module is as Figure 3 shown. The self-attention mechanism actually wants the neural network to notice the correlation between different parts of the input feature image, thereby establishing long-distance dependence relationships between features. Flatten, transpose, and expand the channel information of the query variables, key variables, and value variables to obtain a second global feature map of the infrared image to be detected; perform local feature extraction in the second global feature map to obtain a second local feature map; perform feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map; use a preset activation function to map the third global feature map to obtain a binary map of the infrared image to be detected; obtain the target detection result of the infrared image to be detected in the binary map; the target detection result includes the presence or absence of a target object. Specifically, in the binary map, the value of the target and the region is 1, and the value of the background region is 0. Use the sigmoid function to process the binary map to obtain the target detection result. The specific formula is: target = sigmoid(F final), where target is the target detection result.
[0046] The method in this embodiment encodes the infrared image to be detected to obtain a first global feature map, extracts local features from the first global feature map to obtain a first local feature map, and represents the features of the first local feature map to obtain query variables, key variables, and value variables. The query variables, key variables, and value variables can effectively represent the context information of the feature image, which is beneficial to separating extremely weak targets from the background of the feature image. Therefore, the second global feature map obtained using the query variables, key variables, and value variables can enhance the ability of the neural network model to capture long-distance dependencies of features and achieve the purpose of accurately extracting the semantic information of small targets in the global feature map. The second global feature map is subjected to local feature extraction, feature representation, flattening, transposing, and upsampling operations to obtain a third global feature map. Mapping is performed on the third global feature map to obtain a binary map of the infrared image to be detected, which can make the binary map contain richer semantic information, thereby improving the accuracy of obtaining the detection result of the infrared small target based on the binary map.
[0047] In a specific embodiment, after performing the operations of feature representation, flattening, transposing, and upsampling on the second local feature map to obtain a third global feature map, it further includes: obtaining the number of times the upsampling operation has been completed; determining whether the number of times the upsampling operation has been completed reaches a preset number; if the number of times the upsampling operation has been completed does not reach the preset number, then using the third global feature map as the second global feature map and returning to the step of extracting local features from the second global feature map to obtain a second local feature map until the number of times the upsampling operation has been completed reaches the preset number, and then performing the step of mapping the third global feature map using a preset activation function to obtain the binary map of the infrared image to be detected.
[0048] Specifically, the preset number can be set according to the actual situation, such as 4 times. Feature representation, flattening, transposing, channel information expansion, and channel information expansion (upsampling) are performed on the second local feature map to obtain a third global feature map. At this time, the number of samplings is 1, which does not reach the preset number. Then, local features are extracted from the third global feature map to obtain a third local feature map. Representation, flattening, transposing, channel information expansion, and channel information expansion (upsampling) are performed on the third local feature map to obtain a fourth global feature map. At this time, the number of upsamplings is 2, which does not reach the preset number. The local feature extraction is repeated until the preset number is reached. The last output global feature map is the target global feature map.
[0049] Similarly, when flattening, transposing, and downsampling multiple feature variables to obtain a second global feature map, the number of downsampling operations can also be set. If the number of downsampling operations is not reached, the second global feature map is used as the first global feature map for repeated operations.
[0050] In a specific embodiment, the third global feature map is obtained using the following formula:
[0051] f RCBR (·) = DoubleDCBR(·) + CBR(·)
[0052] F out = f RCBR (f bilinear (F in ))
[0053] where DoubleDCBR is a double-layer depthwise separable convolution, BatchNorm, and Relu activation function, CBR is a single-layer convolution, BatchNorm, and Relu activation function, f bilinear is a bilinear interpolation function, F out is the third global feature map, and F in is the second local feature map.
[0054] In a specific embodiment, a Residual Conv BatchNorm ReLU (RCBR) module is used to extract local feature information. The structural schematic diagram of the RCBR module is as Figure 4 shown. Since Compression Trans can effectively aggregate the global features of an image, in this embodiment, an RCBR module based on convolution is further used to extract the target local features and perform channel expansion, strengthening the contrast information between the target and the background, which is beneficial to restoring the contour and position information of the target. The detailed process is as follows:
[0055] f RCBR (·) = DoubleDCBR(·) + CBR(·)
[0056] where DoubleDCBR represents a double-layer depthwise separable convolution + BatchNorm + Relu activation function; CBR represents a single-layer convolution + BatchNorm + Relu activation function. The size of the convolution kernel is set to 3.
[0057] In a specific embodiment, the operations of flattening, transposing, and downsampling the multiple feature variables to obtain the second global feature map of the infrared image to be detected include: performing low-dimensional encoding on the key variables in the variables to obtain the encoded key variables; performing low-dimensional encoding on the value variables in the variables to obtain the encoded value variables; flattening and transposing the encoded key variables, the encoded value variables, and the query variable to obtain a similarity matrix; and expanding the channel information of the similarity matrix to obtain the second global feature map.
[0058] Specifically, the query variable Q, the key variable K, and the value variable V are flattened and transposed into sequences of size n×d, where n = H×W. The output of self-attention is a scaled dot product, and the specific formula is: where, is called the context aggregation matrix or similarity matrix. Specifically, the i-th similarity matrix calculates the normalized pairwise dot product between each element in qi and ki, and then collects context information from the values. In this way, the self-attention mechanism essentially has a global receptive field and is good at capturing long-range dependencies. Since the self-attention module is equally used in each horizontal encoding, during the low-level encoding process, due to the large number of channels in the feature map, its computational cost to the model is huge. For example, for a 16×16 feature map, n = 256, and for a 256×256 resolution feature map, n = 65536, resulting in the sequence length dominating the self-attention calculation.
[0059] Since images are highly structured data, except for the boundary regions, most background pixels in the high-resolution feature map share similar features. Therefore, the pairwise attention calculation between all pixels is very inefficient. In addition, self-attention is essentially low-rank for long sequences, which means that most of the information is concentrated in the largest singular values. Inspired by this discovery, this embodiment uses two mappings to project the key and value: to low-dimensional encoding: where k = hw << n, and h and w are the sizes of the sampled feature map.
[0060] In a specific embodiment, the similarity matrix is obtained using the following formula:
[0061]
[0062] where P is the similarity matrix, Q is the query variable, is the encoded key variable, is the encoded value variable, T is the transpose, and d is the image depth. By performing low-dimensional encoding, the computational complexity is O(n 2d) It is reduced to O(nkd), enabling the self-attention mechanism to be applied to high-resolution feature maps. In the present invention, this self-attention mechanism is referred to as the Linear Self-Attention (LSA). It should be noted that the projection to the low-dimensional encoding can be any downsampling operation, such as average or max pooling or strided convolution. In the present invention, a 1×1 convolution is used, and then bilinear interpolation is used to downsample the feature map, reducing the size to 16, i.e., k = 16.
[0063] In a specific embodiment, after performing local feature extraction on the second global feature map to obtain a second local feature map, it further includes: expanding the channel information of the second local feature map to obtain an optimized second local feature map; then, the operations of performing feature representation, flattening, transposing, and upsampling on the second local feature map to obtain a third global feature map include: performing feature representation, flattening, transposing, and upsampling on the optimized second local feature map to obtain a third global feature map.
[0064] Specifically, max pooling is used to perform a downsampling operation on the extracted attention features, and a 1×1 convolution is used to further expand the channel information: F c = f 1×1 (maxpool(P)), where F c is the third global feature map, and P is the similarity matrix in the second local feature map.
[0065] In a specific embodiment, bilinear interpolation is used to upsample the extracted features. Bilinear interpolation is a commonly used image interpolation method for estimating pixel values at other positions based on known discrete sampling points. The principle of bilinear interpolation is based on the idea of linear interpolation but performs interpolation in two dimensions:
[0066] f bilinear (P) = f(Q 11 )w 11 + f(Q 21 )w 21 + f(Q 12 )w 12 + f(Q 22 )w 22
[0067] where P is the point to be interpolated, Q are the four positioning points, and w are the linear interpolation weights between the point to be interpolated and the four positioning points. Given the input feature map the feature map is upsampled through bilinear interpolation, and the features are compressed through the convolutional layer of the RCBR module to obtain the upsampled feature map (global feature map) Specifically, F out = f RCBR(f bilinear (F in ))。
[0068] In a specific embodiment, the query variable, key variable, and value variable are obtained using the following formulas:
[0069] Q = W q ·f BN (F i )
[0070] K = W k ·f BN (F i )
[0071] V = W v ·f BN (F i )
[0072] where Q is the query variable, K is the key variable, V is the value variable, W is the width of the feature image, q is the weight matrix corresponding to Q, k is the weight matrix corresponding to K, v is the weight matrix corresponding to V, and f BN is batch normalization processing, and F i is the first local feature map.
[0073] Specifically, through linear mapping, it is divided into three variables: query, key, and value
[0074] Q = W q ·f BN (F i )
[0075] K = W k ·f BN (F i )
[0076] V = W v ·f BN (F i )
[0077] where H and W respectively represent the height and width of F i , d is the dimension embedded in each of the multiple attention heads, and f BN represents Batch Normalization (BatchNorm). It should be noted that, different from most batch normalization operations, in this embodiment, batch normalization is performed on the original dimension instead of using Layer Normalization after linear mapping, which enables the network to more effectively aggregate the channel dimension information of the background.
[0078] In a specific embodiment, another defect in the detection of small and weak infrared targets is the problem of insufficient label constraints. Therefore, the present invention introduces the idea of deep label supervision. The feature maps output by each layer are convolved twice with sigmoid activation and threshold segmentation is performed to obtain the deep prediction Pred containing target information. i . At the same time, the label is downsampled to the same size as the feature maps of each layer to obtain the feature label Lable of each layer. i , and the deep supervision of the network is completed by calculating the similarity between Pred i and Lable i . The loss function selected by the present invention is: loss i = SoftIoU(Pred i , Lable i ). Since the network contains a total of four layers, the final loss function is defined as: Loss = loss0 + loss1 + loss2 + loss3 + loss4, where loss0 represents the output loss of the original size, and loss i (i = 1, 2, 3, 4) respectively represent the losses of each layer of the network from top to bottom.
[0079] Compared with traditional infrared small and weak target detection algorithms, the present method can perform nonlinear modeling of data through a multi-layer neural network. In the detection of small and weak infrared targets, the targets usually have complex texture and shape features, and traditional methods often cannot effectively capture these nonlinear relationships. This method can automatically learn feature representations through the training of large-scale data and provide better nonlinear modeling capabilities.
[0080] The linear self-attention mechanism is beneficial for the neural network to process high-resolution CNN feature maps in the image encoding and decoding stages; the CNN-self-attention mechanism hybrid encoding and decoding structure simultaneously has the advantages of local feature extraction of CNN and the global information interaction advantages of the self-attention mechanism, which is beneficial for extracting the semantic information of the target and combining the background context, so as to separate extremely weak targets from the background.
[0081] By providing supervision at multiple network levels, it is possible to strengthen feature learning, avoid overfitting, and improve the detection accuracy of small and weak targets. The multi-scale supervision promotes the propagation of gradients, accelerates the convergence of the model, and enhances the robustness of the model to background noise. In addition, it helps to improve the positioning accuracy of small targets, enabling the model to more accurately identify and locate targets when detecting micro, weak, dark, and small targets, and reducing the false alarms of the model.
[0082] In a specific embodiment, please refer to Figure 5, which is a schematic structural diagram of an infrared small and weak target detection system provided by the second embodiment of the present application. The system includes: an encoding module 201, a first feature extraction module 202, a feature representation module 203, a downsampling operation module 204, a second feature extraction module 205, an upsampling operation module 206, a mapping module 207, and a detection module 208; the encoding module 201 is used to encode the infrared image to be detected to obtain a first global feature map of the infrared image to be detected; the first feature extraction module 202 is used to perform local feature extraction on the first global feature map to obtain a first local feature map; the feature representation module 203 is used to perform feature representation on the first local feature map to obtain multiple feature variables of the first local feature map; the feature variables include a query variable, a key variable, and a value variable; the downsampling operation module 204 is used to perform flattening, transposing, and downsampling operations on the multiple feature variables to obtain a second global feature map of the infrared image to be detected; the second feature extraction module 205 is used to perform local feature extraction on the second global feature map to obtain a second local feature map; the upsampling operation module 206 is used to perform feature representation, flattening, transposing, and upsampling operations on the second local feature map to obtain a third global feature map; the mapping module 207 is used to map the third global feature map by using a preset activation function to obtain a binary map of the infrared image to be detected; the detection module 208 is used to obtain the target detection result of the infrared image to be detected in the binary map; the target detection result includes the presence or absence of a target object.
[0083] In a specific embodiment, the third embodiment of the present application provides an infrared small and weak target detection device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of the first embodiments of the present application.
[0084] In a specific embodiment, the fourth embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method according to any one of the first embodiments of the present application.
[0085] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
[0086] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for detecting infrared small targets, characterized in that: The method comprises: Encoding the infrared image to be detected to obtain a first global feature map of the infrared image to be detected; Performing local feature extraction on the first global feature map to obtain a first local feature map; Performing feature representation on the first local feature map to obtain a plurality of feature variables of the first local feature map; the feature variables include query variables, key variables and value variables; Flattening, transposing and downsampling the multiple feature variables to obtain a second global feature map of the infrared image to be detected; Performing local feature extraction on the second global feature map to obtain a second local feature map; Performing feature representation, flattening, transposition and upsampling operations on the second local feature map to obtain a third global feature map; Mapping the third global feature map using a preset activation function to obtain a binary image of the infrared image to be detected; The target detection result of the infrared image to be detected is obtained in the binary image; the target detection result includes the existence of the target object or the absence of the target object.
2. The infrared small target detection method according to claim 1, characterized in that: After performing feature representation, flattening, transposition and upsampling operations on the second local feature map to obtain a third global feature map, the method further includes: Get the number of completed upsampling operations; Determine whether the number of completed upsampling operations reaches a preset number; If the number of completed upsampling operations does not reach the preset number, the third global feature map is used as the second global feature map, and the step of extracting local features from the second global feature map to obtain the second local feature map is returned until the number of completed upsampling operations reaches the preset number, and then the step of mapping the third global feature map using a preset activation function to obtain the binary image of the infrared image to be detected is executed.
3. The infrared small target detection method according to claim 2, characterized in that: The third global feature map is obtained using the following formula: f RCBR (·)=DoubleDCBR(·)+CBR(·) F out =f RCBR (f bilinear (F in )) Among them, DoubleDCBR is a double-layer depth-separable convolution, BatchNorm, Relu activation function, CBR is a single-layer convolution, BatchNorm, Relu activation function, f bilinear is the bilinear interpolation function, F out is the third global feature map, F in is the second local feature map.
4. The infrared small target detection method according to claim 1, characterized in that: The flattening, transposing and downsampling operations are performed on the plurality of feature variables to obtain a second global feature map of the infrared image to be detected, comprising: Perform low-dimensional encoding on the key variables in the variables to obtain the encoded key variables; Perform low-dimensional encoding on the value variables in the variables to obtain the encoded value variables; Flattening and transposing the encoded key variable, the encoded value variable, and the query variable to obtain a similarity matrix; Perform channel information expansion on the similarity matrix to obtain the second global feature map.
5. The infrared small target detection method according to claim 4, characterized in that: The similarity matrix is obtained using the following formula: Wherein, P is the similarity matrix, Q is the query variable, is the encoded key variable, is the encoded value variable, T is the transpose, and d is the image depth.
6. The infrared small target detection method according to claim 1, characterized in that: After extracting local features from the second global feature map to obtain a second local feature map, the method further includes: Performing channel information expansion on the second local feature map to obtain an optimized second local feature map; Then, performing feature representation, flattening, transposition and upsampling operations on the second local feature map to obtain a third global feature map includes: Perform feature representation, flattening, transposition and upsampling operations on the optimized second local feature map to obtain a third global feature map.
7. The infrared small target detection method according to claim 1, characterized in that: The query variable, key variable and value variable are obtained using the following formula: Q=W q ·f BN (F i ) K=W k ·f BN (F i ) V=W v ·f BN (F i ) Where Q is the query variable, K is the key variable, V is the value variable, W is the width of the feature image, q is the weight matrix corresponding to Q, k is the weight matrix corresponding to K, v is the weight matrix corresponding to V, and f BN For batch normalization, F i is the first local feature map.
8. An infrared small target detection system, characterized in that: The system comprises: an encoding module, a first feature extraction module, a feature representation module, a downsampling operation module, a second feature extraction module, an upsampling operation module, a mapping module and a detection module; The encoding module is used to encode the infrared image to be detected to obtain a first global feature map of the infrared image to be detected; The first feature extraction module is used to perform local feature extraction on the first global feature map to obtain a first local feature map; The feature representation module is used to perform feature representation on the first local feature map to obtain a plurality of feature variables of the first local feature map; the feature variables include query variables, key variables and value variables; The downsampling operation module is used to perform flattening, transposition and downsampling operations on the multiple feature variables to obtain a second global feature map of the infrared image to be detected; The second feature extraction module is used to extract local features from the second global feature map to obtain a second local feature map; The upsampling operation module is used to perform feature representation, flattening, transposition and upsampling operations on the second local feature map to obtain a third global feature map; The mapping module is used to map the third global feature map using a preset activation function to obtain a binary image of the infrared image to be detected; The detection module is used to obtain the target detection result of the infrared image to be detected in the binary image; the target detection result includes the existence of the target object or the absence of the target object.
9. An infrared small target detection device, comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 7.