An Infrared Target Detection Method Based on Circular Reuse Convolution
Through the multiplexed convolution encoder and bidirectional attention aggregation decoder in the cyclic multiplexing convolution network, small infrared object detection is combined with residual blocks, the problem of high computational complexity is solved and efficient infrared object detection is achieved.
Patent Information
- Application Number
- CN202510480241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing infrared small object detection network has high computational complexity and is difficult to balance detection performance and model complexity.
The circular multiplexing convolution network is adopted, and multiple multiplexing convolution encoders and bidirectional attention aggregation decoders are used for feature fusion, combining residual blocks for detection, reducing the amount of parameters and reducing the computational complexity.
At lower model complexity, the target deep-coded features are refined, multi-scale features are extracted, detection calculation complexity is reduced, detection accuracy is improved, and false alarm rate is reduced.
Smart Images

Figure CN119992077B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to an infrared target detection method based on recycled convolution. Background Art
[0002] The performance of current infrared small target detection networks is constantly improving. However, existing deep learning-based methods rely on densely nested models, and high-precision detection often comes with expensive computational costs, making it difficult to balance detection performance and model complexity for such deep learning methods. The existing U-Net network that has been proposed mainly focuses on modifying existing modules or developing new functional modules to improve performance. For example, it is implemented through nested networks, usually resulting in an increase in model complexity.
[0003] Therefore, how to reduce the computational complexity during detection is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an infrared target detection method based on recycled convolution, which solves the technical problem of high computational complexity in the prior art.
[0005] To solve the above technical problem, the present invention provides an infrared target detection method based on recycled convolution, including:
[0006] Obtain a target image to be detected; wherein, the target image to be detected is an image whose number of pixels occupying the overall image is lower than a set minimum pixel number threshold;
[0007] Process the target image to be detected by using multiple recycled convolution encoders and multiple bidirectional attention aggregation decoders in the recycled convolution network to obtain target detection information; wherein, the convolution kernels of each recycled convolution encoder are the same, and the bidirectional attention aggregation decoder is used to fuse the shallow encoding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow encoding features; each time of fusion, the deep features are deep encoding features or deep decoding features;
[0008] Detect the target detection information by using a residual block in the recycled convolution network to obtain a target detection result.
[0009] Optionally, processing the target image to be detected by using multiple recycled convolution encoders and multiple bidirectional attention aggregation decoders in the recycled convolution network to obtain target detection information includes:
[0010] When the multiplexed convolutional encoder is at the first level, obtain the feature map corresponding to the target image to be detected, and use the multiplexed convolutional encoder to process the feature map to obtain encoded features; wherein, the normalization layers of the convolutional layers in the multiplexed convolutional encoder are different;
[0011] When the multiplexed convolutional encoder is not at the first level, obtain the encoded features processed by the multiplexed convolutional encoder of the previous level, and use the multiplexed convolutional encoder of the current level to process the encoded features.
[0012] Optionally, using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolutional network to process the target image to be detected to obtain target detection information, including:
[0013] Use the bidirectional attention aggregation decoder to enlarge the resolution of the deep features and reduce the number of channels of the deep features to obtain target fusion deep features; wherein, the resolution and number of channels of the target fusion deep features are consistent with the resolution and number of channels of the adjacent shallow encoded features;
[0014] Fuse the target fusion deep features and the adjacent shallow encoded features to obtain multi-scale decoded features, and use the multi-scale decoded features as the input of the previous bidirectional attention aggregation decoder.
[0015] Optionally, fusing the target fusion deep features and the adjacent shallow encoded features to obtain multi-scale decoded features, and using the multi-scale decoded features as the input of the previous bidirectional attention aggregation decoder, including:
[0016] Use efficient channel attention to strengthen the important channels in the target fusion deep features to obtain target enhanced fusion deep features; wherein, the important channels are the channels in which the number of target information is greater than the lowest threshold of the set information number;
[0017] Use pointwise convolution to process the shallow encoded features adjacent to the deep features to obtain low-level attention features;
[0018] Fuse the low-level attention features and the target enhanced fusion deep features to obtain the multi-scale decoded features.
[0019] Optionally, using the bidirectional attention aggregation decoder to enlarge the resolution of the deep features and reduce the number of channels of the deep features to obtain target fusion deep features, including:
[0020] Upsample the resolution of the deep features using the bidirectional attention aggregation decoder, and perform pointwise convolution processing on the number of channels of the deep features to obtain the target fused deep features.
[0021] Optionally, detect the target detection information using the residual blocks in the looped reuse convolutional network to obtain the target detection result, including:
[0022] Determine the target multi-scale decoding information corresponding to each looped reuse; wherein, the looped reuse process includes using the target multi-scale decoding information output in the current loop as the input of the reuse convolutional encoder in the next loop;
[0023] Concatenate all the target multi-scale decoding information in the channel dimension to obtain the target detection information, and detect the target detection information using the residual blocks in the looped reuse convolutional network to obtain the target detection result.
[0024] Optionally, the obtaining of the target image to be detected includes:
[0025] Obtain the infrared image to be detected.
[0026] It can be seen that the present invention obtains a target image to be detected; wherein, the target image to be detected is an image whose number of pixels occupying the overall image is lower than a set minimum pixel number threshold; uses a plurality of reuse convolutional encoders and a plurality of bidirectional attention aggregation decoders in the looped reuse convolutional network to process the target image to be detected to obtain target detection information; wherein, the convolutional kernels of each reuse convolutional encoder are the same, and the bidirectional attention aggregation decoder is used to fuse the shallow encoded features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow encoded features; each time of fusion, the deep features are deep encoded features or deep decoded features; uses the residual blocks in the looped reuse convolutional network to detect the target detection information to obtain the target detection result.
[0027] The beneficial effects of the present invention are as follows: Compared with the current situation where a nested network is required and many nodes need to be incorporated, the encoder in the present invention is an encoder with a reused convolutional kernel, which does not introduce additional parameters, refines the target deep encoded features under a lower model complexity, and adopts bidirectional attention aggregation with low computational complexity in the decoder, effectively guiding the progressive interaction and fusion of multi-scale features. Thus, it can not only extract the deep encoded features of small targets, but also reduce the number of parameters and the computational complexity during detection. Description of the Drawings
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0029] Figure 1 Flowchart of an infrared target detection method based on cyclic reuse convolution provided by an embodiment of the present invention;
[0030] Figure 2 Schematic diagram of a reuse convolution encoder provided by an embodiment of the present invention;
[0031] Figure 3 Schematic diagram of a bidirectional attention aggregation decoder provided by an embodiment of the present invention;
[0032] Figure 4 Flow example diagram of an infrared target detection method based on cyclic reuse convolution provided by an embodiment of the present invention;
[0033] Figure 5 Schematic diagram of a cyclic reuse convolution network provided by an embodiment of the present invention. Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0035] Please refer to Figure 1 , Figure 1 Flowchart of an infrared target detection method based on cyclic reuse convolution provided by an embodiment of the present invention. The method may include:
[0036] S101, obtain an image of a target to be detected; wherein, the image of the target to be detected is an image whose number of pixels occupying the overall image is lower than a set minimum pixel number threshold.
[0037] The execution subject of this embodiment is an electronic device. The electronic device can be a computer, a PC, etc. This embodiment does not limit the specific target image to be detected. For example, the target image to be detected in this embodiment can be an infrared image to be detected, and the infrared image to be detected is a single-frame infrared small and weak target image; or the target image to be detected in this embodiment can be a small and weak target image in an aerial image. This embodiment does not limit the specific minimum pixel number threshold, which can be 20, 15, etc. The characteristics of the target image to be detected in this embodiment can be as follows: few pixels: the number of pixels occupied by the target in the image is very small, usually lower than the set minimum pixel number threshold. For example, if the set threshold is 100 pixels, the number of pixels of the target may be only dozens. Small size: the size of the target is very small relative to the entire image, and it may be only a few pixels wide and a few pixels high. For example, the size of the target may be between 2×2 and 10×10 pixels. Low contrast: the contrast between the target and the background is usually low, which makes the target not obvious in the image and is easily submerged by background noise. Complex background: the background may contain various complex textures and noises, and this background information may interfere with the detection of the target. For example, the background may contain trees, buildings, clouds, waves, etc. Weak signal: the signal intensity of the target is weak, which may be caused by factors such as the physical characteristics of the target, the imaging distance, and the imaging conditions.
[0038] S102, use multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolutional network to process the target image to be detected, and obtain target detection information; wherein, the convolutional kernels of each multiplexed convolutional encoder are the same, and the bidirectional attention aggregation decoder is used to fuse the shallow encoded features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow encoded features; each time of fusion, the deep features are deep encoded features or deep decoded features.
[0039] This embodiment does not limit the specific number of multiplexed convolutional encoders and bidirectional attention aggregation decoders. Since it is a cyclic multiplexed convolutional network, the number of bidirectional attention aggregation decoders is generally one less than the number of multiplexed convolutional encoders. The fact that the convolutional kernels of each multiplexed convolutional encoder in this embodiment are the same means that the encoder includes multiple convolutional layers, and the convolutional kernels used in each convolutional layer are the same, that is, the reuse of convolutional kernels is realized. For easy understanding, please refer to Figure 2 , Figure 2 FIG. 10 is a schematic diagram of a multiplexed convolutional encoder provided by an embodiment of the present invention. Conv represents convolution, BN represents batch normalization, ReLU represents an activation function, and the reuse of convolution operations is performed inside each node of the encoder, as Figure 2 shown. In the nth cycle, the encoder convolutional block and decoder convolutional block of the i-th layer are denoted as and . First, a multiplexed convolution module was designed for the encoder. When n > 0, this module reuses the previous convolution kernels and introduces a new normalization layer to stabilize the data distribution. This method is superior to the previous cascaded multiplexing method. Mainly after the features enter the cyclic structure, the multiplexed convolution kernels can refine the deep semantic features of the target without introducing additional parameters. Its calculation formula is: . Among them, represents the cascaded multiplexed convolution module, represents the max pooling, represents the node 's output, represents the node 's output.
[0040] For ease of understanding, please refer to Figure 3 , Figure 3 which is a schematic diagram of a bidirectional attention aggregation decoder provided by an embodiment of the present invention. By designing a bidirectional attention aggregation module for the encoder to fuse multi-scale features. The shallow features and deep features of adjacent layers are gradually aggregated in this module, which helps to recover multi-scale targets. The deep features are first enlarged twice by bilinear interpolation, and the number of channels is reduced by half by pointwise convolution, so that the deep features have the same dimension as the shallow features. Then, the efficient channel attention strengthens the target response of the important channels in the high-level features. Use the pointwise convolution inner layer features to generate attention, weight this attention to the deep features, and use the shallow detail information to emphasize the deep target information and suppress noise. This module promotes the interactive aggregation of the shallow features with sufficient detail information and the deep features with rich semantic information, and guides the network to focus on multi-scale targets. Its calculation formula is: ; where represents the pointwise convolution, represents the activation function, represents the efficient channel attention, represents the element-wise multiplication, represents the upsampling. Then, a set of decoded features is obtained for each loop multiplexing. The object detection information in this embodiment refers to the information obtained by concatenating the decoded features obtained from each loop multiplexing. It can be understood that the input of the first-layer multiplexed convolution encoder in the first loop during the loop multiplexing process is the feature corresponding to the target image to be detected; the input of the first-layer multiplexed convolution encoder in the remaining loops is the decoded feature obtained in the previous time.
[0041] It should be further noted that, based on any of the above embodiments, the above-mentioned processing of the target image to be detected by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolutional network to obtain target detection information may include: when the multiplexed convolutional encoder is at the first level, obtaining the feature map corresponding to the target image to be detected, and using the multiplexed convolutional encoder to process the feature map to obtain encoded features; wherein, the normalization layers of the convolutional layers in the multiplexed convolutional encoder are different; when the multiplexed convolutional encoder is not at the first level, obtaining the encoded features processed by the multiplexed convolutional encoder of the previous level, and using the multiplexed convolutional encoder of the current level to process the encoded features. In this embodiment, the normalization layers of the convolutional layers in the multiplexed convolutional encoder are different, so as to stabilize the data distribution and improve the accuracy of information acquisition. The number of multiplexed convolutional encoders used in this embodiment is not limited. For example, 4 multiplexed convolutional encoders can be used in this embodiment; or, 6 multiplexed convolutional encoders can be used in this embodiment. When using the multiplexed convolutional encoder to extract features from top to bottom (from shallow to deep) in this embodiment, the input of the first layer is the feature corresponding to the target image to be detected, and the subsequent input is the encoded features obtained by the multiplexed convolutional encoder of the previous layer.
[0042] The processing of the target image to be detected by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolutional network to obtain target detection information includes:
[0043] Using the bidirectional attention aggregation decoder to expand the resolution of the deep features and reduce the number of channels of the deep features to obtain target fusion deep features; wherein, the resolution and the number of channels of the target fusion deep features are consistent with the resolution and the number of channels of the adjacent shallow encoded features;
[0044] Fusing the target fusion deep features and the adjacent shallow encoded features to obtain multi-scale decoded features, and using the multi-scale decoded features as the input of the bidirectional attention aggregation decoder of the previous layer.
[0045] It should be further noted that, based on any of the above embodiments, the above-mentioned fusing the target fusion deep features and the adjacent shallow encoded features to obtain multi-scale decoded features and using the multi-scale decoded features as the input of the bidirectional attention aggregation decoder of the previous layer may include:
[0046] S1021, using efficient channel attention to strengthen the important channels in the target fusion deep features to obtain target enhanced fusion deep features; wherein, the important channels are the channels in which the number of target information is greater than the lowest threshold of the set information number.
[0047] The Efficient Channel Attention (ECA) in this embodiment is a channel attention mechanism used to improve the performance of deep convolutional neural networks. The ECA module captures the dependencies between channels through one-dimensional convolution, avoiding complex dimensionality reduction and dimensionality increase processes, thereby achieving efficient and lightweight characteristics. The important channels in this embodiment refer to those channels that make significant contributions to the current task (such as image classification, object detection, etc.).
[0048] S1022, Use pointwise convolution to process the shallow encoded features adjacent to the deep features to obtain low-level attention features.
[0049] The deep features in this embodiment can be deep encoded features and deep decoded features. It can be understood that when at the deepest layer, the input of the decoder only has the deepest encoded features and the second deepest encoded features (shallow encoded features). When not at the deepest layer, the input of the decoder is the deep decoded features obtained by the decoder of the next layer and the encoded features of the same layer (shallow encoded features).
[0050] S1023, Fuse the low-level attention features and the target-enhanced fusion deep features to obtain multi-scale decoded features.
[0051] This embodiment does not limit the specific fusion method. For example, weighted fusion of features at different scales is performed by learning weights to dynamically adjust the contributions of different feature maps. Or this embodiment can directly perform concatenation fusion on the low-level attention features and the target-enhanced fusion deep features. By strengthening the important channels in this embodiment, the depth of information extraction can be improved, and the accuracy of subsequent object detection can be enhanced.
[0052] It should be further noted that based on any of the above embodiments, expanding the resolution of the deep features and reducing the number of channels of the deep features by using the bidirectional attention aggregation decoder to obtain the target fusion deep features may include: using the bidirectional attention aggregation decoder to perform upsampling on the resolution of the deep features and performing pointwise convolution on the number of channels of the deep features to obtain the target fusion deep features. This embodiment expands the resolution of the deep features by performing upsampling, and this embodiment reduces the number of channels of the deep features by performing pointwise convolution.
[0053] S103, Use the residual block in the recurrent reuse convolutional network to detect the object detection information to obtain the object detection result.
[0054] In this embodiment, in order to make full use of features with different granularities to generate robust predictions, the features obtained by recycling multiple times from coarse-grained to fine-grained are concatenated in the channel dimension to obtain target detection information, and then input into a residual block to generate a target detection result. Its calculation formula is as follows: ; where represents the prediction output, N represents the total number of recycling times, represents the convolutional residual block operation.
[0055] It should be further noted that in order to improve the accuracy of target detection, the above-mentioned use of the residual block in the recycling convolutional network to detect target detection information and obtain the target detection result may include:
[0056] S1031, determining the target multi-scale decoding information corresponding to each recycling; where the recycling process includes using the target multi-scale decoding information output in the current recycling as the input of the recycling convolutional encoder in the next recycling;
[0057] S1032, concatenating all the target multi-scale decoding information in the channel dimension to obtain target detection information, and using the residual block in the recycling convolutional network to detect the target detection information to obtain the target detection result.
[0058] In this embodiment, the multi-scale decoding information of the first recycling will be used as the input of the next time, so as to realize the utilization of features with different granularities. Through continuous recycling, the target multi-scale decoding information is obtained in this embodiment, and then the target detection information is obtained, making full use of features with different granularities to generate robust predictions.
[0059] An infrared target detection method based on cyclic reuse convolution provided by an embodiment of the present invention may include: S101, obtaining a target image to be detected; wherein, the target image to be detected is an image whose number of pixels occupying the overall image is lower than a set minimum pixel number threshold; S102, using a plurality of reuse convolution encoders and a plurality of bidirectional attention aggregation decoders in the cyclic reuse convolution network to process the target image to be detected to obtain target detection information; wherein, the convolution kernels of the plurality of reuse convolution encoders are the same, and the bidirectional attention aggregation decoder is used to fuse the shallow encoding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow encoding features; each time of fusion, the deep features are deep encoding features or deep decoding features; S103, using a residual block in the cyclic reuse convolution network to detect the target detection information to obtain a target detection result. Compared with the current situation where a nested network is required and many nodes need to be incorporated, the encoder in the present invention is an encoder with a reused convolution kernel, so no additional parameters are introduced, effectively refining the deep encoding features of the target at a lower model complexity. In the decoder, bidirectional attention aggregation with low computational complexity is adopted, effectively guiding the progressive interaction and fusion of multi-scale features, so that both the deep encoding features of small targets can be extracted and the number of parameters can be reduced, thereby reducing the computational complexity during detection.
[0060] For the convenience of understanding the present invention, please specifically refer to Figure 4 , Figure 4 FIG. is a flow example diagram of an infrared target detection method based on cyclic reuse convolution provided by an embodiment of the present invention, and specifically may include:
[0061] S201, inputting the infrared image to be detected into the first residual module of the cyclic reuse convolution network for processing to obtain a feature map.
[0062] In this embodiment, the input infrared image is first preliminarily processed by a residual block to obtain a feature map , that is , represents the real number field.
[0063] For the convenience of understanding, please refer to Figure 5 , Figure 5 FIG. is a schematic diagram of a cyclic reuse convolution network provided by an embodiment of the present invention. From Figure 5It can be seen that the cyclic reuse convolutional network includes a residual block (residual module), a reuse convolutional encoder (reuse convolutional module), a bidirectional attention aggregation decoder (bidirectional attention aggregation module), and a feature pair stacking module. From top to bottom are the first layer, the second layer, the third layer, and the fourth layer. The one on the left is the reuse convolutional module, and the three on the right are the bidirectional attention aggregation modules. The deepest layer is the reuse convolutional module. The network model includes i-layer structures. Except for the fourth layer which has only one reuse convolutional module, other layer structures include one reuse convolutional module and one bidirectional attention aggregation module. Denote the reuse convolutional module and the bidirectional attention aggregation module of the i-th layer as and . It should be noted that this network has only 4 layers, and there can be many layers in fact. 4 layers are just an example. In the data flow of the deep learning network, basically, the original infrared image obtained by the infrared device enters , and the first feature map is output and enters , then the second feature map is output. The second feature map enters , and the third feature map is output. The third feature map enters , and the fourth feature map is output. The fourth feature map enters . At the same time, the third feature map jumps and connects in to output the sixth feature map. This process is carried out in turn until the target segmentation image is output. However, the current reuse method of nesting and cascading will increase a lot of parameters.
[0064] S202: Input the feature map into the reuse convolutional encoder of the cyclic reuse convolutional network to perform layer-by-layer feature extraction, and obtain the encoded features corresponding to each layer of the encoder; among them, the reuse convolutional encoder repeatedly uses the convolutional kernel of the first convolutional layer. Except for the first layer encoder, the input of other layer encoders is the encoded features output by the previous encoder.
[0065] Input the backbone composed of the reuse convolutional module to perform layer-by-layer feature extraction. Here, take one encoder node as an example to illustrate. For the i-th layer in the n-th cycle, the encoder node is , and the corresponding input feature is . After being processed by the reuse convolutional block, the output feature of this layer is obtained, and then the downsampling process is performed to obtain the input feature of the next layer, that is, .
[0066] S203. Feed the adjacent deep features and shallow encoded features into the bidirectional attention aggregation decoder of the recurrent reuse convolutional network for progressive interactive fusion to obtain the multi-scale decoding information corresponding to the current cycle. Among them, the bidirectional attention aggregation decoder is used to progressively fuse the shallow encoded features and deep features of adjacent layers. The deep features of the decoder in the deepest layer are deep encoded features, and the deep features of the decoder in non-deepest layers are deep decoded features.
[0067] After the encoder extracts the features in this embodiment, the high-level features and low-level features enter the decoder for progressive interactive fusion. Here, a decoder node is taken as an example to illustrate. For the nth cycle, the decoder node of the ith layer is , and the corresponding inputs are the high-level feature and the low-level feature . . After multiple fusions, the decoded feature is obtained. The decoded feature is returned to the input backbone again to form a cyclic structure, where the reuse module refines the feature depth and decouples the target information.
[0068] S204. Use the multi-scale decoding information obtained last time as the input of the recurrent reuse convolutional encoder in the current cycle to obtain the multi-scale decoding information corresponding to multiple cycles.
[0069] S205. Use the feature stacking module to splice all the multi-scale decoding information to obtain the target detection information.
[0070] In this embodiment, the different fine-grained features extracted by multiple recurrent reuses are stacked together.
[0071] S206. Use the second residual module to fuse the target detection information to obtain the fused feature.
[0072] S207. Use the third residual module to predict the fused feature to obtain the detection result.
[0073] In this embodiment, fusion is performed by residual blocks, and pointwise convolution generates robust predictions. . Among them, Pred represents the prediction output, represents the convolutional residual block. The specific details of the recurrent reuse convolutional network are as follows: The backbone network uses an 18-layer ResUNet network, and the number of recurrent reuses is 2. The specific adjustments are as follows: In the encoder of the first layer, the number of reuse convolutional modules is 3, and the output feature dimension of this layer is . In the encoder of the second layer, the number of reuse convolutional modules is 2, and the output feature dimension of this layer is . In the encoder of the third layer, the number of reuse convolutional modules is 2, and the output feature dimension of this layer is 。In the encoder of the fourth layer, the number of multiplexed convolutional modules is 2, and the output feature dimension of this layer is 。In each decoder, the number of multiplexed interactive attention aggregation modules is 1, which receives the features of this layer and a deeper layer, and the output feature dimension is the same as the feature dimension of this layer.
[0074] In the embodiments of the present invention, when different numbers of loop multiplexing are used, the specific number of parameters is as follows: when there is no loop multiplexing, the number of parameters is 0.184M; when loop multiplexing is performed once, the number of parameters is 0.193M; when loop multiplexing is performed twice, the number of parameters is 0.202M; when loop multiplexing is performed three times, the number of parameters is 0.211M. It can be seen that each time loop multiplexing is performed, there is a parameter increment of 0.009M (not exceeding 5% of the original network parameters), which is brought about by introducing a new batch normalization layer each time. On the other hand, multiplexing convolutional kernels can indeed greatly reduce the introduction of additional parameters. The loop multiplexing convolutional network adjusted in the embodiments of the present invention is tested on the infrared image dataset (IRSTD-1k dataset), and it is found that its network performance is comparable to the state-of-the-art methods, and the number of parameters is greatly reduced. Compared with the UIUNet network (a deep learning model for infrared small target detection), the loop multiplexing convolutional network has a 3% improvement in IoU (the coincidence degree of the predicted box and the ground truth box), a 1% improvement in detection accuracy, and a 50% reduction in the false alarm rate, maintaining good detection accuracy and false alarm suppression performance. In one of the embodiments, the loop multiplexing method is migrated to the ResUNet network (a deep learning model that combines ResNet (residual network) and U-Net (fully convolutional network)). Compared with the original ResUNet network, the number of parameters increases by 7%, the detection IoU increases by 1.8%, the detection rate increases by 2.48%, and the false alarm rate decreases by 11%. In one of the embodiments, the loop multiplexing method is migrated to the DNANet network (a deep learning network for infrared small target detection (SIRST)). Compared with the original DNANet network, the number of parameters remains basically the same, the detection IoU increases by 1.59%, the detection rate remains basically unchanged, and the false alarm rate decreases by 72%.
[0075] It should also be noted that in this article, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0076] The above has introduced in detail an infrared target detection method based on cyclic reuse convolution. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An infrared target detection method based on cyclic multiplexing convolution, characterized in that: include: Acquire a target image to be detected; wherein the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold; The target image to be detected is processed by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexed convolutional network to obtain target detection information; wherein the backbone network of the cyclic multiplexed convolutional network is a ResUNet network, and the connection method is to reintroduce the final features output by the ResUNet network processing the input image, reintroduce the multiplexed convolutional encoder of the ResUNet network, and perform the extracted cyclic connection again, the convolution kernel of each multiplexed convolutional encoder is consistent, and the bidirectional attention aggregation decoder is used to fuse the shallow coding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow coding features; the deep features are deep coding features or deep decoding features; The target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain a target detection result.
2. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: The target image to be detected is processed by using multiple multiplexed convolution encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolution network to obtain target detection information, including: When the multiplexed convolution encoder is at the first level, a feature map corresponding to the target image to be detected is obtained, and the feature map is processed by the multiplexed convolution encoder to obtain a coding feature; wherein the normalization layer of the multiplexed convolution encoder is different each time it is cyclically multiplexed; When the multiplexed convolutional encoder is not at the first level, the encoding features processed by the multiplexed convolutional encoder at the previous level are obtained, and the encoding features are processed using the multiplexed convolutional encoder at the current level.
3. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: The target image to be detected is processed by using multiple multiplexed convolution encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolution network to obtain target detection information, including: The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain target fused deep features; wherein the resolution and the number of channels of the target fused deep features are consistent with the resolution and the number of channels of the adjacent shallow coding features; The target fusion deep features and the adjacent shallow coding features are fused to obtain multi-scale decoding features, and the multi-scale decoding features are used as input of the previous bidirectional attention aggregation decoder.
4. The infrared target detection method based on cyclic multiplexing convolution according to claim 3 is characterized in that: The target fusion deep layer feature and the adjacent shallow layer encoding feature are fused to obtain a multi-scale decoding feature, and the multi-scale decoding feature is used as the input of the previous bidirectional attention aggregation decoder, including: Using efficient channel attention to strengthen the important channels in the target fusion deep features, to obtain the target strengthened fusion deep features; wherein the important channels are channels in which the amount of target information in the channels is greater than the set minimum threshold of the amount of information; Processing the shallow coding features adjacent to the deep features using point-by-point convolution to obtain low-level attention features; The low-level attention features and the target enhanced fusion deep features are fused to obtain the multi-scale decoding features.
5. The infrared target detection method based on cyclic multiplexing convolution according to claim 3 is characterized in that: The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain the target fused deep features, including: The bidirectional attention aggregation decoder is used to upsample the resolution of the deep features, and the number of channels of the deep features is convolved point by point to obtain the target fused deep features.
6. The infrared target detection method based on cyclic multiplexing convolution according to any one of claims 1 to 5, characterized in that: The target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain a target detection result, including: Determine the target multi-scale decoding information corresponding to each cyclic multiplexing; wherein the cyclic multiplexing process includes using the target multi-scale decoding information output by the current cycle as the input of the multiplexing convolution encoder in the next cycle; All the target multi-scale decoding information is spliced in the channel dimension to obtain the target detection information, and the target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result.
7. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: Get the target image to be detected, including: Acquire the infrared image to be detected.
Citation Information
Patent Citations
Transmission opportunity sharing method in LTE-U network
CN105933982A
Quick magnetic resonance imaging method based on recursive residual U-type network
CN110151181A