Infrared target detection method based on cyclic multiplexing convolution
By using a cyclic multiplexing convolution network and a bidirectional attention aggregation decoder method in infrared object detection, the problem of high computational complexity in the prior art is solved, and efficient infrared small object detection at lower model complexity is achieved.
Patent Information
- Application Number
- CN202510480241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing infrared small object detection network based on deep learning improves detection performance while being too complex in computing, making it difficult to balance detection performance and model complexity.
The infrared object detection method based on cyclic multiplexing convolution is adopted, and the target image to be detected is processed through multiple multiplexing convolution encoders and bidirectional attention aggregation decoders to reduce the computational complexity.
At lower model complexity, deep-coded features of small targets can be extracted, and multi-scale feature interaction fusion can be guided through bidirectional attention aggregation with low computational complexity to improve detection performance.
Smart Images

Figure CN119992077A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an infrared target detection method based on cyclic multiplexing convolution. Background Art
[0002] The performance of current infrared small target detection networks is constantly improving. However, existing deep learning-based methods rely on densely nested models, and high-precision detection is often accompanied by expensive computational costs, making it difficult for such deep learning methods to balance detection performance and model complexity. The existing U-Net networks that have been proposed mainly focus on modifying existing modules or developing new functional modules to improve performance, for example, through nested network implementation, which usually leads to increased model complexity.
[0003] Therefore, how to reduce the complexity of calculation during detection is a technical problem that technical personnel in this field urgently need to solve. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an infrared target detection method based on cyclic multiplexing convolution, which solves the technical problem of high computational complexity in the prior art.
[0005] In order to solve the above technical problems, the present invention provides an infrared target detection method based on cyclic multiplexing convolution, comprising:
[0006] Acquire a target image to be detected; wherein the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold;
[0007] The target image to be detected is processed by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexed convolutional network to obtain target detection information; wherein the convolution kernel of each multiplexed convolutional encoder is consistent, and the bidirectional attention aggregation decoder is used to fuse the shallow coding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow coding features; each time the fusion is performed, the deep features are deep coding features or deep decoding features;
[0008] The target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain a target detection result.
[0009] Optionally, the target image to be detected is processed using multiple multiplexed convolution encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexed convolution network to obtain target detection information, including:
[0010] When the multiplexed convolution encoder is at the first level, a feature map corresponding to the target image to be detected is obtained, and the feature map is processed by the multiplexed convolution encoder to obtain a coding feature; wherein the normalization layers of the convolution layers in the multiplexed convolution encoder are different;
[0011] When the multiplexed convolutional encoder is not at the first level, the encoding features processed by the multiplexed convolutional encoder at the previous level are obtained, and the encoding features are processed using the multiplexed convolutional encoder at the current level.
[0012] Optionally, the method of processing the target image to be detected by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexed convolutional network to obtain target detection information includes:
[0013] The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain target fused deep features; wherein the resolution and the number of channels of the target fused deep features are consistent with the resolution and the number of channels of the adjacent shallow coding features;
[0014] The target fusion deep features and the adjacent shallow coding features are fused to obtain multi-scale decoding features, and the multi-scale decoding features are used as input of the previous bidirectional attention aggregation decoder.
[0015] Optionally, the target fusion deep feature and the adjacent shallow encoding feature are fused to obtain a multi-scale decoding feature, and the multi-scale decoding feature is used as the input of the previous bidirectional attention aggregation decoder, including:
[0016] Using efficient channel attention to strengthen the important channels in the target fusion deep features, to obtain the target strengthened fusion deep features; wherein the important channels are channels in which the amount of target information in the channels is greater than the set minimum threshold of the amount of information;
[0017] Processing the shallow coding features adjacent to the deep features using point-by-point convolution to obtain low-level attention features;
[0018] The low-level attention features and the target enhanced fusion deep features are fused to obtain the multi-scale decoding features.
[0019] Optionally, the resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain the target fused deep features, including:
[0020] The bidirectional attention aggregation decoder is used to upsample the resolution of the deep features, and the number of channels of the deep features is convolved point by point to obtain the target fused deep features.
[0021] Optionally, the target detection information is detected using a residual block in the cyclic multiplexing convolutional network to obtain a target detection result, including:
[0022] Determine the target multi-scale decoding information corresponding to each cyclic multiplexing; wherein the cyclic multiplexing process includes using the target multi-scale decoding information output by the current cycle as the input of the multiplexing convolution encoder in the next cycle;
[0023] All the target multi-scale decoding information is spliced in the channel dimension to obtain the target detection information, and the target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result.
[0024] Optionally, acquiring the target image to be detected includes:
[0025] Acquire the infrared image to be detected.
[0026] It can be seen that the present invention obtains a target image to be detected; wherein, the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold; multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexing convolutional network are used to process the target image to be detected to obtain target detection information; wherein, the convolution kernel of each multiplexed convolutional encoder is consistent, and the bidirectional attention aggregation decoder is used to fuse the shallow coding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow coding features; each time the fusion occurs, the deep features are deep coding features or deep decoding features; the target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result.
[0027] The beneficial effect of the present invention is that compared with the current nested network that needs to integrate many nodes, the encoder in the present invention is an encoder that reuses convolution kernels, does not introduce additional parameters, refines the deep coding features of the target under lower model complexity, and adopts low computational complexity bidirectional attention aggregation in the decoder to effectively guide the progressive interactive fusion of multi-scale features, thereby extracting deep coding features of small targets, reducing the amount of parameters, and reducing the complexity of calculation during detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0029] Figure 1 A flowchart of an infrared target detection method based on cyclic multiplexing convolution provided by an embodiment of the present invention;
[0030] Figure 2 A schematic diagram of a multiplexed convolutional encoder provided by an embodiment of the present invention;
[0031] Figure 3 A schematic diagram of a bidirectional attention aggregation decoder provided by an embodiment of the present invention;
[0032] Figure 4 An example flow chart of an infrared target detection method based on cyclic multiplexing convolution provided in an embodiment of the present invention;
[0033] Figure 5 A schematic diagram of a recurrent multiplexing convolutional network provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] Please refer to Figure 1 , Figure 1 A flowchart of an infrared target detection method based on cyclic multiplexing convolution is provided in an embodiment of the present invention. The method may include:
[0036] S101, obtaining a target image to be detected; wherein the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold.
[0037] The execution subject of this embodiment is an electronic device. The electronic device can be a computer, etc. This embodiment does not limit the specific target image to be detected. For example, the target image to be detected in this embodiment can be an infrared image to be detected, and the infrared image to be detected is a single-frame infrared weak target image; or the target image to be detected in this embodiment can be a weak target image in an aerial image. This embodiment does not limit the specific minimum pixel number threshold to 20, 15, etc. The feature of the target image to be detected in this embodiment can be a small number of pixels: the number of pixels occupied by the target in the image is very small, usually lower than the set minimum pixel number threshold. For example, if the threshold is set to 100 pixels, the number of pixels of the target may be only dozens. Small size: The size of the target is very small relative to the entire image, and may be only a few pixels wide and a few pixels high. For example, the size of the target may be between 2×2 and 10×10 pixels. Low contrast: The contrast between the target and the background is usually low, which makes the target not obvious in the image and easily submerged by background noise. Complex background: The background may contain a variety of complex textures and noises, and these background information may interfere with the detection of the target. For example, the background may contain trees, buildings, clouds, waves, etc. Weak signal: The signal strength of the target is weak, which may be caused by factors such as the physical characteristics of the target, imaging distance, and imaging conditions.
[0038] S102, using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the recurrent multiplexed convolutional network to process the target image to be detected and obtain target detection information; wherein the convolution kernel of each multiplexed convolutional encoder is consistent, and the bidirectional attention aggregation decoder is used to fuse the shallow coding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow coding features; each time the fusion occurs, the deep features are deep coding features or deep decoding features.
[0039] This embodiment does not limit the specific number of multiplexed convolutional encoders and bidirectional attention aggregation decoders. Since it is a cyclic multiplexed convolutional network, the number of bidirectional attention aggregation decoders is generally one less than the number of multiplexed convolutional encoders. In this embodiment, the convolution kernel of each multiplexed convolutional encoder is consistent, which means that the encoder includes multiple convolutional layers, and the convolution kernel used in each convolutional layer is the same, that is, the convolution kernel is reused. For ease of understanding, please refer to Figure 2 , Figure 2 A schematic diagram of a multiplexed convolutional encoder provided in an embodiment of the present invention, Conv represents convolution, BN represents batch normalization, ReLU represents activation function, and convolution operations are multiplexed inside each node of the encoder, such as Figure 2 As shown. In the nth cycle, the encoder convolution block and decoder convolution block of the i-th layer are expressed as and First, a reused convolution module is designed for the encoder. When n>0, the module reuses the previous convolution kernel and introduces a new normalization layer to stabilize the data distribution. This method is superior to the previous cascade reuse, mainly because after the features enter the loop structure, the reused convolution kernel can also refine the target deep semantic features without introducing additional parameters. Its calculation formula is: .in, represents a cascaded multiplexed convolutional module, represents the maximum pooling, Representation Node The output, Representation Node Output.
[0040] For easier understanding, please refer to Figure 3 , Figure 3 A schematic diagram of a bidirectional attention aggregation decoder provided for an embodiment of the present invention. Multi-scale features are fused by designing a bidirectional attention aggregation module for the encoder. Shallow features and deep features of adjacent layers are aggregated step by step in this module, which helps to recover multi-scale targets. The deep features are first doubled by bilinear interpolation, and the number of channels is reduced by half by point-by-point convolution, so that the deep features and shallow features maintain the same dimension. Then, efficient channel attention strengthens the target response of important channels in high-level features. Attention is generated using point-by-point convolution inner layer features, and the attention is weighted to deep features, using shallow detail information to emphasize deep target information and suppress noise. This module promotes the interactive aggregation of shallow features with sufficient detail information and deep features with rich semantic information, and guides the network to focus on multi-scale targets. Its calculation formula is: ;in, represents point-wise convolution, represents the activation function, represents efficient channel attention, represents element-wise multiplication, Indicates upsampling. Then, each loop multiplexing will obtain a set of decoded features. The target detection information in this embodiment refers to the information obtained by splicing the decoded features obtained by each loop multiplexing. It can be understood that in the first loop of the loop multiplexing process, the input of the first layer of multiplexing convolutional encoder is the feature corresponding to the target image to be detected; in the remaining loops, the input of the first layer of multiplexing convolutional encoder is the decoded feature obtained last time.
[0041] It should be further explained that, based on any of the above embodiments, the above-mentioned use of multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolutional network to process the target image to be detected and obtain target detection information may include: when the multiplexed convolutional encoder is the first level, obtaining the feature map corresponding to the target image to be detected, and using the multiplexed convolutional encoder to process the feature map to obtain the encoding feature; wherein, the normalization layer of the convolutional layer in the multiplexed convolutional encoder is different; when the multiplexed convolutional encoder is not the first level, obtaining the encoding feature obtained by the multiplexed convolutional encoder of the previous level, and using the multiplexed convolutional encoder of the current level to process the encoding feature. In this embodiment, the normalization layers of the convolutional layers in the multiplexed convolutional encoder are different, so that the distribution of data can be stabilized and the accuracy of information acquisition can be improved. The number of multiplexed convolutional encoders used is not limited in this embodiment. For example, 4 multiplexed convolutional encoders can be used in this embodiment; or, 6 multiplexed convolutional encoders can be used in this embodiment. In this embodiment, when the multiplexed convolution encoder is used to extract features from top to bottom (from shallow to deep), the input of the first layer is the features corresponding to the target image to be detected, and the subsequent input is the encoded features obtained by the multiplexed convolution encoder of the previous layer.
[0042] The method uses a plurality of multiplexed convolutional encoders and a plurality of bidirectional attention aggregation decoders in a cyclic multiplexed convolutional network to process the target image to be detected to obtain target detection information, including:
[0043] The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain target fused deep features; wherein the resolution and the number of channels of the target fused deep features are consistent with the resolution and the number of channels of the adjacent shallow coding features;
[0044] The target fusion deep features and the adjacent shallow coding features are fused to obtain multi-scale decoding features, and the multi-scale decoding features are used as input of the previous bidirectional attention aggregation decoder.
[0045] It should be further explained that, based on any of the above embodiments, the above-mentioned fusion of the target deep features and the adjacent shallow coding features to obtain multi-scale decoding features, and using the multi-scale decoding features as the input of the previous bidirectional attention aggregation decoder, may include:
[0046] S1021, using efficient channel attention to strengthen the important channels in the target fusion deep features, to obtain the target strengthened fusion deep features; wherein the important channel is a channel in which the amount of target information in the channel is greater than the set minimum threshold of the amount of information.
[0047] The Efficient Channel Attention (ECA) in this embodiment is a channel attention mechanism used to improve model performance in deep convolutional neural networks. The ECA module captures the dependencies between channels through one-dimensional convolution, avoiding the complex dimensionality reduction and dimensionality increase process, thereby achieving high efficiency and lightweight characteristics. The important channels in this embodiment refer to those channels that have made significant contributions to the current task (such as image classification, object detection, etc.).
[0048] S1022, using point-by-point convolution to process shallow coding features adjacent to deep features to obtain low-level attention features.
[0049] The deep features in this embodiment may be deep coding features and deep decoding features. It is understandable that when at the deepest layer, the input of the decoder only has the coding features of the deepest layer and the coding features of the second deepest layer (shallow coding features), and when not at the deepest layer, the input of the decoder is the deep decoding features obtained by the decoder of the next layer and the coding features of the same layer (shallow coding features).
[0050] S1023, fuse the low-level attention features and the target enhanced fusion deep features to obtain multi-scale decoding features.
[0051] This embodiment does not limit the specific method of fusion. For example, by learning weights, features of different scales are weighted and fused, and the contribution of different feature maps is dynamically adjusted. Or this embodiment can directly splice and fuse the low-level attention features and the target enhanced fusion deep features. This embodiment can improve the depth of information extraction and the accuracy of subsequent target detection by strengthening important channels.
[0052] It should be further explained that, based on any of the above embodiments, the above-mentioned use of the bidirectional attention aggregation decoder to expand the resolution of the deep features and reduce the number of channels of the deep features to obtain the target fused deep features may include: using the bidirectional attention aggregation decoder to upsample the resolution of the deep features and perform point-by-point convolution on the number of channels of the deep features to obtain the target fused deep features. This embodiment expands the resolution of the deep features by upsampling, and this embodiment reduces the number of channels of the deep features by point-by-point convolution.
[0053] S103, using the residual block in the cyclic multiplexing convolutional network to detect the target detection information to obtain the target detection result.
[0054] In this embodiment, in order to make full use of features of different granularities to generate robust predictions, the features from coarse granularity to fine granularity acquired by multiple cycles are spliced in the channel dimension to obtain target detection information, which is then input into the residual block to generate the target detection result. The calculation formula is: ;in, Represents the predicted output, N represents the total number of loop reuses, Represents a convolutional residual block operation.
[0055] It should be further explained that, in order to improve the accuracy of target detection, the target detection information is detected by using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result, which may include:
[0056] S1031, determining target multi-scale decoding information corresponding to each cyclic multiplexing; wherein the cyclic multiplexing process includes using the target multi-scale decoding information output by the current cycle as the input of the multiplexing convolution encoder in the next cycle;
[0057] S1032, all target multi-scale decoding information is spliced in the channel dimension to obtain target detection information, and the target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result.
[0058] In this embodiment, the first cyclic multi-scale decoding information is used as the input for the next cycle, thereby realizing the utilization of features of different granularities. The obtained embodiment continuously obtains target multi-scale decoding information through cyclic multiplexing, thereby obtaining target detection information, and fully utilizing features of different granularities to generate robust predictions.
[0059] An infrared target detection method based on cyclic multiplexing convolution provided by an embodiment of the present invention may include: S101, obtaining a target image to be detected; wherein the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold; S102, using multiple multiplexing convolution encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexing convolution network to process the target image to be detected to obtain target detection information; wherein the convolution kernels of the multiple multiplexing convolution encoders are consistent, and the bidirectional attention aggregation decoder is used to fuse shallow coding features and deep features of adjacent layers; the level of deep features is greater than the level of shallow coding features; each time the fusion is performed, the deep features are deep coding features or deep decoding features; S103, using the residual block in the cyclic multiplexing convolution network to detect the target detection information to obtain a target detection result. Compared with the current nested network that needs to integrate many nodes, the encoder in the present invention is an encoder that reuses convolution kernels, so no additional parameters are introduced, and the deep coding features of the target are effectively refined under lower model complexity. The decoder adopts low computational complexity bidirectional attention aggregation to effectively guide the progressive interactive fusion of multi-scale features, so as to extract the deep coding features of small targets and reduce the amount of parameters, thereby reducing the complexity of calculation during detection.
[0060] In order to make the present invention easier to understand, please refer to Figure 4 , Figure 4 An example flow chart of an infrared target detection method based on cyclic multiplexing convolution provided in an embodiment of the present invention may specifically include:
[0061] S201, input the infrared image to be detected into the first residual module of the recurrent multiplexing convolutional network for processing to obtain a feature map.
[0062] The infrared image input in this embodiment First, the residual block is used for preliminary processing to obtain the feature map ,Right now , Represents the field of real numbers.
[0063] For easier understanding, please refer to Figure 5 , Figure 5 A schematic diagram of a recurrent multiplexing convolutional network provided by an embodiment of the present invention, from Figure 5As can be seen in the figure, the recurrent multiplexing convolutional network includes residual blocks (residual modules), multiplexing convolutional encoders (multiplexing convolutional modules), bidirectional attention aggregation decoders (bidirectional attention aggregation modules) and feature pair stacking modules. From top to bottom, they are the first layer, the second layer, the third layer and the fourth layer. The one on the left is a multiplexing convolutional module, the three on the right are bidirectional attention aggregation modules, and the deepest layer is a multiplexing convolutional module. The network model includes an i-layer structure. Except for the fourth layer, which has only one multiplexing convolutional module, the other layer structures include a multiplexing convolutional module and a bidirectional attention aggregation module. The multiplexing convolutional module and the bidirectional attention aggregation module of the i-th layer are represented as and This network has only 4 layers. It should be noted that there can be many layers. 4 layers is just an example. The data flow in the deep learning network is basically the original infrared image obtained by the infrared device entering , Output the first feature map into , Output the second feature map, the second feature map enters , Output the third feature map, the third feature map enters , Output the fourth feature map, the fourth feature map enters , while the third feature map is The jump connection outputs the sixth feature map, and this is done in sequence until Output the target segmented image. The current nested and cascaded reuse method will increase a lot of parameters.
[0064] S202, input the feature map into the multiplexed convolutional encoder of the recurrent multiplexing convolutional network for layer-by-layer feature extraction to obtain the encoding features corresponding to each layer of the encoder; wherein the multiplexed convolutional encoder reuses the convolution kernel of the first convolutional layer, and the input of the encoders of other layers except the first layer is the encoding features output by the previous encoder.
[0065] The input multiplexed convolutional module is used to extract features layer by layer. Here we take one of the encoder nodes as an example. The encoder node is , the corresponding input features are , and then processed by the multiplexed convolution block to obtain the output features of this layer, and then downsampled to obtain the input features of the next layer ,Right now, .
[0066] S203, the adjacent deep features and shallow coding features are fed into the bidirectional attention aggregation decoder of the recurrent multiplexing convolutional network for step-by-step asymptotic interactive fusion to obtain multi-scale decoding information corresponding to the current cycle; wherein the bidirectional attention aggregation decoder is used to step-by-step fuse the shallow coding features and deep features of adjacent layers; the deep features of the deepest decoder are deep coding features, and the deep features of the non-deepest decoders are deep decoding features.
[0067] In this embodiment, after the encoder extracts the features, the high-level features and low-level features enter the decoder for step-by-step interactive fusion. Here, a decoder node is used as an example. The decoder node is , the corresponding input has high-level features and low-level features , . After multiple fusions, the decoded features are obtained , the decoded features are returned to the input backbone again, forming a loop structure, in which the reuse module deeply refines the features and decouples the target information.
[0068] S204, using the multi-scale decoding information obtained last time as the input of the multiplexing convolution encoder in the current cycle, to obtain multi-scale decoding information corresponding to multiple cycles.
[0069] S205, using a feature stacking module to concatenate all multi-scale decoding information to obtain target detection information.
[0070] In this embodiment, different fine-grained features extracted by multiple cycles are stacked together.
[0071] S206, using the second residual module to fuse the target detection information to obtain a fusion feature.
[0072] S207, using the third residual module to predict the fusion features to obtain a detection result.
[0073] In this embodiment, the robust prediction is generated by residual block fusion and point-by-point convolution. . Where Pred represents the predicted output, Represents the convolution residual block. The specific details of the cyclic multiplexing convolutional network are that the backbone network uses an 18-layer ResUNet network and the number of cyclic multiplexing is 2. The specific adjustments are as follows: In the encoder of the first layer, the number of multiplexing convolutional modules is 3, and the output feature dimension of this layer is In the encoder of the second layer, the number of multiplexed convolution modules is 2, and the output feature dimension of this layer is In the encoder of the third layer, the number of multiplexed convolution modules is 2, and the output feature dimension of this layer is In the encoder of the fourth layer, the number of multiplexed convolution modules is 2, and the output feature dimension of this layer is In each layer of the decoder, the number of reused interactive attention aggregation modules is 1, which receives features from this layer and a deeper layer, and the output feature dimension is consistent with the feature dimension of this layer.
[0074] When the embodiment of the present invention uses different times of cyclic multiplexing, the specific parameter amounts are as follows: when there is no cyclic multiplexing, the parameter amount is 0.184M; when cyclic multiplexing is used once, the parameter amount is 0.193M; when cyclic multiplexing is used twice, the parameter amount is 0.202M; when cyclic multiplexing is used three times, the parameter amount is 0.211M. It can be seen that each time the cyclic multiplexing is used, there is a parameter increment of 0.009M (no more than 5% of the original network parameter amount). This is caused by the introduction of a new batch normalization layer each time. On the other hand, reusing convolution kernels can indeed greatly reduce the amount of additional parameters introduced. The cyclic multiplexing convolutional network adjusted by the embodiment of the present invention was tested on the infrared image dataset (IRSTD-1k dataset), and it was found that its network performance was comparable to the most advanced methods, and the parameter amount was greatly reduced. Compared with the UIUNet network (a deep learning model for infrared small target detection), the cyclic multiplexing convolutional network has a 3% improvement in IoU (the degree of overlap between the predicted box and the real box), a 1% improvement in detection accuracy, and a 50% reduction in false alarm rate, maintaining good detection accuracy and false alarm suppression performance. In one embodiment, the migration cyclic multiplexing method is applied to the ResUNet network (a deep learning model that combines ResNet (residual network) and U-Net (full convolutional network)). Compared with the original ResUNet network, the number of parameters increases by 7%, the detection IoU increases by 1.8%, the detection rate increases by 2.48%, and the false alarm rate decreases by 11%. In one embodiment, the migration cyclic multiplexing method is applied to the DNANet network (a deep learning network for infrared small target detection (SIRST)). Compared with the original DNANet network, the number of parameters remains basically the same, the detection IoU increases by 1.59%, the detection rate remains basically unchanged, and the false alarm rate decreases by 72%.
[0075] It should also be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or device.
[0076] The above is a detailed introduction to the infrared target detection method based on cyclic multiplexing convolution provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. An infrared target detection method based on cyclic multiplexing convolution, characterized in that: include: Acquire a target image to be detected; wherein the target image to be detected is an image whose number of pixels occupying the entire image is lower than a set minimum pixel number threshold; The target image to be detected is processed by using multiple multiplexed convolutional encoders and multiple bidirectional attention aggregation decoders in a cyclic multiplexed convolutional network to obtain target detection information; wherein the convolution kernel of each multiplexed convolutional encoder is consistent, and the bidirectional attention aggregation decoder is used to fuse the shallow coding features and deep features of adjacent layers; the level of the deep features is greater than the level of the shallow coding features; the deep features are deep coding features or deep decoding features; The target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain a target detection result.
2. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: The target image to be detected is processed by using multiple multiplexed convolution encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolution network to obtain target detection information, including: When the multiplexed convolution encoder is at the first level, a feature map corresponding to the target image to be detected is obtained, and the feature map is processed by the multiplexed convolution encoder to obtain a coding feature; wherein the normalization layers of the convolution layers in the multiplexed convolution encoder are different; When the multiplexed convolutional encoder is not at the first level, the encoding features processed by the multiplexed convolutional encoder at the previous level are obtained, and the encoding features are processed using the multiplexed convolutional encoder at the current level.
3. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: The target image to be detected is processed by using multiple multiplexed convolution encoders and multiple bidirectional attention aggregation decoders in the cyclic multiplexed convolution network to obtain target detection information, including: The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain target fused deep features; wherein the resolution and the number of channels of the target fused deep features are consistent with the resolution and the number of channels of the adjacent shallow coding features; The target fusion deep features and the adjacent shallow coding features are fused to obtain multi-scale decoding features, and the multi-scale decoding features are used as input of the previous bidirectional attention aggregation decoder.
4. The infrared target detection method based on cyclic multiplexing convolution according to claim 3 is characterized in that: The target fusion deep layer feature and the adjacent shallow layer encoding feature are fused to obtain a multi-scale decoding feature, and the multi-scale decoding feature is used as the input of the previous bidirectional attention aggregation decoder, including: Using efficient channel attention to strengthen the important channels in the target fusion deep features, to obtain the target strengthened fusion deep features; wherein the important channels are channels in which the amount of target information in the channels is greater than the set minimum threshold of the amount of information; Processing the shallow coding features adjacent to the deep features using point-by-point convolution to obtain low-level attention features; The low-level attention features and the target enhanced fusion deep features are fused to obtain the multi-scale decoding features.
5. The infrared target detection method based on cyclic multiplexing convolution according to claim 3 is characterized in that: The resolution of the deep features is enlarged by using the bidirectional attention aggregation decoder, and the number of channels of the deep features is reduced to obtain the target fused deep features, including: The bidirectional attention aggregation decoder is used to upsample the resolution of the deep features, and the number of channels of the deep features is convolved point by point to obtain the target fused deep features.
6. The infrared target detection method based on cyclic multiplexing convolution according to any one of claims 1 to 5, characterized in that: The target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain a target detection result, including: Determine the target multi-scale decoding information corresponding to each cyclic multiplexing; wherein the cyclic multiplexing process includes using the target multi-scale decoding information output by the current cycle as the input of the multiplexing convolution encoder in the next cycle; All the target multi-scale decoding information is spliced in the channel dimension to obtain the target detection information, and the target detection information is detected using the residual block in the cyclic multiplexing convolutional network to obtain the target detection result.
7. The infrared target detection method based on cyclic multiplexing convolution according to claim 1 is characterized in that: Get the target image to be detected, including: Acquire the infrared image to be detected.
Citation Information
Patent Citations
Transmission opportunity sharing method in LTE-U network
CN105933982A
Quick magnetic resonance imaging method based on recursive residual U-type network
CN110151181A
Trained image processing for DWI and / or TSE with focus on body applications
EP3798662A1
Multichannel deep learning reconstruction of multiple repetitions
US20240036138A1