Detection Method and Device for Multispectral Occluded Targets, Storage Medium, Electronic Device

By fusing infrared images of different spectral segments and using reverse cascade Transformer and Transformer encoder with deformable attention module for feature extraction and encoding, combined with the detection head without anchor frame, the problem of low target detection accuracy in complex occlusion environments is solved, and higher robustness and accuracy are achieved.

CN119942095BActive Publication Date: 2025-06-20XIAN ORDNANCE IND TECH IND DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510425895.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-20
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the case of complex and diverse occlusions, existing object detection models are difficult to improve detection accuracy, especially when facing highly overlapping objects and strong interference environments.

Method used

The infrared images of different spectral segments in the target occlusion scene are obtained for fusion processing, and multi-scale feature extraction is performed using the backbone network of the reverse cascade Transformer. At the same time, based on the Transformer encoder with a deformable attention module, the feature codes are performed on the large target layer feature map and multi-scale fusion is performed, and the target detection head is finally used to detect the target.

Benefits of technology

The accuracy of object detection is improved, especially in strong interference and complex occlusion environments, the robustness and stability of the model are enhanced, while missing and missed detection are reduced, and redundant boxes are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942095B_ABST
    Figure CN119942095B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for detecting multi-spectral occluded targets, a storage medium, and an electronic device, which relate to the technical field of image processing. The main purpose is to solve the problem of improving the target detection accuracy in complex and diverse occlusion situations. The method includes obtaining infrared images in different spectral bands in a target occlusion scene, and performing fusion processing on the infrared images in different spectral bands to obtain a fused infrared image; performing multi-scale feature extraction processing on the fused infrared image based on the backbone network of the inverse cascaded Transformer to obtain multi-level image features; performing feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; performing multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fused feature image; and performing target detection processing on the preprocessed fused feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for detecting multi-spectral occluded targets, a storage medium, and an electronic device. Background Art

[0002] The task of object detection is to quickly and accurately locate the objects in an image, and object detection in an occlusion scenario is an important topic in the direction of object detection. When an object is occluded by other objects, inter-class occlusion occurs. Intra-class occlusion usually occurs in crowded scenes and seriously affects the detection performance.

[0003] The current occlusion object algorithms are divided into two types: traditional algorithms and deep learning algorithms. The traditional occlusion object detection algorithms mainly rely on prior information such as the saliency of the object and the consistency of the background to achieve detection. Most of these methods rely on prior information and have poor robustness and versatility in the face of complex occlusions in a strong interference environment. The object detection method based on deep learning uses training to obtain the features of each layer of image data and then learns information to output the result through a classifier. The deep learning method has better robustness and accuracy compared with the traditional algorithm and can also effectively handle multi-object and scale changes. Advanced typical models such as SSD, YOLOv8, sparse-R-CNN, etc. have good performance in the field of object detection. However, when facing occluded objects, highly overlapping objects have similar features, which makes it difficult for the detector to extract effective features. Secondly, because the objects overlap severely, the NMS algorithm may wrongly suppress some predictions. Therefore, for complex and diverse occlusion situations, the current detection models are difficult to learn all occlusion types, and the accuracy of occluded object recognition is not high. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for detecting multi-spectral occluded targets, a storage medium, and an electronic device, mainly aiming to solve the problem of how to improve the accuracy of object detection in complex and diverse occlusion situations.

[0005] According to one aspect of the present invention, a method for detecting multi-spectral occluded targets is provided, including:

[0006] Obtaining infrared images of different spectral bands in a target occlusion scenario, and performing fusion processing on the multiple infrared images of different spectral bands to obtain a fused infrared image;

[0007] Performing multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; the multi-level image features include small target layer feature maps, medium target layer feature maps, and large target layer feature maps;

[0008] Feature encoding processing is performed on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map;

[0009] The encoded feature map, the small target layer feature map, and the medium target layer feature map are subjected to multi-scale fusion to obtain a preprocessed fusion feature image;

[0010] Based on an anchor-free detection head, target detection processing is performed on the preprocessed fusion feature image to obtain the target detection result in the target occlusion scenario.

[0011] Further, the backbone network based on the reverse cascaded Transformer performs multi-scale feature extraction processing on the fused infrared image to obtain multi-level image features, including:

[0012] A multi-layer perceptron neural network is used to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1;

[0013] The backbone network is used to perform convolution and max pooling processing on the feature map y1 to obtain a feature map y2;

[0014] A reverse cascaded Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3;

[0015] A reverse cascaded Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4;

[0016] A reverse cascaded Transformer module is used to perform third-scale feature extraction processing on the medium target layer feature map y4 to obtain a large target feature map y5.

[0017] Further, the execution steps of the reverse cascaded Transformer module include:

[0018] A multi-layer perceptron neural network is used to perform dimensionality increase processing on the input image to obtain a dimensionality-increased input image;

[0019] A cascaded group attention mechanism is used to perform feature enhancement processing on the dimensionality-increased input image to obtain an enhanced feature image;

[0020] An inverted multi-layer perceptron neural network is used to perform dimensionality reduction processing on the enhanced feature image to obtain a target feature representation.

[0021] Further, the feature encoding processing of the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map includes:

[0022] Perform a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector;

[0023] Use a Transformer encoder to add position information to the large target feature vector, and input the large target feature vector with added position information into a deformable attention module;

[0024] The deformable attention module performs an offset process on the large target feature vector with added position information to obtain an offset feature vector corresponding to the large target feature vector;

[0025] Adjust the offset feature vector to a two-dimensional form to obtain the encoded feature map.

[0026] Furthermore, the deformable attention module includes multi-head cascaded group attention and a convolution operation;

[0027] The multi-head cascaded group attention inputs different segments of the large target feature vector with added position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascaded manner to optimize the feature representation, and calculates each attention feature map.

[0028] Furthermore, the multi-scale fusion of the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image includes:

[0029] Map the number of channels of the encoded feature map, the small target layer feature map, and the medium target layer feature map to the same number to obtain an encoded feature map, a small target layer feature map, and a medium target layer feature map respectively;

[0030] Use a convolution and reparameterization module to extract the features of the encoded feature map, the small target layer feature map, and the medium target layer feature map respectively, and perform a splicing process on the obtained features to obtain the preprocessed fusion feature image.

[0031] Furthermore, the object detection process of the preprocessed fusion feature image based on an anchor-free detection head to obtain the object detection result in the object occlusion scenario includes:

[0032] Use the query screening module in the detection head to perform object matching on the preprocessed fusion feature map to obtain multiple object prediction results; the query screening module is optimized based on the loss value for object prediction; the loss value is calculated based on the object prediction result, the real object, the adjacent real object, and the adjacent non-homogeneous prediction object;

[0033] The Top-K algorithm is used to filter multiple target prediction results to obtain filtered target prediction results;

[0034] The Transformer decoder in the detection head is used to decode the filtered target prediction results to obtain the target detection results in the target occlusion scenario.

[0035] According to another aspect of the present invention, a multi-spectral occluded target detection device is provided, including:

[0036] A first fusion module for acquiring infrared images of different spectral bands in a target occlusion scenario and performing fusion processing on the multiple infrared images of different spectral bands to obtain a fused infrared image;

[0037] A feature extraction module for performing multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map;

[0038] A feature encoding module for performing feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map;

[0039] A second fusion module for performing multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image;

[0040] A target detection module for performing target detection processing on the preprocessed fusion feature image based on an anchor-free detection head to obtain the target detection results in the target occlusion scenario.

[0041] Further, the feature extraction module is further configured to:

[0042] Use a multi-layer perceptron neural network to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1;

[0043] Use the backbone network to perform convolution and max pooling processing on the feature map y1 to obtain a feature map y2;

[0044] Use the reverse cascaded Transformer module to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3;

[0045] Use the reverse cascaded Transformer module to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4;

[0046] The third-scale feature extraction process is performed on the feature map y4 of the target layer using a reverse cascaded Transformer module to obtain a large target feature map y5.

[0047] Furthermore, the execution steps of the reverse cascaded Transformer module in the feature extraction module include:

[0048] The input image is dimensionally upsampled using a multi-layer perceptron neural network to obtain an upsampled input image;

[0049] The feature enhanced processing is performed on the upsampled input image using a cascaded group attention mechanism to obtain an enhanced feature image;

[0050] The enhanced feature image is dimensionally downsampled using an inverted multi-layer perceptron neural network to obtain a target feature representation.

[0051] Furthermore, the feature encoding module is also used for:

[0052] A convolution operation is performed on the large target layer feature map to obtain a corresponding large target feature vector;

[0053] The Transformer encoder is used to add position information to the large target feature vector and input the large target feature vector with added position information into the deformable attention module;

[0054] The deformable attention module performs an offset process on the large target feature vector with added position information to obtain an offset feature vector corresponding to the large target feature vector;

[0055] The offset feature vector is adjusted to a two-dimensional form to obtain the encoded feature map.

[0056] Furthermore, the deformable attention module in the feature encoding module includes multi-head cascaded group attention and convolution operations;

[0057] The multi-head cascaded group attention inputs different segments of the large target feature vector with added position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascaded manner to optimize the feature representation and calculate each attention feature map.

[0058] Furthermore, the second fusion module is also used for:

[0059] The channel numbers of the encoded feature map, the small target layer feature map, and the medium target layer feature map are mapped to the same number to obtain an encoded feature map, a small target layer feature map, and a medium target layer feature map respectively;

[0060] The convolution and reparameterization modules are used to extract the features of the encoded feature map, the small target layer feature map, and the medium target layer feature map respectively, and the obtained features are concatenated to obtain the preprocessed fusion feature image.

[0061] Further, the object detection module is further configured to:

[0062] Use the query and screening module in the detection head to perform object matching on the preprocessed fusion feature map to obtain multiple object prediction results; the query and screening module is obtained by optimizing the model based on the loss value and is used for object prediction; the loss value is calculated based on the object prediction results, the real objects, the adjacent real objects, and the adjacent non-homogeneous prediction objects;

[0063] Use the Top-K algorithm to filter multiple object prediction results to obtain the filtered object prediction results;

[0064] Use the Transformer decoder in the detection head to decode the filtered object prediction results to obtain the object detection results in the object occlusion scenario.

[0065] According to another aspect of the present invention, there is provided a storage medium storing at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above multi-spectral occluded object detection method.

[0066] According to another aspect of the present invention, there is provided an electronic device including a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0067] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above multi-spectral occluded object detection method.

[0068] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0069] The present invention provides a method and apparatus for detecting multi-spectral occluded targets, a storage medium, and an electronic device. Compared with the prior art, the present invention obtains infrared images in different spectral bands in a target occlusion scenario, and performs fusion processing on the multiple infrared images in different spectral bands to obtain a fused infrared image; performs multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; wherein, the backbone network has both local modeling of convolution and relies on the global modeling ability of the Transformer, which can effectively extract the features of occluded targets while reducing the resources used in model calculation. In the face of a strong interference occlusion environment, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; performs multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image; wherein, the Transformer encoder with a deformable attention module combines the attention calculations of local and global receptive fields, which helps the model learn strong features, and the obtained offset features enable the network to selectively focus on more important regions. In the face of occluded targets, the model will pay more attention to the visible parts of the targets, enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the reparameterization module is used to improve the inference speed, save memory occupancy, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the preprocessed fusion feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scenario. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted by this method causes fewer missed detections and false detections than the detection heads of existing algorithms when facing occluded targets. Only one prediction box is generated for each target in the detection head part, avoiding a large number of redundant boxes generated by the algorithm and improving the accuracy of the detection model when facing occluded targets.

[0070] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. Brief Description of the Drawings

[0071] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0072] Figure 1Shows a schematic flowchart of a method for detecting multi - spectral occluded targets provided by an embodiment of the present invention;

[0073] Figure 2 Shows a schematic flowchart of another method for detecting multi - spectral occluded targets provided by an embodiment of the present invention;

[0074] Figure 3 Shows a schematic flowchart of yet another method for detecting multi - spectral occluded targets provided by an embodiment of the present invention;

[0075] Figure 4 Shows a schematic flowchart of still another method for detecting multi - spectral occluded targets provided by an embodiment of the present invention;

[0076] Figure 5 Shows a schematic structural diagram of a cascaded group attention mechanism provided by an embodiment of the present invention;

[0077] Figure 6 Shows a schematic structural diagram of a deformable attention module provided by an embodiment of the present invention;

[0078] Figure 7 Shows a schematic structural diagram of a device for detecting multi - spectral occluded targets provided by an embodiment of the present invention;

[0079] Figure 8 Shows a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0080] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0081] An embodiment of the present invention provides a method for detecting multi - spectral occluded targets. As Figure 1 shown, the method includes:

[0082] 101. Obtain infrared images of different spectral bands in a target occlusion scene, and perform fusion processing on the multiple infrared images of different spectral bands to obtain a fused infrared image;

[0083] In an embodiment of the present invention, the current execution end acquires infrared images of different spectral bands in a target occlusion scene. Among them, the infrared images of different spectral bands represent infrared images captured based on different wavelength ranges. Since the wavelength ranges of the infrared images are different, more and richer information in the target occlusion scene can be obtained from the infrared images of different spectral bands. After the current execution end obtains the infrared images of different spectral bands, it fuses multiple infrared images of different spectral bands to obtain a fused infrared image. For example, after obtaining three infrared images of different spectral bands with a size of 640×640, the three single-channel images are fused into a three-channel image, that is, the three-channel image is the fused infrared image, which is not specifically limited in the embodiment of the present invention.

[0084] 102. Perform multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map;

[0085] In an embodiment of the present invention, the current execution end performs multi-scale feature extraction processing on the fused infrared image obtained in step 101 based on the backbone network of the reverse cascaded Transformer. Among them, the backbone network of the reverse cascaded Transformer is obtained by reversely cascading Transformer modules on the basis of the backbone network. The backbone network of the reverse cascaded Transformer has both local modeling of convolution and relies on the global modeling ability of the Transformer. It can effectively extract the features of the occluded target while reducing the resources used in model calculation. In the face of a strong interference occlusion environment, it has strong robustness and stability. The multi-level image features obtained by feature extraction in this embodiment include a small target layer feature map, a medium target layer feature map, and a large target layer feature map. Among them, the image dimensions of the small target layer feature map, the medium target layer feature map, and the large target layer feature map are arranged in a decreasing order. For example, if the dimension of the original image is 32, the dimension of the small target layer feature map can be 1 / 8 of the dimension of the original image, the dimension of the medium target layer feature map can be 1 / 16 of the dimension of the original image, and the dimension of the large target layer feature map can be 1 / 32 of the dimension of the original image, etc., which are not specifically limited in the embodiment of the present invention.

[0086] 103. Perform feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map;

[0087] In an embodiment of the present invention, the current execution end performs feature encoding processing on the feature map of the large target layer in the multi-level image features based on a Transformer encoder with a deformable attention module. Among them, the deformable attention module uses deformable convolution. Compared with ordinary convolution, deformable convolution adds an offset for dilation, so that the sampling position becomes an irregular position, making the receptive field closer to the shape of the actual target. In addition, the deformable attention module can model the global relationship. This method combines the attention calculations of the local and global receptive fields, which helps the model learn strong features. And this module obtains the offset feature through the offset network, enabling the network to selectively focus on more important regions. When facing occluded targets, the model will pay more attention to the visible part of the target, enhancing the modeling ability of the model.

[0088] 104. Perform multi-scale fusion on the encoded feature map, the feature map of the small target layer, and the feature map of the medium target layer to obtain a preprocessed fused feature image;

[0089] 105. Perform target detection processing on the preprocessed fused feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scenario.

[0090] In an embodiment of the present invention, the current execution end performs target detection processing on the preprocessed fused feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scenario. Among them, the anchor-free detection head includes a query screening module and a Transformer decoder. The query screening module includes an assignment stage and a loss calculation stage. The assignment stage is used for target matching processing, and the loss calculation stage is used for calculating the loss value of the prediction result, which is convenient for optimizing the model.

[0091] It should be noted that in an embodiment of the present invention, the current execution end initializes the anchor box as the position query for the subsequent decoder. Therefore, the final output consists of the classification results predicted by the refined anchor box and the refined content features.

[0092] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to extract richer features of occluded targets, another method for detecting multi-spectral occluded targets is provided, as Figure 2 shown. The steps are to perform multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features, including:

[0093] 201. Use a multi-layer perceptron neural network to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1;

[0094] In the embodiment of the present invention, the current execution end uses a Multilayer Perceptron (MLP) neural network to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1. The specific formula is as follows:

[0095]

[0096] Wherein, MLP e () represents the MLP model, X is the fused infrared image input to the model, represents the output / input ratio, C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map, is the feature map after dimensionality increase of the feature map X, which is the feature map y1 obtained after dimensionality increase in the embodiment of the present invention.

[0097] 202. Use the backbone network to perform convolution and max pooling processing on the feature map y1 to obtain a feature map y2;

[0098] In the embodiment of the present invention, the current execution end uses the backbone network to perform convolution and max pooling processing on the feature map y1 obtained in step 201. Among them, the backbone network preferably adopts the basic architecture of Resnet-18. The feature map y1 passes through a 7×7 convolution of Resnet-18 and then passes through a max pooling layer to obtain the feature map y2. The embodiment of the present invention does not make specific limitations.

[0099] 203. Use the reverse cascaded Transformer module to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3;

[0100] In the embodiment of the present invention, the current execution end uses the reverse cascaded Transformer module to perform first-scale feature extraction processing on the feature map y2 obtained in step 202. Specifically, the feature map y2 passes through two layers of reverse cascaded Transformer modules to obtain a small target layer feature map y3 with a size of 1 / 8 of the original image dimension. The embodiment of the present invention does not make specific limitations.

[0101] 204. Use the reverse cascaded Transformer module to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4;

[0102] In the embodiment of the present invention, the current execution end uses the reverse cascaded Transformer module to perform second-scale feature extraction processing on the small target layer feature map y3 obtained in step 203. Specifically, the small target layer feature map y3 passes through one layer of reverse cascaded Transformer module to obtain a medium target layer feature map y4 with a size of 1 / 16 of the original image dimension. The embodiment of the present invention does not make specific limitations.

[0103] 205. Use the reverse cascaded Transformer module to perform third-scale feature extraction processing on the medium target layer feature map y4 to obtain the large target feature map y5.

[0104] In the embodiment of the present invention, the current execution end uses the reverse cascaded Transformer module to perform third-scale feature extraction processing on the medium target layer feature map y4 obtained in step 204. Specifically, the medium target layer feature map y4 passes through one layer of the reverse cascaded Transformer module to obtain the large target layer feature map y5 with the dimension size of 1 / 32 of the original image dimension. The embodiment of the present invention does not make specific limitations.

[0105] It should be noted that, in the above embodiment, the execution steps of the reverse cascaded Transformer module include:

[0106] (1) Use a multi-layer perceptron neural network to perform dimensionality increase processing on the input image to obtain the input image after dimensionality increase; among them, the formula for the multi-layer perceptron (MLP) neural network to perform dimensionality increase processing on the input image is as follows:

[0107]

[0108] Among them, MLP e () represents the multi-layer perceptron model, X is the input image, represents the ratio of output / input, C is the number of channels of the feature map, H is the height of the feature map, W is the width of the feature map, is the feature map after the input image X is dimensionally increased. In step 203 of the embodiment of the present invention, the input image is the feature map y2; in step 204 of the embodiment of the present invention, the input image is the small target layer feature map y3; in step 205 of the embodiment of the present invention, the input image is the medium target layer feature map y4.

[0109] (2) Use the cascaded group attention mechanism to perform feature enhancement processing on the input image after dimensionality increase to obtain the enhanced feature image; among them, the cascaded group attention mechanism performs feature enhancement on the input image after dimensionality increase, and the specific formula is as follows:

[0110]

[0111] Among them, is the feature map after passing through operation. The operation is an efficient inverse difference structure, which has both the local modeling of convolution and relies on the global modeling ability of Transformer, as shown in the following formula:

[0112]

[0113] Among them, DW-Conv is depthwise separable convolution, Skip is skip connection operation, and CG-MHSA is cascaded multi-head group attention. It includes cascaded multi-head group attention and convolution operation. To save computational overhead, depthwise separable convolution is used to implement convolution, and multi-head attention uses cascaded group attention to expand the diversity of feature maps and optimize feature representation.

[0114] The cascaded group attention is as Figure 5 shown. The multi-head self-attention mechanism of Transformer can be used to capture the context information of features. In the occluded object detection task, it will pay more attention to the feature information of the visible part of the object and ignore the influence of the occluded part on detection. However, the attention heads of the multi-head self-attention mechanism are prone to redundancy, resulting in low computational efficiency. The cascaded group attention inputs each head with different partitions of the complete feature, thus explicitly decomposing the attention calculation into different heads. As shown in the following formula:

[0115]

[0116]

[0117] Among them, Attn is the attention operation, , , is the weight matrix, Concat is the concatenation operation, represents the projected matrix after concatenation, is the j-th partition of the input feature , is the feature after passing through the attention module, is composed of these partitions of the input feature , is the total number of attention heads. Splitting the features can make the input of attention heads more efficient and save computational resources, and can control the number of split features to improve the diversity of attention feature maps. Finally, the outputs of each attention are added to the subsequent attention by cascading to optimize feature representation and calculate each attention feature map.

[0118] (3) Use an inverted multi-layer perceptron neural network to perform dimensionality reduction on the enhanced feature image to obtain the target feature representation. Among them, the specific formula for the inverted multi-layer perceptron neural network to perform dimensionality reduction on the enhanced feature image is as follows:

[0119]

[0120] Among them, Indicates that the reverse input / output ratio is The multi-layer perceptron neural network of is the feature map The feature map after dimension contraction.

[0121] (4) Finally, the final output can also be obtained through the residual structure: , which is not specifically limited in the embodiments of the present invention.

[0122] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to better focus on the feature information of the target visible part and ignore the influence of the occluded part on the detection, another detection method for multi-spectral occluded targets is provided, such as Figure 3 shown, the steps are to perform feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map, including:

[0123] 301. Perform a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector;

[0124] In the embodiments of the present invention, the current execution end performs a convolution operation on the large target layer feature map y5 obtained in steps 201 to 205, so that the graphic information is reduced in dimension to vector information, and a corresponding large target feature vector is obtained. Then it is handed over to the Transformer encoder with a deformable attention module for processing.

[0125] 302. Use the Transformer encoder to add position information to the large target feature vector, and input the large target feature vector with added position information into the deformable attention module;

[0126] 303. The deformable attention module performs an offset processing on the large target feature vector with added position information to obtain an offset feature vector corresponding to the large target feature vector;

[0127] 304. Adjust the offset feature vector to a two-dimensional form to obtain the encoded feature map.

[0128] In the embodiments of the present invention, the current execution end uses a Transformer encoder to add positional information to the input large target feature vector, and then inputs the large target feature vector with positional information added into a deformable attention module. The output of the previous step is added to a residual structure for summation regularization, and the residual output is added through a feed-forward module. Among them, the deformable attention module uses deformable convolution. Compared with ordinary convolution, deformable convolution adds an offset for dilation, so that the sampling positions become irregular positions, making the receptive field closer to the shape of the actual target. The deformable attention module is as shown in Figure 6 shown, the input feature image , and its size is . For the input image , a reference grid is generated by downsampling with a scale factor . Among them, the reference point P represents the coordinate value, which is obtained by normalizing the coordinate value. The feature image X linearly projects the features to obtain , is obtained by an offset network with query as the input, and the obtained is added to the reference point to obtain the offset position information, as shown in the following formula:

[0129]

[0130] Among them, x represents the large target feature vector, Wq is the projection matrix, and q is the linear projection result; represents the offset network, is the bilinear interpolation operation, and are the weight matrices. The offset network uses the query feature to output the offset value of the reference point. The input feature first passes through a depth convolution to capture local features. Then, the GELU activation function and 1×1 convolution are used to obtain a two-dimensional offset. The bilinear interpolation method is used to sample the deformed reference point to obtain the sampling result . For the sampling result , multi-head self-attention calculation is performed, and relative position offset embedding is added at the same time. Finally, the output of the final attention module is obtained through the projection matrix. The deformable attention module can model global relationships. This method combines the attention calculations of local and global receptive fields, which helps the model learn strong features. And this module calculates the offset network through , and the obtained offset features can enable the network to selectively focus on more important regions. When facing occluded targets, the model will pay more attention to the visible parts of the targets, enhancing the modeling ability of the model.

[0131] It should be noted that in the above embodiment, the deformable attention module includes multi-head cascaded group attention and convolution operations; the multi-head cascaded group attention inputs different segments of the large target feature vector after adding position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascaded manner to optimize the feature representation and calculate each attention feature map.

[0132] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to better reflect the extracted feature information on a single feature map, another detection method for multi-spectral occluded targets is provided. The steps include multi-scale fusion of the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image, including:

[0133] Map the number of channels of the encoded feature map, the small target layer feature map, and the medium target layer feature map to the same number to obtain an encoded feature map, a small target layer feature map, and a medium target layer feature map respectively;

[0134] Use a convolution and reparameterization module to extract the features of the encoded feature map, the small target layer feature map, and the medium target layer feature map respectively, and perform splicing processing on the obtained features to obtain the preprocessed fusion feature image.

[0135] In the embodiment of the present invention, the current execution end performs multi-scale fusion on the encoded feature map F5 obtained in steps 301 to 304, the small target layer feature map y3 obtained in step 203, and the medium target layer feature map y4 obtained in step 204. Specifically, the multi-scale fusion process uses two paths, from top to bottom and from bottom to top, for feature fusion. Among them, features from different levels use a 1×1 convolutional layer to map the number of channels to the same number, and then perform feature fusion. The fusion module uses two paths. One path uses convolution to adjust the number of channels, and the other path uses a convolution and reparameterization module to extract features. Finally, the two paths are spliced and output. The reparameterization module uses two different network structure models during training and inference, which can improve the inference speed, save memory occupancy, and facilitate model deployment and acceleration.

[0136] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to avoid the algorithm from generating a large number of redundant boxes and improve the accuracy of the detection model for occluded targets, another detection method for multi-spectral occluded targets is provided, as Figure 4 shown. The steps include performing target detection processing on the preprocessed fusion feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scenario, including:

[0137] 401. Use the query and screening module in the detection head to perform target matching on the preprocessed fusion feature map to obtain multiple target prediction results. The query and screening module is obtained by optimizing the model based on the loss value and is used for target prediction. The loss value is calculated based on the target prediction results, true targets, adjacent true targets, and adjacent non-homogeneous prediction targets.

[0138] 402. Use the Top-K algorithm to filter multiple target prediction results to obtain filtered target prediction results.

[0139] 403. Use the Transformer decoder in the detection head to decode the filtered target prediction results to obtain the target detection results in the target occlusion scenario.

[0140] In the embodiment of the present invention, the current execution end inputs the preprocessed fusion feature map into the detection head part. The detection head includes a query and screening module query and a Transformer decoder. When the existing detection model using the NMS algorithm faces occluded targets, due to the possible severe overlap between prediction boxes during occlusion, the prediction boxes of different targets may be wrongly suppressed by the NMS algorithm as the prediction of one target, resulting in missed detections. And the threshold scores adopted during training and actual operation are not the same, so that the true level of the network is not fully utilized during network inference. In the embodiment of the present invention, the detection head part does not adopt the operations of the NMS algorithm and threshold screening, but only performs a TOPK operation on the final prediction results. The post-processing part is simpler and faster, generating only one prediction box for each target, avoiding the generation of a large number of redundant boxes by the algorithm.

[0141] Specifically, the preprocessed fusion feature map is input into the query and screening module query. The query and screening module includes an assignment stage and a loss calculation stage. The assignment stage performs target matching on the preprocessed fusion feature map to obtain multiple target prediction results. The loss calculation stage uses the Repulsion Loss loss function, and the specific formula is as follows:

[0142]

[0143] The above loss function is divided into three parts. The first part is the loss value (attraction term) generated by the prediction box and the true target box, denoted as ; the second part is the loss value (repulsion term (RepGT)) generated by the prediction box and the adjacent true target box, denoted as ; the third part is the loss value (repulsion Box (RepBox)) generated by the prediction box and the prediction box that does not predict the same true target as the adjacent one, denoted as . and is the correlation coefficient, which is used to balance the repulsion loss values of the two parts.

[0144] It should be noted that in the embodiment of the present invention, the current execution end initializes the anchor box as the position query for the subsequent decoder, and this parameter is learnable. During the training process, the loss function for occluding the target is used to regress and update the parameters. In the subsequent decoder, the denoising idea of DINO HEAD is adopted, and two hyperparameters are used to generate positive and negative samples. The noise scale of the positive sample is smaller than the smaller hyperparameter, and the noise scale of the negative sample is between the two hyperparameters. The positive sample is expected to predict the presence of a target, while the negative sample is expected to have no target. By generating difficult negative samples with a small difference from the positive samples, repeated predictions are avoided and confusion is suppressed. The final output consists of the refined anchor box and the classification result predicted by the refined content features. Finally, the TOPK operation is performed to sort the targets of this class to obtain the final object detection result.

[0145] An embodiment of the present invention provides a method for detecting multi-spectral occluded targets. Compared with the prior art, the present invention obtains infrared images of different spectral bands in a target occlusion scene, and performs fusion processing on multiple infrared images of different spectral bands to obtain a fused infrared image; performs multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; among them, the backbone network has both local modeling of convolution and relies on the global modeling ability of the Transformer, which can effectively extract the features of occluded targets while reducing the resources used in model calculation. In the face of a strongly interfering occlusion environment, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; performs multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fused feature image; among them, the Transformer encoder with a deformable attention module combines the attention calculations of local and global receptive fields, which helps the model learn strong features. The obtained offset features enable the network to selectively focus on more important regions. When facing occluded targets, the model will pay more attention to the visible parts of the targets, enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the reparameterization module is used to improve the inference speed, save memory occupancy, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the preprocessed fused feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scene. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted by this method causes fewer missed detections and false detections when facing occluded targets than the detection heads of existing algorithms. Only one prediction box is generated for each target in the detection head part, avoiding a large number of redundant boxes generated by the algorithm and improving the accuracy of the detection model when facing occluded targets.

[0146] As an implementation of the method described above Figure 1 shown, an embodiment of the present invention provides a multi-spectral occluded target detection device, as Figure 7 shown, the device includes:

[0147] A first fusion module 51, configured to obtain infrared images of different spectral bands in a target occlusion scene, and perform fusion processing on multiple infrared images of different spectral bands to obtain a fused infrared image;

[0148] A feature extraction module 52, configured to perform multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascaded Transformer to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map;

[0149] A feature encoding module 53, configured to perform feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map;

[0150] A second fusion module 54, configured to perform multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image;

[0151] A target detection module 55, configured to perform target detection processing on the preprocessed fusion feature image based on an anchor-free detection head to obtain a target detection result in the target occlusion scenario.

[0152] Further, the feature extraction module 52 is further configured to:

[0153] Use a multi-layer perceptron neural network to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1;

[0154] Use the backbone network to perform convolution and max pooling processing on the feature map y1 to obtain a feature map y2;

[0155] Use a reverse cascaded Transformer module to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3;

[0156] Use a reverse cascaded Transformer module to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4;

[0157] Use a reverse cascaded Transformer module to perform third-scale feature extraction processing on the medium target layer feature map y4 to obtain a large target feature map y5.

[0158] Further, the execution steps of the reverse cascaded Transformer module in the feature extraction module 52 include:

[0159] Use a multi-layer perceptron neural network to perform dimensionality increase processing on the input image to obtain a dimensionality-increased input image;

[0160] Use a cascaded group attention mechanism to perform feature enhancement processing on the dimensionality-increased input image to obtain an enhanced feature image;

[0161] Use an inverted multi-layer perceptron neural network to perform dimensionality reduction processing on the enhanced feature image to obtain a target feature representation.

[0162] Further, the feature encoding module 53 is further configured to:

[0163] Perform a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector;

[0164] Use a Transformer encoder to add position information to the large target feature vector and input the large target feature vector with added position information into a deformable attention module;

[0165] The deformable attention module performs an offset process on the large target feature vector with added position information to obtain an offset feature vector corresponding to the large target feature vector;

[0166] Adjust the offset feature vector into a two-dimensional form to obtain the encoded feature map.

[0167] Further, the deformable attention module in the feature encoding module 53 includes multi-head cascaded group attention and a convolution operation;

[0168] The multi-head cascaded group attention inputs different segments of the large target feature vector with added position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascaded manner to optimize the feature representation and calculate each attention feature map.

[0169] Further, the second fusion module 54 is also used for:

[0170] Map the channel numbers of the encoded feature map, the small target layer feature map, and the medium target layer feature map to the same number to obtain an encoded feature map, a small target layer feature map, and a medium target layer feature map respectively;

[0171] Use a convolution and reparameterization module to extract the features of the encoded feature map, the small target layer feature map, and the medium target layer feature map respectively, and perform a splicing process on the obtained features to obtain the preprocessed fusion feature image.

[0172] Further, the target detection module 55 is also used for:

[0173] Use the query screening module in the detection head to perform target matching on the preprocessed fusion feature map to obtain multiple target prediction results; the query screening module is optimized based on the loss value for target prediction; the loss value is calculated based on the target prediction results, the true target, the adjacent true target, and the adjacent non-homogeneous prediction targets;

[0174] Use the Top-K algorithm to filter multiple target prediction results to obtain the filtered target prediction results;

[0175] The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scenario.

[0176] An embodiment of the present invention provides a detection device for multi-spectral occluded targets. Compared with the prior art, the present invention obtains infrared images of different spectral bands in a target occlusion scenario, and performs fusion processing on the multiple infrared images of different spectral bands to obtain a fused infrared image; performs multi-scale feature extraction processing on the fused infrared image based on a backbone network of a reverse cascaded Transformer to obtain multi-level image features; among them, the backbone network has both local modeling of convolution and relies on the global modeling ability of the Transformer, which can effectively extract the features of occluded targets while reducing the resources used in model calculation. In the face of a strongly interfering occlusion environment, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; performs multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fusion feature image; among them, the Transformer encoder with a deformable attention module combines the attention calculations of local and global receptive fields, which helps the model learn strong features, and the obtained offset features enable the network to selectively focus on more important regions. In the face of occluded targets, the model will pay more attention to the visible part of the target, enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the reparameterization module is used to improve the inference speed, save memory occupancy, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the preprocessed fusion feature image based on an anchor-free detection head to obtain the target detection result in the target occlusion scenario. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted in this method causes fewer missed detections and false detections when facing occluded targets than the detection heads of existing algorithms. Only one prediction box is generated for each target in the detection head part, avoiding a large number of redundant boxes generated by the algorithm, and improving the accuracy of the detection model when facing occluded targets.

[0177] According to an embodiment of the present invention, a storage medium is provided, and the storage medium stores at least one executable instruction, and the computer executable instruction can execute the multi-spectral occluded target detection method in any of the above method embodiments.

[0178] Figure 8 The structural schematic diagram of an electronic device provided according to an embodiment of the present invention is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present invention.

[0179] As Figure 8As shown in the figure, the electronic device may include: a processor 602, a communications interface 604, a memory 606, and a communication bus 608.

[0180] Among them: The processor 602, the communications interface 604, and the memory 606 communicate with each other through the communication bus 608.

[0181] The communications interface 604 is used to communicate with network elements of other devices such as clients or other servers.

[0182] The processor 602 is used to execute the program 610, and specifically can execute the relevant steps of the above-mentioned multi-spectral occlusion target detection method.

[0183] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.

[0184] The processor 602 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0185] The memory 606 is used to store the program 610. The memory 606 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0186] The program 610 is specifically used to cause the processor 602 to perform the following operations:

[0187] Obtain infrared images of different spectral bands in the target occlusion scene, and perform fusion processing on the multiple infrared images of different spectral bands to obtain a fused infrared image;

[0188] Perform multi-scale feature extraction processing on the fused infrared image based on the backbone network of the inverse cascaded Transformer to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map;

[0189] Perform feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map;

[0190] Perform multi-scale fusion on the encoded feature map, the small target layer feature map, and the medium target layer feature map to obtain a preprocessed fused feature image;

[0191] Perform object detection processing on the preprocessed fused feature image based on an anchor-free detection head to obtain the object detection result in the target occlusion scenario.

[0192] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0193] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-spectral obstruction target detection method, characterized in that: include: Acquire infrared images of different spectral bands in a target occlusion scene, and fuse a plurality of the infrared images of different spectral bands to obtain a fused infrared image; Based on the backbone network of the reverse cascade Transformer, multi-scale feature extraction processing is performed on the fused infrared image to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map and a large target layer feature map; Performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; Performing multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; Performing target detection processing on the preprocessed fusion feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene; The backbone network based on the reverse cascade Transformer performs multi-scale feature extraction processing on the fused infrared image to obtain multi-level image features, including: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

2. The method according to claim 1, characterized in that The execution steps of the reverse cascade Transformer module include: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading; the input image is one of the feature map y2, the small target layer feature map y3 and the medium target layer feature map y4; Using a cascade group attention mechanism to perform feature enhancement processing on the dimensionally upgraded input image to obtain an enhanced feature image; The enhanced feature image is processed by using an inverted multi-layer perceptron neural network to reduce the dimension, thereby obtaining a target feature representation.

3. The method according to claim 1, characterized in that The Transformer encoder with a deformable attention module performs feature encoding processing on the large target layer feature map to obtain an encoded feature map, including: Performing a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; A Transformer encoder is used to add position information to the large target feature vector, and the large target feature vector after adding the position information is input into a deformable attention module; The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; The offset feature vector is adjusted to a two-dimensional form to obtain the encoding feature map.

4. The method according to claim 3, characterized in that The deformable attention module includes multi-head cascade group attention and convolution operations; The multi-head cascade group attention inputs different cuts of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

5. The method according to claim 1, characterized in that The multi-scale fusion of the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image includes: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

6. The method according to any one of claims 1 to 5, characterized in that The anchor-free detection head performs target detection processing on the preprocessed fusion feature image to obtain a target detection result in the target occlusion scene, including: The query and screening module in the detection head is used to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; Using the Top-K algorithm to filter the plurality of target prediction results to obtain filtered target prediction results; The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scene.

7. A multi-spectral obstruction target detection device, characterized in that: include: The first fusion module is used to obtain infrared images of different spectral bands in a target occlusion scene, and fuse multiple infrared images of different spectral bands to obtain a fused infrared image; A feature extraction module is used to perform multi-scale feature extraction processing on the fused infrared image based on a reverse cascade Transformer backbone network to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map; A feature encoding module, used for performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; A second fusion module is used to perform multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; A target detection module, used for performing target detection processing on the pre-processed fusion feature image based on a detection head without an anchor frame, to obtain a target detection result in the target occlusion scene; The feature extraction module is also used for: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

8. A storage medium, characterized in that: The storage medium stores at least one executable instruction, and the executable instruction executes an operation corresponding to the multi-spectral obstruction target detection method according to any one of claims 1-6.

9. An electronic device, characterized in that: It includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the multi-spectral obstruction target detection method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Remote sensing image target fine granularity identification method, system and device and storage medium

    CN115019182A