Multispectral shielding target detection method and device, storage medium and electronic equipment

By fusing infrared images of different spectral segments and combining technical means of reverse cascade Transformer and deformable attention module, the problem of low target detection accuracy in complex occlusion situations is solved, and more efficient feature extraction and detection accuracy is achieved.

CN119942095AActive Publication Date: 2025-05-06XIAN ORDNANCE IND TECH IND DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510425895.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art is difficult to improve the accuracy of object detection under complex and diverse occlusion conditions, especially in strong interference environments, with poor robustness and general applicability.

Method used

By obtaining infrared images of different spectral segments in the target occlusion scene for fusion processing, combining with the reverse cascading Transformer backbone network for multi-scale feature extraction, and using the Transformer encoder with a deformable attention module for feature encoding and multi-scale fusion, finally target detection is performed based on the detection head without anchor frame.

Benefits of technology

It improves the accuracy and robustness of object detection, can more effectively extract occlude target features, reduce model computing resources, enhance model modeling capabilities, and avoid generating a large number of redundant boxes, improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942095A_ABST
    Figure CN119942095A_ABST
Patent Text Reader

Abstract

The invention discloses a multispectral shielding target detection method and device, a storage medium and electronic equipment, relates to the technical field of image processing, and mainly aims to solve the problem of improving the target detection precision under complex and diversified shielding conditions, and the method comprises the steps: obtaining infrared images of different spectral bands in a target shielding scene, and obtaining the infrared images of different spectral bands in the target shielding scene; carrying out fusion processing on the infrared images of different spectral bands to obtain a fused infrared image; performing multi-scale feature extraction processing on the fused infrared image based on a backbone network of a reverse cascade Transform to obtain multi-level image features; performing feature coding processing on the large target layer feature map based on a Transform encoder with a deformable attention module to obtain a coded feature map; performing multi-scale fusion on the coding feature map, the small target layer feature map and the middle target layer feature map to obtain a preprocessed fusion feature image; and performing target detection processing on the preprocessed fusion feature image based on an anchor-frame-free detection head to obtain a target detection result in the target shielding scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a detection method and device for a multi-spectral occluded target, a storage medium, and an electronic device. Background Art

[0002] The task of object detection is to quickly and accurately locate the object in the image, and object detection in occluded scenes is an important topic in the field of object detection. When the object is occluded by other objects, inter-class occlusion occurs. Intra-class occlusion usually occurs in crowded scenes and seriously affects the detection performance.

[0003] At present, there are two types of occluded target algorithms: traditional algorithms and deep learning algorithms. Traditional occluded target detection algorithms mainly rely on prior information such as target saliency and background consistency to achieve detection. Most of these methods need to rely on prior information, and their robustness and versatility are poor in the face of complex occlusion in strong interference environments. The target detection method based on deep learning uses training to obtain the features of each layer of image data and then learns information to output the results through the classifier. Compared with traditional algorithms, deep learning methods have better robustness and accuracy, and can also effectively handle multiple targets and scale changes. Advanced typical models such as SSD, YOLOv8, sparse-R-CNN, etc. have good performance in the field of target detection, but when facing occluded targets, highly overlapping objects have similar features, which makes it difficult for the detector to extract effective features. Secondly, because of the serious overlap of objects, the NMS algorithm may mistakenly suppress some predictions. Therefore, for complex and diverse occlusion situations, it is difficult for current detection models to learn all occlusion types, and the accuracy of occluded target recognition is not high. Summary of the invention

[0004] In view of this, the present invention provides a multi-spectral obstructed target detection method and device, storage medium, and electronic device, the main purpose of which is to solve the problem of how to improve the accuracy of target detection under complex and diverse obstruction situations.

[0005] According to one aspect of the present invention, a method for detecting a multi-spectral obstructed target is provided, comprising: Acquire infrared images of different spectral bands in a target occlusion scene, and fuse a plurality of the infrared images of different spectral bands to obtain a fused infrared image; Based on the backbone network of the reverse cascade Transformer, multi-scale feature extraction processing is performed on the fused infrared image to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map and a large target layer feature map; Performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; Performing multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; The detection head based on the anchor-free frame performs target detection processing on the preprocessed fused feature image to obtain the target detection result in the target occlusion scene.

[0006] Furthermore, the backbone network based on the reverse cascade Transformer performs multi-scale feature extraction processing on the fused infrared image to obtain multi-level image features, including: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

[0007] Furthermore, the execution steps of the reverse cascade Transformer module include: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading; Using a cascade group attention mechanism to perform feature enhancement processing on the dimensionally upgraded input image to obtain an enhanced feature image; The enhanced feature image is processed by using an inverted multi-layer perceptron neural network to reduce the dimension, thereby obtaining a target feature representation.

[0008] Furthermore, the Transformer encoder with a deformable attention module performs feature encoding processing on the large target layer feature map to obtain an encoded feature map, including: Performing a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; A Transformer encoder is used to add position information to the large target feature vector, and the large target feature vector after adding the position information is input into a deformable attention module; The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; The offset feature vector is adjusted to a two-dimensional form to obtain the encoding feature map.

[0009] Further, the deformable attention module includes a multi-head cascade group attention and convolution operation; The multi-head cascade group attention inputs different cuts of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

[0010] Furthermore, the multi-scale fusion of the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image includes: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

[0011] Furthermore, the anchor-free detection head performs target detection processing on the preprocessed fusion feature image to obtain a target detection result in the target occlusion scene, including: The query and screening module in the detection head is used to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; Using the Top-K algorithm to filter the plurality of target prediction results to obtain a filtered target prediction result; The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scene.

[0012] According to another aspect of the present invention, a multi-spectral obstruction target detection device is provided, comprising: The first fusion module is used to obtain infrared images of different spectral bands in a target occlusion scene, and fuse multiple infrared images of different spectral bands to obtain a fused infrared image; A feature extraction module is used to perform multi-scale feature extraction processing on the fused infrared image based on a reverse cascade Transformer backbone network to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map; A feature encoding module, used for performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; A second fusion module is used to perform multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; The target detection module is used to perform target detection processing on the pre-processed fused feature image based on a detection head without an anchor frame to obtain a target detection result in the target occlusion scene.

[0013] Furthermore, the feature extraction module is also used for: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

[0014] Furthermore, the execution steps of the reverse cascade Transformer module in the feature extraction module include: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading; Using a cascade group attention mechanism to perform feature enhancement processing on the dimensionally upgraded input image to obtain an enhanced feature image; The enhanced feature image is processed by using an inverted multi-layer perceptron neural network to reduce the dimension, thereby obtaining a target feature representation.

[0015] Furthermore, the feature encoding module is also used for: Performing a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; A Transformer encoder is used to add position information to the large target feature vector, and the large target feature vector after adding the position information is input into a deformable attention module; The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; The offset feature vector is adjusted to a two-dimensional form to obtain the encoding feature map.

[0016] Further, the deformable attention module in the feature encoding module includes a multi-head cascade group attention and convolution operation; The multi-head cascade group attention inputs different cuts of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

[0017] Furthermore, the second fusion module is also used for: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

[0018] Furthermore, the target detection module is also used for: The query and screening module in the detection head is used to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; Using the Top-K algorithm to filter the plurality of target prediction results to obtain a filtered target prediction result; The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scene.

[0019] According to another aspect of the present invention, a storage medium is provided, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to execute operations corresponding to the above-mentioned multi-spectral obstruction target detection method.

[0020] According to another aspect of the present invention, there is provided an electronic device, comprising a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned multi-spectral obstruction target detection method.

[0021] By means of the above technical solution, the technical solution provided by the embodiment of the present invention has at least the following advantages: The present invention provides a detection method and device for multi-spectral occluded targets, a storage medium, and an electronic device. Compared with the prior art, the present invention obtains infrared images of different spectrum bands in the target occlusion scene, and fuses multiple infrared images of different spectrum bands to obtain a fused infrared image; a backbone network based on a reverse cascaded Transformer performs multi-scale feature extraction on the fused infrared image to obtain multi-level image features; wherein the backbone network has both local modeling of convolution and global modeling capabilities of Transformer, which can effectively extract the features of the occluded target while reducing the resources used for model calculation. In the face of strong interference and occlusion environments, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; the encoded feature map, the small target layer feature map and the medium target layer feature map are multi-scale fused to obtain a pre-processed fused feature image; wherein, the Transformer encoder with a deformable attention module combines the attention calculation of local and global receptive fields at the same time, which helps the model learn strong features, and the offset features can enable the network to selectively focus on more important areas. When facing occluded targets, the model will pay more attention to the visible part of the target, thereby enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the re-parameterization module is used to improve the reasoning speed, save memory usage, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the pre-processed fused feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted in this method is less likely to cause missed detection and false detection when facing occluded targets than the detection head of the existing algorithm. In the detection head part, only one prediction box is generated for each target, which avoids the algorithm from generating a large number of redundant boxes and improves the accuracy of the detection model when facing occluded targets.

[0022] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 A schematic flow chart of a multi-spectral obstruction target detection method provided by an embodiment of the present invention is shown; Figure 2 A schematic flow chart of another multi-spectral obstruction target detection method provided by an embodiment of the present invention is shown; Figure 3 A schematic flow chart of another multi-spectral obstruction target detection method provided by an embodiment of the present invention is shown; Figure 4 A schematic flow chart of another multi-spectral obstruction target detection method provided by an embodiment of the present invention is shown; Figure 5 A schematic diagram of the structure of a cascade group attention mechanism provided by an embodiment of the present invention is shown; Figure 6 A schematic diagram of the structure of a deformable attention module provided by an embodiment of the present invention is shown; Figure 7 A schematic structural diagram of a multi-spectral obstruction target detection device provided by an embodiment of the present invention is shown; Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0025] The embodiment of the present invention provides a method for detecting a multi-spectral obstructed target, such as Figure 1 As shown, the method includes: 101. Acquire infrared images of different spectral bands in a target occlusion scene, and fuse a plurality of the infrared images of different spectral bands to obtain a fused infrared image; In an embodiment of the present invention, the current execution end obtains infrared images of different spectral bands in a target occlusion scene. Among them, infrared images of different spectral bands represent infrared images taken based on different wavelength ranges. Since the wavelength ranges of infrared images are different, infrared images of different spectral bands can obtain more and richer information in the target occlusion scene. After obtaining infrared images of different spectral bands, the current execution end fuses multiple infrared images of different spectral bands to obtain a fused infrared image. For example, after obtaining three infrared images of different spectral bands with a size of 640×640, the three single-channel images are fused into a three-channel image, that is, the three-channel image is a fused infrared image, which is not specifically limited in the embodiment of the present invention.

[0026] 102. Perform multi-scale feature extraction processing on the fused infrared image based on a reverse cascade Transformer backbone network to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map; In the embodiment of the present invention, the current execution end performs multi-scale feature extraction processing on the fused infrared image obtained in step 101 based on the backbone network of the reverse cascade transformer, wherein the backbone network of the reverse cascade transformer is obtained by reverse cascading the transformer module on the basis of the backbone network. The backbone network based on the reverse cascade transformer has both local modeling of convolution and relies on the global modeling ability of the transformer, which can effectively extract the features of the occluded target while reducing the resources used for model calculation. In the face of strong interference occlusion environment, it has strong robustness and stability. The multi-level image features obtained by feature extraction in this embodiment include a small target layer feature map, a medium target layer feature map and a large target layer feature map. Among them, the image dimensions of the small target layer feature map, the medium target layer feature map and the large target layer feature map are arranged in descending order. For example, if the original image dimension is 32, the dimension of the small target layer feature map can be 1 / 8 of the original image dimension, the dimension of the medium target layer feature map can be 1 / 16 of the original image dimension, and the dimension of the large target layer feature map can be 1 / 32 of the original image dimension, etc., which is not specifically limited in the embodiment of the present invention.

[0027] 103. Performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; In an embodiment of the present invention, the current execution end performs feature encoding processing on the large target layer feature map in the multi-level image features based on the Transformer encoder with a deformable attention module. Among them, the deformable attention module adopts deformable convolution. Compared with ordinary convolution, the deformable convolution adds an offset for expansion, so that the sampling position becomes an irregular position, making the receptive field closer to the shape of the actual target. In addition, the deformable attention module can model global relationships. This method combines the attention calculation of local and global receptive fields at the same time, which helps the model learn strong features. The module obtains the offset features through the offset network, which enables the network to selectively focus on more important areas. When facing occluded targets, the model will pay more attention to the visible part of the target, enhancing the modeling ability of the model.

[0028] 104. Perform multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; 105. Perform target detection processing on the preprocessed fused feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene.

[0029] In an embodiment of the present invention, the current execution end performs target detection processing on the preprocessed fusion feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene. Among them, the detection head without anchor frame includes a query screening module and a Transformer decoder. The query screening module includes an allocation stage and a loss calculation stage. The allocation stage is used for target matching processing, and the loss calculation stage is used to calculate the loss value of the prediction result, so as to facilitate the optimization of the model.

[0030] It should be noted that in the embodiment of the present invention, the current execution end initializes the anchor frame as a position query of the subsequent decoder, so the final output consists of the refined anchor frame and the refined classification results of the content feature prediction.

[0031] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to extract more abundant features of the occluding target, another multi-spectral occluding target detection method is provided, such as Figure 2 As shown, the step performs multi-scale feature extraction processing on the fused infrared image based on the backbone network of the reverse cascade Transformer to obtain multi-level image features, including: 201. Perform dimensionality-upgrading processing on the fused infrared image using a multi-layer perceptron neural network to obtain a feature map y1; In the embodiment of the present invention, the current execution end uses a multilayer perceptron (MLP) neural network to perform dimensionality increase processing on the fused infrared image to obtain a feature map y1. The specific formula is as follows:

[0032] in, MLP e () represents the multi-layer perceptron model, X is the fused infrared image input to the model, represents the ratio of output to input, C is the number of feature map channels, H is the feature map height, and W is the feature map width. It is the feature map after the feature map X is upgraded in dimension. In the embodiment of the present invention, it is the feature map y1 obtained after the dimension upgrade.

[0033] 202. Use the backbone network to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; In an embodiment of the present invention, the current execution end uses a backbone network to perform convolution and maximum pooling on the feature map y1 obtained in step 201, wherein the backbone network preferably uses the basic architecture of Resnet-18, and the feature map y1 is subjected to 7×7 convolution of Resnet-18 and then to a maximum pooling layer to obtain the feature map y2, which is not specifically limited in the embodiment of the present invention.

[0034] 203. Perform first-scale feature extraction processing on the feature map y2 using a reverse cascade Transformer module to obtain a small target layer feature map y3; In an embodiment of the present invention, the current execution end uses a reverse cascade Transformer module to perform a first-scale feature extraction process on the feature map y2 obtained in step 202. Specifically, the feature map y2 passes through two layers of reverse cascade Transformer modules to obtain a small target layer feature map y3 with a dimension size of 1 / 8 of the original image, which is not specifically limited in the embodiment of the present invention.

[0035] 204. Perform second-scale feature extraction processing on the small target layer feature map y3 using a reverse cascade Transformer module to obtain a medium target layer feature map y4. In an embodiment of the present invention, the current execution end uses a reverse cascade Transformer module to perform a second-scale feature extraction process on the small target layer feature map y3 obtained in step 203. Specifically, the small target layer feature map y3 is subjected to a layer of reverse cascade Transformer module to obtain a medium target layer feature map y4 with a dimension size of 1 / 16 of the original image. The embodiment of the present invention does not make any specific limitation.

[0036] 205. Use a reverse cascade Transformer module to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

[0037] In an embodiment of the present invention, the current execution end uses a reverse cascade Transformer module to perform third-scale feature extraction processing on the middle target layer feature map y4 obtained in step 204. Specifically, the middle target layer feature map y4 is subjected to a layer of reverse cascade Transformer module to obtain a large target layer feature map y5 with a dimension size of 1 / 32 of the original image, which is not specifically limited in the embodiment of the present invention.

[0038] It should be noted that in the above embodiment, the execution steps of reverse cascading Transformer modules include: (1) A multilayer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading. The formula for performing dimension-upgrading processing on the input image by the multilayer perceptron (MLP) neural network is as follows:

[0039] in, MLP e () represents the multi-layer perceptron model, X is the input image, represents the ratio of output to input, C is the number of feature map channels, H is the feature map height, and W is the feature map width. It is the feature map after the input image X is dimensionally upgraded. In step 203 of the embodiment of the present invention, the input image is the feature map y2; in step 204 of the embodiment of the present invention, the input image is the small target layer feature map y3; in step 205 of the embodiment of the present invention, the input image is the medium target layer feature map y4.

[0040] (2) using a cascade group attention mechanism to perform feature enhancement processing on the input image after dimensionality increase to obtain an enhanced feature image; wherein the cascade group attention mechanism performs feature enhancement processing on the input image after dimensionality increase The specific formula for feature enhancement is as follows:

[0041] in, The feature map go through Feature map of the operation. The operation is an efficient inverse difference structure, which has both local modeling of convolution and global modeling capability of Transformer, as shown in the following formula:

[0042] Among them, DW-Conv is a depthwise separable convolution, Skip is a skip connection operation, and CG-MHSA is a multi-head cascade group attention. It includes cascaded multi-head cascade group attention and convolution operations. In order to save computational overhead, depthwise separable convolution is used to implement convolution, and multi-head attention uses cascade group attention to expand the diversity of feature maps and optimize feature representation.

[0043] Cascade group attention Figure 5 As shown. Transformer's multi-head self-attention mechanism can be used to capture the contextual information of features. In the occluded target detection task, it will pay more attention to the feature information of the visible part of the target and ignore the impact of the occluded part on the detection. However, the attention heads of the multi-head self-attention mechanism are prone to redundancy, resulting in low computational efficiency. The cascade group attention inputs each head with different cuts of the complete feature, thereby explicitly decomposing the attention calculation into different heads. As shown in the following formula:

[0044]

[0045] Among them, Attn is the attention operation, , , is the weight matrix, Concat is the concatenation operation, represents the projection matrix after stitching, is the input feature The jth segmentation of is the feature after the attention module, Input features These divisions consist of is the total number of attention heads. Splitting features can make attention head input more efficient and save computing resources, and can control the number of split features to improve the diversity of attention feature maps. Finally, the output of each attention is added to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

[0046] (3) Using an inverted multi-layer perceptron neural network to perform dimensionality reduction processing on the enhanced feature image to obtain a target feature representation. The specific formula for dimensionality reduction is as follows:

[0047] in, The input / output ratio representing the inversion is A multi-layer perceptron neural network is used to shrink the channel dimension. The feature map Feature map after dimension reduction.

[0048] (4) Finally, the final output can be obtained through the residual structure: , the embodiments of the present invention are not specifically limited.

[0049] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to better focus on the feature information of the visible part of the target and ignore the influence of the occluded part on the detection, another multi-spectral occluded target detection method is provided, such as Figure 3 As shown, the step performs feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map, including: 301. Perform a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; In the embodiment of the present invention, the current execution end performs a convolution operation on the large target layer feature map y5 obtained from step 201 to step 205, so that the graphic information is reduced to vector information, and the corresponding large target feature vector is obtained. Then it is handed over to the Transformer encoder with a deformable attention module for processing.

[0050] 302. Use a Transformer encoder to add position information to the large target feature vector, and input the large target feature vector after adding the position information into a deformable attention module; 303. The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; 304. Adjust the offset feature vector into a two-dimensional form to obtain the encoding feature map.

[0051] In the embodiment of the present invention, the current execution end uses the Transformer encoder to add position information to the input large target feature vector, and then inputs the large target feature vector with the added position information into the deformable attention module. The output of the previous step is added with the residual structure for addition regularization, and the residual output is added through the feedforward module. Among them, the deformable attention module uses deformable convolution. Compared with ordinary convolution, the deformable convolution adds an offset for expansion, so that the sampling position becomes an irregular position, making the receptive field closer to the shape of the actual target. The deformable attention module is as follows Figure 6 As shown, the input feature image , whose size is For the input image With proportionality factor Downsampling generates a reference grid, where the reference point P represents the coordinate value, which is obtained by normalizing the coordinate value. The feature image X is obtained by linearly projecting the feature , It is derived from the offset network with query as input and the obtained Add it to the reference point to get the offset position information, as shown in the following formula:

[0052] Among them, x represents the large target feature vector, Wq is the projection matrix, and q is the linear projection result; represents the offset network, is a bilinear interpolation operation, and is the weight matrix, biasing the network Use the query feature to output the offset value of the reference point The input features are first subjected to a deep convolution to capture local features. Then the GELU activation function and 1×1 convolution are used to obtain the two-dimensional offset. The deformed reference points are sampled using bilinear interpolation to obtain the sampling result. In the sampling results Multi-head self-attention calculations are performed, and relative position offset embedding is added. Finally, the final attention module output is obtained through the projection matrix. The deformable attention module can model global relationships. This method combines the attention calculations of local and global receptive fields to help the model learn strong features, and this module is Calculating the offset network and obtaining the offset features can enable the network to selectively focus on more important areas. When facing occluded targets, the model will pay more attention to the visible part of the target, enhancing the modeling ability of the model.

[0053] It should be noted that in the above embodiment, the deformable attention module includes multi-head cascade group attention and convolution operations; the multi-head cascade group attention inputs different splits of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

[0054] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to better reflect the extracted feature information on a feature map, another multi-spectral occluded target detection method is provided, wherein the coding feature map, the small target layer feature map and the medium target layer feature map are multi-scale fused to obtain a pre-processed fused feature image, including: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

[0055] In an embodiment of the present invention, the current execution end performs multi-scale fusion on the encoded feature map F5 obtained from step 301 to step 304, the small target layer feature map y3 obtained from step 203, and the medium target layer feature map y4 obtained from step 204. Specifically, the multi-scale fusion process uses two paths from top to bottom and from bottom to top to perform feature fusion. Among them, the features from different levels use a 1×1 convolution layer to map the number of channels to the same number, and then perform feature fusion. The fusion module uses two paths, one uses convolution to adjust the number of channels, and the other uses convolution and re-parameterization module to extract features, and finally the two paths are spliced ​​and output. The re-parameterization module uses two network structure models during training and inference, respectively, to improve the inference speed, save memory usage, and facilitate model deployment and acceleration.

[0056] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to avoid the algorithm from generating a large number of redundant frames and improve the accuracy of the detection model facing the occluded target, another multi-spectral occluded target detection method is provided, such as Figure 4 As shown, the step performs target detection processing on the preprocessed fusion feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene, including: 401. Use the query and screening module in the detection head to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; 402. Using a Top-K algorithm to filter the plurality of target prediction results to obtain a filtered target prediction result; 403. Use the Transformer decoder in the detection head to decode the filtered target prediction result to obtain the target detection result in the target occlusion scenario.

[0057] In an embodiment of the present invention, the current execution end inputs the preprocessed fusion feature map into the detection head part, and the detection head includes a query screening module query and a Transformer decoder. When the existing detection model using the NMS algorithm faces an occluded target, the prediction boxes may overlap seriously during occlusion, so the prediction boxes of different targets may be regarded as the prediction of one target by the NMS algorithm and mistakenly suppressed, resulting in missed detection. In addition, the threshold scores adopted during training and actual operation are not the same, so that the network does not perform at its true level during reasoning. In an embodiment of the present invention, the detection head part does not use the NMS algorithm and threshold screening operations, and only performs a TOPK operation on the final prediction result. It is simpler and faster in the post-processing part, and only generates one prediction box for each target, avoiding the algorithm from generating a large number of redundant boxes.

[0058] Specifically, the preprocessed fusion feature map is input into the query screening module query. The query screening module includes an allocation stage and a loss calculation stage. The allocation stage performs target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the loss calculation stage adopts the Repulsion Loss loss function, and the specific formula is as follows:

[0059] The above loss function is divided into three parts. The first part is the loss value (attraction term) generated by the predicted box and the true target box, which is recorded as ; The second part is the loss value generated by the predicted box and the adjacent real target box (repulsion term (RepGT)), recorded as ; The third part is the loss value generated by the prediction box and the adjacent prediction box that does not predict the same real target (repulsion Box (RepBox)), recorded as . and is the correlation coefficient, which is used to balance the two parts of the repulsion loss value.

[0060] It should be noted that in the embodiment of the present invention, the current execution end initializes the anchor frame as the position query of the subsequent decoder, and this parameter is learnable. The loss function used for occluding the target during training regresses and updates the parameters. In the subsequent decoder, the denoising idea of ​​DINO HEAD is adopted, and two hyperparameters are used to generate positive samples and negative samples. The noise scale of the positive sample is smaller than the smaller hyperparameter, and the noise scale of the negative sample is between the two hyperparameters. Positive samples are expected to predict targets, while negative samples are expected to have no targets. By generating difficult negative samples with a small gap with the positive samples, repeated predictions are avoided and confusion is suppressed. The final output consists of the classification results predicted by the refined anchor frame and the refined content feature. Finally, the TOPK operation is performed to sort the targets of this type to obtain the final target detection result.

[0061] The embodiment of the present invention provides a method for detecting multi-spectral occluded targets. Compared with the prior art, the present invention obtains infrared images of different spectrum bands in the target occlusion scene, and fuses multiple infrared images of different spectrum bands to obtain a fused infrared image; a backbone network based on the reverse cascade Transformer performs multi-scale feature extraction on the fused infrared image to obtain multi-level image features; wherein the backbone network has both local modeling of convolution and global modeling capabilities of Transformer, which can effectively extract the features of the occluded target while reducing the resources used for model calculation. In the face of strong interference occlusion environment, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; the encoded feature map, the small target layer feature map and the medium target layer feature map are multi-scale fused to obtain a pre-processed fused feature image; wherein, the Transformer encoder with a deformable attention module combines the attention calculation of local and global receptive fields at the same time, which helps the model learn strong features, and the offset features can enable the network to selectively focus on more important areas. When facing occluded targets, the model will pay more attention to the visible part of the target, thereby enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the re-parameterization module is used to improve the reasoning speed, save memory usage, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the pre-processed fused feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted in this method is less likely to cause missed detection and false detection when facing occluded targets than the detection head of the existing algorithm. In the detection head part, only one prediction box is generated for each target, which avoids the algorithm from generating a large number of redundant boxes and improves the accuracy of the detection model when facing occluded targets.

[0062] As the above Figure 1 The embodiment of the present invention provides a multi-spectral occlusion target detection device, such as Figure 7 As shown, the device comprises: The first fusion module 51 is used to obtain infrared images of different spectrum bands in the target occlusion scene, and fuse multiple infrared images of different spectrum bands to obtain a fused infrared image; A feature extraction module 52 is used to perform multi-scale feature extraction processing on the fused infrared image based on a reverse cascade Transformer backbone network to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map; A feature encoding module 53 is used to perform feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; A second fusion module 54 is used to perform multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a pre-processed fused feature image; The target detection module 55 is used to perform target detection processing on the pre-processed fused feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene.

[0063] Furthermore, the feature extraction module 52 is also used for: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

[0064] Furthermore, the execution steps of the reverse cascade Transformer module in the feature extraction module 52 include: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading; Using a cascade group attention mechanism to perform feature enhancement processing on the dimensionally upgraded input image to obtain an enhanced feature image; The enhanced feature image is processed by using an inverted multi-layer perceptron neural network to reduce the dimension, thereby obtaining a target feature representation.

[0065] Furthermore, the feature encoding module 53 is also used for: Performing a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; A Transformer encoder is used to add position information to the large target feature vector, and the large target feature vector after adding the position information is input into a deformable attention module; The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; The offset feature vector is adjusted to a two-dimensional form to obtain the encoding feature map.

[0066] Further, the deformable attention module in the feature encoding module 53 includes multi-head cascade group attention and convolution operations; The multi-head cascade group attention inputs different cuts of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

[0067] Furthermore, the second fusion module 54 is also used for: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

[0068] Furthermore, the target detection module 55 is also used for: The query and screening module in the detection head is used to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; Using the Top-K algorithm to filter the plurality of target prediction results to obtain a filtered target prediction result; The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scene.

[0069] The embodiment of the present invention provides a detection device for multi-spectral occluded targets. Compared with the prior art, the present invention obtains infrared images of different spectrum bands in the target occlusion scene, and fuses multiple infrared images of different spectrum bands to obtain a fused infrared image; a backbone network based on the reverse cascade Transformer performs multi-scale feature extraction on the fused infrared image to obtain multi-level image features; wherein the backbone network has both local modeling of convolution and global modeling capabilities of Transformer, which can effectively extract the features of the occluded target while reducing the resources used for model calculation. In the face of strong interference and occlusion environments, it has strong robustness and stability. Secondly, the present invention performs feature encoding processing on the large target layer feature map based on the Transformer encoder with a deformable attention module to obtain an encoded feature map; the encoded feature map, the small target layer feature map and the medium target layer feature map are multi-scale fused to obtain a pre-processed fused feature image; wherein, the Transformer encoder with a deformable attention module combines the attention calculation of local and global receptive fields at the same time, which helps the model learn strong features, and the offset features can enable the network to selectively focus on more important areas. When facing occluded targets, the model will pay more attention to the visible part of the target, thereby enhancing the modeling ability of the model. The multi-scale fusion module is used to fuse deep and shallow information, and the re-parameterization module is used to improve the reasoning speed, save memory usage, and facilitate model deployment and acceleration. Furthermore, the present invention performs target detection processing on the pre-processed fused feature image based on the detection head without anchor frame to obtain the target detection result in the target occlusion scene. Compared with a specific multi-spectral occluded target detection method, the anchor-free detection head adopted in this method is less likely to cause missed detection and false detection when facing occluded targets than the detection head of the existing algorithm. In the detection head part, only one prediction box is generated for each target, which avoids the algorithm from generating a large number of redundant boxes and improves the accuracy of the detection model when facing occluded targets.

[0070] According to one embodiment of the present invention, a storage medium is provided, wherein the storage medium stores at least one executable instruction, and the computer executable instruction can execute the multi-spectral obstruction target detection method in any of the above method embodiments.

[0071] Figure 8 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the electronic device.

[0072] like Figure 8As shown, the electronic device may include: a processor (processor) 602 , a communication interface (Communications Interface) 604 , a memory (memory) 606 , and a communication bus 608 .

[0073] The processor 602 , the communication interface 604 , and the memory 606 communicate with each other via a communication bus 608 .

[0074] The communication interface 604 is used to communicate with other devices such as clients or other servers.

[0075] The processor 602 is used to execute the program 610, and specifically can execute the relevant steps of the above-mentioned multi-spectral obstruction target detection method.

[0076] Specifically, the program 610 may include program codes, which include computer operation instructions.

[0077] The processor 602 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present invention. The one or more processors included in the electronic device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0078] The memory 606 is used to store the program 610. The memory 606 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0079] The program 610 may be specifically configured to enable the processor 602 to perform the following operations: Acquire infrared images of different spectral bands in a target occlusion scene, and fuse a plurality of the infrared images of different spectral bands to obtain a fused infrared image; Based on the backbone network of the reverse cascade Transformer, multi-scale feature extraction processing is performed on the fused infrared image to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map and a large target layer feature map; Performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; Performing multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; The detection head based on the anchor-free frame performs target detection processing on the preprocessed fused feature image to obtain the target detection result in the target occlusion scene.

[0080] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0081] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-spectral obstruction target detection method, characterized in that: include: Acquire infrared images of different spectral bands in a target occlusion scene, and fuse a plurality of the infrared images of different spectral bands to obtain a fused infrared image; Based on the backbone network of the reverse cascade Transformer, multi-scale feature extraction processing is performed on the fused infrared image to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map and a large target layer feature map; Performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; Performing multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; The detection head based on the anchor-free frame performs target detection processing on the preprocessed fused feature image to obtain the target detection result in the target occlusion scene.

2. The method according to claim 1, characterized in that The backbone network based on the reverse cascade Transformer performs multi-scale feature extraction processing on the fused infrared image to obtain multi-level image features, including: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the fused infrared image to obtain a feature map y1; The backbone network is used to perform convolution and maximum pooling processing on the feature map y1 to obtain a feature map y2; A reverse cascade Transformer module is used to perform first-scale feature extraction processing on the feature map y2 to obtain a small target layer feature map y3; A reverse cascade Transformer module is used to perform second-scale feature extraction processing on the small target layer feature map y3 to obtain a medium target layer feature map y4; The reverse cascade Transformer module is used to perform third-scale feature extraction processing on the middle target layer feature map y4 to obtain a large target feature map y5.

3. The method according to claim 2, characterized in that The execution steps of the reverse cascade Transformer module include: A multi-layer perceptron neural network is used to perform dimension-upgrading processing on the input image to obtain an input image after dimension-upgrading; the input image is one of the feature map y2, the small target layer feature map y3 and the medium target layer feature map y4; Using a cascade group attention mechanism to perform feature enhancement processing on the dimensionally upgraded input image to obtain an enhanced feature image; The enhanced feature image is processed by using an inverted multi-layer perceptron neural network to reduce the dimension, thereby obtaining a target feature representation.

4. The method according to claim 1, characterized in that The Transformer encoder with a deformable attention module performs feature encoding processing on the large target layer feature map to obtain an encoded feature map, including: Performing a convolution operation on the large target layer feature map to obtain a corresponding large target feature vector; A Transformer encoder is used to add position information to the large target feature vector, and the large target feature vector after adding the position information is input into a deformable attention module; The deformable attention module performs an offset process on the large target feature vector after adding the position information to obtain an offset feature vector corresponding to the large target feature vector; The offset feature vector is adjusted to a two-dimensional form to obtain the encoding feature map.

5. The method according to claim 4, characterized in that The deformable attention module includes multi-head cascade group attention and convolution operations; The multi-head cascade group attention inputs different cuts of the large target feature vector after adding the position information into each head for attention calculation, and adds the output of each attention to the subsequent attention in a cascade manner to optimize the feature representation and calculate each attention feature map.

6. The method according to claim 1, characterized in that The multi-scale fusion of the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image includes: Mapping the number of channels of the coded feature map, the small target layer feature map and the medium target layer feature map to the same number, and obtaining a coded feature map, a small target layer feature map and a medium target layer feature map respectively; The convolution and re-parameter modules are used to respectively extract the features of the encoding feature map, the small target layer feature map and the medium target layer feature map, and the obtained features are spliced ​​to obtain the pre-processed fused feature image.

7. The method according to any one of claims 1 to 6, characterized in that: The anchor-free detection head performs target detection processing on the preprocessed fusion feature image to obtain a target detection result in the target occlusion scene, including: The query and screening module in the detection head is used to perform target matching processing on the preprocessed fusion feature map to obtain multiple target prediction results; the query and screening module is obtained by model optimization based on the loss value and is used for target prediction; the loss value is calculated based on the target prediction result, the real target, the adjacent real target and the adjacent non-same predicted target; Using the Top-K algorithm to filter the plurality of target prediction results to obtain filtered target prediction results; The Transformer decoder in the detection head is used to decode the filtered target prediction result to obtain the target detection result in the target occlusion scene.

8. A multi-spectral obstruction target detection device, characterized in that: include: The first fusion module is used to obtain infrared images of different spectral bands in a target occlusion scene, and fuse multiple infrared images of different spectral bands to obtain a fused infrared image; A feature extraction module is used to perform multi-scale feature extraction processing on the fused infrared image based on a reverse cascade Transformer backbone network to obtain multi-level image features; the multi-level image features include a small target layer feature map, a medium target layer feature map, and a large target layer feature map; A feature encoding module, used for performing feature encoding processing on the large target layer feature map based on a Transformer encoder with a deformable attention module to obtain an encoded feature map; A second fusion module is used to perform multi-scale fusion on the encoding feature map, the small target layer feature map and the medium target layer feature map to obtain a preprocessed fused feature image; The target detection module is used to perform target detection processing on the pre-processed fused feature image based on a detection head without an anchor frame to obtain a target detection result in the target occlusion scene.

9. A storage medium, characterized in that: The storage medium stores at least one executable instruction, and the executable instruction executes an operation corresponding to the multi-spectral obstruction target detection method as described in any one of claims 1-7.

10. An electronic device, characterized in that: It includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the multi-spectral obstruction target detection method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Remote sensing image target fine granularity identification method, system and device and storage medium

    CN115019182A

  • Infrared target detection method and device, computer equipment and storage medium

    CN118229961A

  • Infrared image target detection method based on Transform architecture

    CN118314333A

  • Remote sensing image target detection method based on deformable and depth separation convolution

    CN118865093A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1