Lightweight aviation multispectral target detection method

By adopting multi-scale extraction network, neck network and head network in lightweight aviation multi-spectral object detection, combined with the efficient cross attention fusion module and SIEM module, the problem of insufficient fusion of multimodal features is solved, and efficient and low-complex multi-spectral object detection is achieved.

CN120032196AActive Publication Date: 2025-05-23CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510520981.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing multimodal feature fusion is insufficient in lightweight aviation multispectral object detection, and the small-scale fusion feature map lacks semantic information and global receptive fields, resulting in limited detection performance.

Method used

A lightweight aviation multispectral object detection method is designed, using multi-scale extraction network, neck network and head network, and the interaction and fusion of different modal features are achieved through the efficient cross-attention fusion module and the SIEM module, reducing the computational complexity and enhancing semantic information and global perception capabilities.

Benefits of technology

It realizes full interaction and fusion of features between different modes, improves detection accuracy and efficiency, reduces calculation complexity and memory usage, and is suitable for lightweight aviation multi-spectral multi-object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032196A_ABST
    Figure CN120032196A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a lightweight aviation multispectral target detection method. The method comprises the following steps: S1, acquiring a visible light image and an infrared image in the same shooting scene; s2, constructing a multispectral target detection network, wherein the multispectral target detection network comprises a multi-scale extraction network, a neck network and a head network; s3, inputting the visible light image and the infrared image into a multispectral target detection network for training to obtain a multispectral target detection model; and S4, inputting the to-be-detected visible light image and the to-be-detected infrared image in the same shooting scene into the multispectral target detection model for detection to obtain a target position and a target category. The method has the advantages of low calculation complexity and small video memory occupation, and can meet the actual application requirements of lightweight aviation multispectral multi-target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a lightweight aviation multi-spectral target detection method. Background Art

[0002] Target detection algorithms can quickly extract information from images or videos, accurately locate targets of interest, and determine their categories. However, the application of single-modal target detection is limited by environmental influences and has certain limitations. For example, under conditions of insufficient lighting such as at night or in bad weather, target detectors based on visible light images find it difficult to extract effective features, thus affecting detection performance; and relying solely on infrared images, target detectors lack color and texture information and find it difficult to accurately determine target categories. In order to solve this problem, multispectral target detection significantly improves the performance of target detectors in all-weather target detection by fusing features of images of different modalities, and exhibits excellent robustness and stability in different scenarios, which has important research significance and application value.

[0003] In multispectral target detection, how to effectively fuse different modal information is a key issue. Based on the fusion position, multispectral target detection algorithms can be divided into pixel-level fusion, feature-level fusion, and decision-level fusion. Pixel-level fusion fuses images of different modalities at the original data layer. Pixel-level fusion requires additional fusion steps and cannot achieve end-to-end detection. Decision-level fusion uses two models to detect visible light images and infrared images respectively, and screens the detection results through methods such as non-maximum suppression. The operation is simple, but there is a lack of information interaction between different modalities. At the same time, the computational complexity is greatly increased. Feature-level fusion fuses feature maps at different stages in the target detection model by designing various fusion modules, making full use of the multi-scale features of different modalities, and has high detection accuracy and efficiency. Summary of the invention

[0004] In view of this, the present invention aims to provide a lightweight aerial multispectral target detection method to solve the problems of insufficient existing multimodal feature fusion, lack of semantic information and global receptive field in small-scale fusion feature maps, etc. The present invention is aimed at the practical application needs of aerial remote sensing all-weather, multi-scene and high-precision target detection, and realizes the feature interaction between different modalities and feature mining and automatic fusion detection within each modality. The present invention has the advantages of low computational complexity and small video memory occupancy, and can meet the practical application needs of lightweight aerial multispectral multi-target detection.

[0005] To achieve the above object, the technical solution created by the present invention is implemented as follows: A lightweight aviation multispectral target detection method specifically comprises the following steps: S1: Acquire visible light images and infrared images of the same shooting scene; S2: Construct a multispectral target detection network, which includes a multi-scale extraction network, a neck network, and a head network; The multi-scale extraction network is used to extract, interact and fuse information from visible light images and infrared images. The fused features are input into the head network via the neck network. S3: Input the visible light image and the infrared image into the multispectral target detection network for training to obtain a multispectral target detection model; S4: Input the visible light image to be detected and the infrared image to be detected in the same shooting scene into the multispectral target detection model for detection to obtain the target position and target category.

[0006] Furthermore, the multi-scale extraction network includes a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a seventh feature extraction module, an eighth feature extraction module, a ninth feature extraction module, a tenth feature extraction module, a first efficient cross-attention fusion module, a second efficient cross-attention fusion module, a third efficient cross-attention fusion module, a fourth efficient cross-attention fusion module and an SPPF module; wherein the infrared image is processed by the first feature extraction module and the second feature extraction module to obtain a feature map A1; the visible light image is processed by the third feature extraction module and the fourth feature extraction module to obtain a feature map A2; the feature map A1 and the feature map A2 are input into the first efficient cross-attention fusion module for processing to obtain feature maps A3 and A4 respectively; the feature map A3 is input into the fifth feature extraction module for processing to obtain Feature map A5, input feature map A4 to the sixth feature extraction module for processing to obtain feature map A6, input feature map A5 and feature map A6 to the second efficient cross-attention fusion module for processing, and obtain feature map A7 and feature map A8 accordingly; input feature map A7 to the seventh feature extraction module for processing to obtain feature map A9, input feature map A8 to the eighth feature extraction module for processing to obtain feature map A10, input feature map A9 and feature map A10 to the third efficient cross-attention fusion module for processing, and obtain feature map A11 and feature map A12 accordingly; input feature map A11 to the ninth feature extraction module for processing to obtain feature map A13, input feature map A12 to the tenth feature extraction module for processing to obtain feature map A14, input feature map A13 and feature map A14 to the fourth efficient cross-attention fusion module for processing, and obtain feature map A15 and feature map A16 accordingly; Add feature map A3 and feature map A4 to obtain the first fusion feature; add feature map A7 and feature map A8 to obtain the second fusion feature; add feature map A11 and feature map A12 to obtain the third fusion feature; add feature map A15 and feature map A16 to obtain feature map A17; input feature map A17 into the SPPF module for processing to obtain the fourth fusion feature.

[0007] Furthermore, the first feature extraction module and the third feature extraction module have the same network structure, and the first feature extraction module includes a first CBS module and a second CBS module having the same network structure; The network structures of the second feature extraction module, the fourth feature extraction module, the fifth feature extraction module, the sixth feature extraction module, the seventh feature extraction module, the eighth feature extraction module, the ninth feature extraction module and the tenth feature extraction module are the same, and the second feature extraction module includes a third CBS module and a first C2F module connected in sequence.

[0008] Furthermore, the network structures of the first efficient cross-attention fusion module, the second efficient cross-attention fusion module, the third efficient cross-attention fusion module and the fourth efficient cross-attention fusion module are the same. The first efficient cross-attention fusion module includes a fourth CBS module, a fifth CBS module, a first downsampling layer, a second downsampling layer, a third downsampling layer, a first Conv module, a second Conv module, a third Conv module, a fourth Conv module, a fifth Conv module, a sixth Conv module, a visible light cross-attention mechanism, an infrared cross-attention mechanism, a first upsampling layer, a second upsampling layer, and a first MLP module; wherein, the feature map A1 is processed by the first downsampling layer and the first Conv module to obtain the feature map B1; the feature map A2 is processed by the second downsampling layer and the second Conv module to obtain the feature map B2; the feature map A1 and the feature map A2 are spliced ​​to obtain the feature map B3, and the feature map B3 is processed by the fourth CBS module, the third downsampling layer and the third Conv module The feature map B4 is processed by the fourth Conv module to obtain the feature map B7, the feature map A2 is input to the fifth Conv module for processing to obtain the feature map B8, the feature map B7 is added to the visible light attention matrix to obtain the feature map B9, the feature map B8 is added to the infrared attention matrix to obtain the feature map B10, the feature map B9 and the feature map B10 are spliced ​​to obtain the feature map B11, the feature map B11 is input to the sixth Conv module for processing to obtain the feature map B12, the feature map A1 and the feature map A2 are spliced ​​and processed by CBS in the channel dimension to obtain the feature map F f , the feature map F f Input to the fifth CBS module for processing to obtain feature map B13, add feature map B12 and feature map B13 to obtain feature map B14, input feature map B14 to the first MLP module for processing to obtain feature map B15, split feature map B15 into feature map B17 and feature map B18, and add feature map B17 and feature map A1 to obtain feature map A3, add feature map B18 and feature map A2 to obtain feature map A4, add feature map A3 and feature map A4 to obtain the first fusion feature.

[0009] Furthermore, the network structures of the first CBS module, the second CBS module, the third CBS module, the fourth CBS module and the fifth CBS module are the same, the first CBS module includes a 2d convolution layer, a batch normalization layer and a SiLU activation function connected in sequence, and the convolution kernel of the 2d convolution layer is 3×3 and the stride is 2; The network structures of the first downsampling layer, the second downsampling layer and the third downsampling layer are the same, wherein the first downsampling layer includes a maximum pooling layer and an average pooling layer. After the feature map A1 is processed by the maximum pooling layer and the average pooling layer respectively, the processing results of the two are multiplied by their respective assigned weights and then added, and the added result is input into the first Conv module for processing. The weight assigned to the branch where the maximum pooling layer is located is 0.5, and the weight assigned to the branch where the average pooling layer is located is 0.5; The feature scaling factor of the downsampling operation in the first efficient cross-attention fusion module is set to 8, the feature scaling factor of the downsampling operation in the second efficient cross-attention fusion module is set to 4, the feature scaling factor of the downsampling operation in the third efficient cross-attention fusion module is set to 2, and the feature scaling factor of the downsampling operation in the fourth efficient cross-attention fusion module is set to 1.

[0010] Furthermore, the neck network includes a second C2F module, a third C2F module, a fourth C2F module, a fifth C2F module, a sixth C2F module and a seventh C2F module, and the fourth fusion feature is upsampled and concatenated with the third fusion feature to obtain a feature map A18, and the feature map A18 is input to the second C2F module for processing to obtain a feature map A19; the feature map A19 is upsampled and concatenated with the second fusion feature to obtain a feature map A20, and the feature map A20 is input to the third C2F module for processing to obtain a feature map A21; the feature map A21 is upsampled and concatenated with the first fusion feature to obtain a feature map A22, and the feature map A23 is input to the third C2F module for processing to obtain a feature map A24. The feature map A22 is input to the fourth C2F module for processing to obtain a feature map A23; the feature map A23 is downsampled and concatenated with the feature map A21 to obtain a feature map A24, and the feature map A24 is input to the fifth C2F module for processing to obtain a scale feature map P3, and the scale feature map P3 is downsampled and concatenated with the feature map A19 to obtain a feature map A25, and the feature map A25 is input to the sixth C2F module for processing to obtain a scale feature map P4, and the scale feature map P4 is downsampled and concatenated with the feature map A17 to obtain a feature map A26, and the feature map A26 is input to the seventh C2F module for processing to obtain a scale feature map P5.

[0011] Furthermore, the head network includes a SIEM module, a first detection head, a second detection head and a third detection head. The scale feature map P3, the scale feature map P4 and the scale feature map P5 are input into the SIEM module for processing to obtain a small-scale feature map. The small-scale feature map is input into the first detection head for processing, the scale feature map P4 is input into the second detection head for processing, and the scale feature map P5 is input into the third detection head for processing. The target position and target category are obtained according to the processing results of the first detection head, the second detection head and the third detection head.

[0012] Furthermore, the first detection head corresponds to the P3 detection layer, the second detection head corresponds to the P4 detection layer, and the third detection head corresponds to the P5 detection layer.

[0013] Further, the SIEM module includes a fourth downsampling layer, a fifth downsampling layer, a sixth CBS module, a seventh CBS module, a seventh Conv module, an attention mechanism, a third upsampling layer, an eighth Conv module, a ninth Conv module and a second MLP module, wherein the scale feature map P3 is input into the fourth downsampling layer for processing to obtain a feature map C1, the scale feature map P4 is input into the fifth downsampling layer for processing to obtain a feature map C2, the scale feature map P5 is input into the sixth CBS module for processing to obtain a feature map C3, the feature maps C1, C2 and C3 are concatenated to obtain a feature map C4, The feature map C4 is processed by the seventh CBS module, the seventh Conv module and the attention mechanism to obtain the attention matrix, and the attention matrix is ​​input into the third upsampling layer for upsampling operation to obtain the feature map C5, and the scale feature map P3 is input into the eighth Conv module for processing to obtain the feature map C6, and the feature map C5 is added to the feature map C6 to obtain the feature map C7, and the feature map C7 is input into the ninth Conv module for processing to obtain the feature map C8, and the feature map C8 is added to the scale feature map P3 to obtain the feature map C9, and the feature map C9 is input into the second MLP module for processing to obtain the small-scale feature map .

[0014] Furthermore, the network structure of the second MLP module is the same as that of the first MLP module. The first MLP module includes an eighth CBS module and a ninth CBS module. After the feature map B14 is input into the MLP module, it is processed by the eighth CBS module and the ninth CBS module to obtain the feature map D1. After adding the feature map D1 to the feature map B14, the feature map B15 is obtained. The channel expansion rate between the eighth CBS module and the ninth CBS module is set to 2.

[0015] Compared with the prior art, the invention can achieve the following beneficial effects: (1) The present invention creates the lightweight aerial multispectral target detection method. Considering the sufficiency of feature fusion and the semantic information and global receptive field of small-scale fusion features, the present invention designs a lightweight multispectral target detection model based on the YOLOv8 architecture, wherein the backbone network uses the C2F module as the basic unit to construct two independent networks, which are respectively used to extract the multi-scale feature map of visible light and the multi-scale feature map of infrared image, and adopts an efficient cross-attention fusion module to fuse the multimodal feature map of the same stage. The present invention introduces a feature scaling method for local information enhancement (setting the feature scaling factor of the downsampling operation in each efficient cross-attention fusion module), which effectively reduces the computational complexity and video memory occupancy, while reducing the information loss caused by the feature scaling process. After the feature aggregation network, the present invention also designs a SIEM module (small-scale information enhancement module) to fuse multi-scale feature maps and model long-distance dependencies, providing semantic information and global receptive field for small-scale branches.

[0016] (2) The present invention creates the lightweight aerial multispectral target detection method and proposes a new multimodal feature fusion module, namely, an efficient cross-attention fusion module. The module consists of a cross-attention mechanism and a feature scaling method for local information enhancement. The cross-attention mechanism can effectively extract, interact and fuse multimodal features, while the feature scaling method for local information enhancement not only reduces the negative impact caused by the downsampling-upsampling process, but also significantly reduces the computational complexity of the fusion module.

[0017] (3) In the lightweight aerial multispectral target detection method created by the present invention, the SIEM module (small-scale information enhancement module) is designed to fuse multi-scale feature maps and calculate the self-attention mechanism, aiming to enhance the semantic information and global perception ability of small-scale branches. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings constituting part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation on the present invention. In the drawings: Figure 1 A schematic diagram of a lightweight aviation multi-spectral target detection method according to an embodiment of the present invention; Figure 2 A schematic diagram of the network structure of the multi-spectral target detection model described in the embodiment of the present invention; Figure 3 A schematic diagram of the network structure of the first efficient cross-attention fusion module described in an embodiment of the present invention; Figure 4 A schematic diagram of the network structure of the SIEM module described in the embodiment of the present invention; Figure 5 A schematic diagram of the detection results of the present invention on the DroneVehicle dataset according to an embodiment of the present invention; Figure 6 A schematic diagram of the detection results of the present invention on the LLVIP data set described in the embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the invention more clear, the invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described here are only used to explain the invention and do not constitute a limitation of the invention.

[0020] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0021] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0022] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the invention can be understood according to specific circumstances.

[0023] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0024] like Figure 1As shown, the present invention proposes a lightweight aviation multi-spectral target detection method, which specifically includes the following steps: S1: Obtain visible light images and infrared images in the same shooting scene; S2: Construct a multi-spectral target detection network, which includes a multi-scale extraction network, a neck network, and a head network; the multi-scale extraction network is used to extract, interact, and fuse information from visible light images and infrared images, and the fused features are input into the head network through the neck network; S3: Input the visible light images and infrared images into the multi-spectral target detection network for training to obtain a multi-spectral target detection model; S4: Input the visible light image to be detected and the infrared image to be detected in the same shooting scene into the multi-spectral target detection model for detection to obtain the target position and target category.

[0025] It should be noted that the present invention first reads visible light images and infrared images, and the shooting scenes of the visible light images and infrared images are the same; the visible light images and infrared images are respectively input into two independent multi-scale feature extraction networks, and through an efficient cross-attention fusion module, information extraction, interaction, and fusion between different modality feature maps are realized; the multi-scale fusion feature maps are input into the neck network, and the semantic information and localization information contained in the visible light images and infrared images are respectively transmitted along the top-down and bottom-up paths; a decoupled detection head is used to predict the three-scale feature maps (P3, P4, P5) generated by the neck network, and the target position and target category are output.

[0026] In some embodiments, such as Figure 2As shown, the multi-scale extraction network includes a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a seventh feature extraction module, an eighth feature extraction module, a ninth feature extraction module, a tenth feature extraction module, a first efficient cross-attention fusion module, a second efficient cross-attention fusion module, a third efficient cross-attention fusion module, a fourth efficient cross-attention fusion module and an SPPF module; wherein the infrared image is processed by the first feature extraction module and the second feature extraction module to obtain a feature map A1; the visible light image is processed by the third feature extraction module and the fourth feature extraction module to obtain a feature map A2; the feature map A1 and the feature map A2 are input into the first efficient cross-attention fusion module for processing, and feature maps A3 and A4 are obtained correspondingly; the feature map A3 is input into the fifth feature extraction module for processing to obtain a feature map A1. Feature map A5, feature map A4 is input to the sixth feature extraction module for processing to obtain feature map A6, feature map A5 and feature map A6 are input to the second efficient cross attention fusion module for processing, and feature map A7 and feature map A8 are obtained accordingly; feature map A7 is input to the seventh feature extraction module for processing to obtain feature map A9, feature map A8 is input to the eighth feature extraction module for processing to obtain feature map A10, feature map A9 and feature map A10 are input to the third efficient cross attention fusion module for processing, and feature map A11 and feature map A12 are obtained accordingly; feature map A11 is input to the ninth feature extraction module for processing to obtain feature map A13, feature map A12 is input to the tenth feature extraction module for processing to obtain feature map A14, feature map A13 and feature map A14 are input to the fourth efficient cross attention fusion module for processing, and feature map A15 and feature map A16 are obtained accordingly; Add feature map A3 and feature map A4 to obtain the first fusion feature; add feature map A7 and feature map A8 to obtain the second fusion feature; add feature map A11 and feature map A12 to obtain the third fusion feature; add feature map A15 and feature map A16 to obtain feature map A17; input feature map A17 into the SPPF module for processing to obtain the fourth fusion feature.

[0027] It should be noted that the visible light image and the infrared image are processed by two independent feature extraction modules respectively, and then processed by the first efficient cross-attention fusion module at the same time to obtain the first visible light feature (feature map A3), the first infrared feature (feature map A4) and the first fusion feature. The first visible light feature and the first infrared feature are processed by independent feature extraction modules respectively, and processed by the second efficient cross-attention fusion module at the same time to obtain the second visible light feature (feature map A7), the second infrared feature (feature map A8) and the second fusion feature. The second visible light feature and the second infrared feature are processed by independent feature extraction modules respectively, and processed by the third efficient cross-attention fusion module at the same time to obtain the third visible light feature (feature map A11), the third infrared feature (feature map A12) and the third fusion feature. The third visible light feature and the third infrared feature are processed by independent feature extraction modules respectively, and processed by the fourth efficient cross-attention fusion module and the SPPF module at the same time to obtain the fourth fusion feature. The first fusion feature, the second fusion feature, the third fusion feature and the fourth fusion feature are all input into the neck network. The SPPF module is a commonly used structure in the YOLO series of models, which is used to extract multi-scale information in the feature map.

[0028] In some embodiments, the first feature extraction module and the third feature extraction module have the same network structure, and the first feature extraction module includes a first CBS module and a second CBS module having the same network structure; The network structures of the second feature extraction module, the fourth feature extraction module, the fifth feature extraction module, the sixth feature extraction module, the seventh feature extraction module, the eighth feature extraction module, the ninth feature extraction module and the tenth feature extraction module are the same, and the second feature extraction module includes a third CBS module and a first C2F module connected in sequence; It should be noted that the first CBS module, the second CBS module and the third CBS module all include a 2D convolution layer, a batch normalization layer and a SiLU activation function connected in sequence, and the convolution kernel of the 2D convolution layer is 3×3 and the step size is 2.

[0029] In some embodiments, the network structures of the first efficient cross-attention fusion module, the second efficient cross-attention fusion module, the third efficient cross-attention fusion module and the fourth efficient cross-attention fusion module are the same, and the first efficient cross-attention fusion module includes a fourth CBS module, a fifth CBS module, a first downsampling layer, a second downsampling layer, a third downsampling layer, a first Conv module, a second Conv module, a third Conv module, a fourth Conv module, a fifth Conv module, a sixth Conv module, a visible light cross-attention mechanism, an infrared cross-attention mechanism, a first upsampling layer, a second upsampling layer, and a first MLP module; wherein, the feature map A1 is processed by the first downsampling layer and the first Conv module to obtain the feature map B1; the feature map A2 is processed by the second downsampling layer and the second Conv module to obtain the feature map B2; the feature map A1 and the feature map A2 are spliced ​​to obtain the feature map B3, and the feature map B3 is processed by the fourth CBS module, the third downsampling layer and the third Conv module Processing is performed to obtain feature map B4; feature map B4 is split to obtain feature map B5 and feature map B6; feature map B1 and feature map B5 are processed by the visible light cross attention mechanism and the first upsampling layer to obtain a visible light attention matrix; feature map B2 and feature map B6 are processed by the infrared cross attention mechanism and the second upsampling layer to obtain an infrared attention matrix; feature map A1 is input into the fourth Conv module for processing to obtain feature map B7; feature map A2 is input into the fifth Conv module for processing to obtain feature map B8; feature map B7 and the visible light attention matrix are added to obtain feature map B9; feature map B8 and the infrared attention matrix are added to obtain feature map B10; feature map B9 and feature map B10 are spliced ​​to obtain feature map B11; feature map B11 is input into the sixth Conv module for processing to obtain feature map B12; feature map A1 and feature map A2 are spliced ​​and processed by CBS in the channel dimension to obtain feature map F f , the feature map F f Input the feature map B12 to the fifth CBS module for processing to obtain feature map B13, add the feature map B12 and the feature map B13 to obtain feature map B14, input the feature map B14 to the first MLP module for processing to obtain feature map B15, split the feature map B15 into feature map B17 and feature map B18, and add the feature map B17 to the feature map A1 to obtain feature map A3, add the feature map B18 to the feature map A2 to obtain feature map A4, add the feature map A3 to the feature map A4 to obtain the first fusion feature.

[0030] It should be noted that the convolution kernels of the first Conv module, the second Conv module, the third Conv module and the sixth Conv module are 1×1, and the convolution kernels of the fourth Conv module and the fifth Conv module are 3×3.

[0031] Furthermore, the visible light features and infrared features are processed by the downsampling layer and the Conv module respectively to obtain the visible light downsampling features (such as feature map B1) and the infrared downsampling features (such as feature map B2). At the same time, the visible light features and the infrared features are spliced ​​and then passed through the CBS module to obtain the primary fusion feature map (feature map F). The primary fusion feature map is subjected to the downsampling layer, the Conv module and the split operation to obtain the primary fusion visible light downsampling features (feature map B5) and the primary fusion infrared downsampling features (feature map B6). The visible light downsampling features and the primary fusion visible light downsampling features are subjected to the visible light cross attention mechanism and the upsampling layer, and the infrared downsampling features and the primary fusion infrared downsampling features are subjected to the infrared cross attention mechanism and the upsampling layer to obtain the visible light attention matrix and the infrared attention matrix respectively. The sum of the visible light features after the convolution operation and the visible light attention matrix is ​​concatenated and convolved with the sum of the infrared features after the convolution operation and the added attention matrix, and then added to the primary fused feature map after the downsampling layer and the convolution operation. After passing through the first MLP module, it is split into the visible light feature map and infrared feature map of this stage, and the split visible light feature map and infrared feature map are added to the input visible light feature map and infrared feature map respectively, and the visible light features and infrared features of this layer are output. The visible light features and infrared features of this layer are added, and the fused features of this layer are output.

[0032] In some embodiments, the network structures of the first CBS module, the second CBS module, the third CBS module, the fourth CBS module and the fifth CBS module are the same, the first CBS module includes a 2d convolution layer, a batch normalization layer and a SiLU activation function connected in sequence, and the convolution kernel of the 2d convolution layer is 3×3 and the step size is 2; The network structures of the first downsampling layer, the second downsampling layer and the third downsampling layer are the same, wherein the first downsampling layer includes a maximum pooling layer and an average pooling layer. After the feature map A1 is processed by the maximum pooling layer and the average pooling layer respectively, the processing results of the two are multiplied by their respective assigned weights and then added, and the added result is input into the first Conv module for processing. The weight assigned to the branch where the maximum pooling layer is located is 0.5, and the weight assigned to the branch where the average pooling layer is located is 0.5; The feature scaling factor of the downsampling operation in the first efficient cross-attention fusion module is set to 8, the feature scaling factor of the downsampling operation in the second efficient cross-attention fusion module is set to 4, the feature scaling factor of the downsampling operation in the third efficient cross-attention fusion module is set to 2, and the feature scaling factor of the downsampling operation in the fourth efficient cross-attention fusion module is set to 1.

[0033] In some embodiments, the neck network includes a second C2F module, a third C2F module, a fourth C2F module, a fifth C2F module, a sixth C2F module and a seventh C2F module. The fourth fusion feature is upsampled and concatenated with the third fusion feature to obtain a feature map A18. The feature map A18 is input to the second C2F module for processing to obtain a feature map A19. The feature map A19 is upsampled and concatenated with the second fusion feature to obtain a feature map A20. The feature map A20 is input to the third C2F module for processing to obtain a feature map A21. The feature map A21 is upsampled and concatenated with the first fusion feature to obtain a feature map A22. The feature map A22 is input to the fourth C2F module for processing to obtain the feature map A23; the feature map A23 is downsampled and concatenated with the feature map A21 to obtain the feature map A24, and the feature map A24 is input to the fifth C2F module for processing to obtain the scale feature map P3, and the scale feature map P3 is downsampled and concatenated with the feature map A19 to obtain the feature map A25, and the feature map A25 is input to the sixth C2F module for processing to obtain the scale feature map P4, and the scale feature map P4 is downsampled and concatenated with the feature map A17 to obtain the feature map A26, and the feature map A26 is input to the seventh C2F module for processing to obtain the scale feature map P5.

[0034] It should be noted that the first fusion feature, the second fusion feature, the third fusion feature, and the fourth fusion feature are respectively input into the neck network. The fourth fusion feature is spliced ​​with the third fusion feature through an upsampling operation, and then the first scale fusion feature (such as feature map A19) is obtained through the C2F layer. Then, it is spliced ​​with the second fusion feature through an upsampling operation, and the second scale fusion feature (such as feature map A21) is obtained through the C2F layer. Then, it is spliced ​​with the first fusion feature through an upsampling operation, and the second scale fusion feature is spliced ​​through the C2F layer and downsampling operation. After the C2F layer, the P3 scale feature map is obtained. After the downsampling operation, it is spliced ​​with the first scale fusion feature, and the P4 scale feature map is obtained through the C2F layer. After the downsampling operation, it is spliced ​​with the fourth fusion feature, and the P5 scale feature map is obtained through the C2F layer. The P3 scale feature map, the P4 scale feature map, and the P5 scale feature map are simultaneously input into the small-scale information enhancement module (SIEM module) to obtain the small-scale feature map.

[0035] like Figure 3 As shown, and They represent the visible light feature map and the infrared feature map respectively. By performing splicing and 3×3 convolution operations in the channel dimension, the information carried by the two modalities is adaptively integrated together to obtain a simple fusion feature map : ; Among them, CBS represents the processing function of the CBS module (convolution kernel is 3×3).

[0036] Due to the computational complexity and spatial dimension of self-attention The square of the length of is proportional to the three feature maps.

[0037] ; in, Represents the feature scaling factor, and Down represents the downsampling operation.

[0038] Specifically, the downsampling operation It consists of maximum pooling and average pooling: ; in, and is the regulating factor, Represents the input feature map, where and Both are set to 0.5.

[0039] The calculation process of cross attention is as follows: the downsampled feature map is respectively calculated through three 1×1 convolutions to calculate the query matrix of visible light , infrared interrogation matrix , the fused key-value matrix and , .

[0040] ; Then, adjust , , and , and calculate and Relative to Visible light attention matrix and infrared attention matrix : ; ; Among them, the symbol represents matrix multiplication, is the number of channels of the feature map, and SoftMax is the SoftMax function.

[0041] In order to reduce the information loss in the feature scaling process, and Transform the size and upsample the two matrices using a 3×3 convolution from and Extract local information from and Combined, , .

[0042] ; In obtaining the feature maps of the two modalities relative to the fusion feature map After the information is obtained, it is spliced ​​in the channel dimension and , and use 1x1 convolution for linear projection. In addition, we introduce The residual path of is used to alleviate the model degradation problem, and the 1×1 convolution is used to The number of channels can be adjusted. .

[0043] ; in, is the channel dimension concatenation operation, It is a 1x1 convolutional layer.

[0044] In the feedforward network, the MLP module consists of two 1×1 convolutions and residual connections to achieve information interaction between different modalities. Considering the amount of computation, the channel expansion rate between the two convolutions is set to 2:

[0045] Split along the channel dimension , and add them to the original two modal feature maps to obtain the visible feature map for feature extraction in the next stage and infrared signature Finally, add these two feature maps to get the fusion feature map of this stage , .

[0046] ; ; in, To evenly divide the feature map in the channel dimension.

[0047] In some embodiments, the head network includes a SIEM module, a first detection head, a second detection head and a third detection head. The scale feature map P3, the scale feature map P4 and the scale feature map P5 are input into the SIEM module for processing to obtain a small-scale feature map. The small-scale feature map is input into the first detection head for processing, the scale feature map P4 is input into the second detection head for processing, and the scale feature map P5 is input into the third detection head for processing. The target position and target category are obtained according to the processing results of the first detection head, the second detection head and the third detection head.

[0048] In some embodiments, the first detection head corresponds to the P3 detection layer, the second detection head corresponds to the P4 detection layer, and the third detection head corresponds to the P5 detection layer.

[0049] It should be noted that the three detection heads are arranged in parallel and correspond to the output feature maps of each scale. The detection heads are used to obtain prediction information at different scales from the multi-scale fusion features, generate target categories and target coordinate positions, and mark target attributes and areas in infrared images and visible light images.

[0050] In some embodiments, the SIEM module includes a fourth downsampling layer, a fifth downsampling layer, a sixth CBS module, a seventh CBS module, a seventh Conv module, an attention mechanism, a third upsampling layer, an eighth Conv module, a ninth Conv module and a second MLP module, wherein the scale feature map P3 is input to the fourth downsampling layer for processing to obtain a feature map C1, the scale feature map P4 is input to the fifth downsampling layer for processing to obtain a feature map C2, the scale feature map P5 is input to the sixth CBS module for processing to obtain a feature map C3, the feature maps C1, C2 and C3 are spliced ​​to obtain a feature map C4 , the feature map C4 is processed by the seventh CBS module, the seventh Conv module and the attention mechanism to obtain the attention matrix, the attention matrix is ​​input to the third upsampling layer for upsampling operation, and the feature map C5 is obtained. The scale feature map P3 is input to the eighth Conv module for processing to obtain the feature map C6, the feature map C5 is added to the feature map C6 to obtain the feature map C7, the feature map C7 is input to the ninth Conv module for processing to obtain the feature map C8, the feature map C8 is added to the scale feature map P3 to obtain the feature map C9, the feature map C9 is input to the second MLP module for processing to obtain the small-scale feature map .

[0051] In the SIEM module (small-scale information enhancement module), the input scale feature map P3 and scale feature map P4 are downsampled with feature scaling factors of 4 and 2 respectively, and the scale feature map P5 is input to the sixth CBS module for processing. After the processing results of the three are spliced, the number of channels is adjusted through the seventh CBS module and the seventh Conv module, and then the scaled dot product attention is calculated to obtain the attention matrix. The attention matrix is ​​upsampled and added to the P3 scale feature map processed by the eighth Conv module. After being processed by the ninth Conv module, it is added to the P3 scale feature map, and linear projection and feedforward network calculation are performed through the second MLP module to obtain the small-scale feature map. .

[0052] Furthermore, the seventh Conv module is a 1×1 convolution, and the eighth Conv module and the ninth Conv module are 3×3 convolutions.

[0053] In some embodiments, the network structure of the second MLP module is the same as that of the first MLP module, the first MLP module includes an eighth CBS module and a ninth CBS module, and after the feature map B14 is input into the MLP module, it is processed by the eighth CBS module and the ninth CBS module to obtain the feature map D1, and after adding the feature map D1 to the feature map B14, the feature map B15 is obtained; the channel expansion rate between the eighth CBS module and the ninth CBS module is set to 2.

[0054] It should be noted that the MLP module consists of two 1×1 convolutional layers and residual connections, and the channel expansion rate between the two convolutional layers is set to 2. The feature scaling factors of the downsampling operations (the downsampling operations consist of the maximum pooling layer and the average pooling layer) in these four efficient cross-attention fusion modules are set to 8, 4, 2, and 1, respectively, so that the size of the downsampled feature map at each stage is controlled to 1 / 32 of the input image. The network structure of the eighth CBS module and the ninth CBS module is the same, both of which include sequentially connected convolutional layers, batch normalization layers, and SiLU activation functions. The convolution kernels of the convolutional layers contained in the eighth CBS module and the ninth CBS module are both 1×1.

[0055] like Figure 4 As shown, the multi-scale feature map after selecting the path aggregation network , and As the input of the SIEM module, a similar feature scaling method is used to reduce the computational complexity. In order to fuse information of different scales, and The downsampling operation with feature scaling factors of 4 and 2 was applied respectively, and A 1x1 convolution is performed: ; , and By splicing in the channel dimension and adjusting the number of channels using 1×1 convolution, a fused feature map is obtained. .

[0056] ; Then, we calculate the scaled dot product attention of DP and get the attention matrix :

[0057]

[0058] In order to reduce the information loss caused by the downsampling-upsampling process, is upsampled in advance and added with the small-scale local information extracted by 3×3 convolution: ; in, is an upsampling operation.

[0059] Finally, Perform linear projection and feedforward network calculations to obtain , used to replace the original Serves as the input to the small-scale branch of the detection head.

[0060] ; .

[0061] Example 1 The present invention proposes a lightweight aviation multispectral target detection method, which specifically includes the following steps: Step 1: Input a visible light image and an infrared image with a size of H×W respectively.

[0062] In step 2, the visible light image and infrared image are respectively input into two independent multi-scale feature extraction networks, and the information extraction, interaction and fusion between feature maps of different modalities are realized through an efficient cross-attention fusion module.

[0063] The present invention is based on the YOLOv8 framework for optimization and improvement. The entire multispectral target detection network is optimized by the stochastic gradient descent algorithm, with a learning rate of 0.01, a momentum of 0.937, a weight decay of 0.0005, an epoch of 100, and an early stopping mechanism. The pre-trained weight method is not used. Each image is horizontally flipped with a probability of 0.5, HSV transformation, and Mosaic enhancement to increase diversity. The same data enhancement is performed on visible light images and infrared images. Therefore, the image pairs sent to the multispectral target detection network for enhancement are still aligned. The DroneVehicle dataset and the LLVIP dataset are selected for training and testing on two NVIDIA RTX3090 GPUs, and the batchsize is set to 16.

[0064] In step 3, the multi-scale fusion feature map is input into the neck network to transfer semantic information and positioning information along the top-down and bottom-up paths respectively.

[0065] The top-down path refers to the branch where the upsampling process is located, which gradually transfers the semantic information of the high-level feature map to the low-level feature map; the bottom-up path refers to the branch where the downsampling process is located, which gradually transfers the positioning information of the low-level feature map to the high-level feature map.

[0066] Step 4: Use the decoupled detection head to predict the three-scale feature maps (P3, P4, P5) generated by the neck network and output the target location and target category.

[0067] In order to verify the effectiveness of the present invention, the detection results of the present invention are compared with those of the existing methods. For the evaluation of target detection, some common indicators are introduced, such as selecting accuracy, recall rate, mAP50 and mAP as the evaluation criteria of model performance, and the detection speed of the model is measured by the model parameters Paramrs, mAP50 and mAP. The effectiveness and superiority of the algorithm are verified from multiple aspects such as detection accuracy, model size and real-time performance.

[0068] The experimental results of different methods on the DroneVehicle dataset are shown in Table 1. The present invention achieves 83.3% mAP50 and 63.4% mAP respectively. With fewer parameters, it greatly surpasses the two simple fusion methods YOLOv8n-Add and YOLOv8n-Cat. Compared with the deeper and wider YOLOv8s-Add, it improves by 1.0mAP50 and 0.8mAP respectively, proving that YOLO-FT (the multispectral target detection model designed by the present invention) can effectively fuse features of different modalities under the guidance of global information. Compared with the ICAFusion method, the multispectral target detection model of the present invention improves by 0.5 and 2.5 in mAP50 and mAP respectively, while reducing the number of parameters by 12.95M. Figure 5 This is a result diagram of detection using the present invention on the DroneVehicle dataset. It can be found that YOLO-FT has fewer missed detections and false detections.

[0069] Table 1

[0070] The experimental results of different methods on the LLVIP dataset are shown in Table 2. The present invention achieves 97.9% mAP50 and 68.1% mAP, respectively, exceeding the 0.9mAP50 and 1.0mAP of YOLOv8n in the infrared modality. The present invention has the same mAP50 as LRAF-Net, but the mAP is improved by 1.8. The present invention obtains the best detection accuracy with the least number of parameters, indicating that the predicted box of the multispectral target detection model has a higher degree of overlap with the real box. Figure 6 As a result of using the present invention to perform detection on the LLVIP dataset, it can be found that YOLO-FT has fewer missed detections and false detections.

[0071] Table 2

[0072] In summary, the present invention improves the YOLOv8 network architecture and proposes a lightweight multispectral target detection model, which is mainly composed of an efficient cross-attention module and a small-scale information enhancement module. The cross-attention mechanism fully extracts, interacts and fuses multimodal features by calculating the attention of different modal feature maps to simple fusion feature maps. The feature scaling method of local information enhancement reduces the computational complexity of the efficient cross-attention module and reduces the negative impact caused by the downsampling-upsampling process. The SIEM module captures the long-range dependencies of multi-scale fusion feature maps, provides rich semantic information and global perception capabilities for small-scale branches, and can effectively improve the all-weather, multi-scene, and high-precision target detection capabilities of aerial remote sensing.

[0073] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the disclosure of the present invention can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.

[0074] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A lightweight aviation multispectral target detection method, characterized by: The specific steps include: S1: Acquire visible light images and infrared images of the same shooting scene; S2: construct a multispectral target detection network, wherein the multispectral target detection network includes a multi-scale extraction network, a neck network and a head network; The multi-scale extraction network is used to extract, interact and fuse information from visible light images and infrared images, and the fused features are input into the head network via the neck network; S3: Input the visible light image and the infrared image into the multispectral target detection network for training to obtain a multispectral target detection model; S4: Input the visible light image to be detected and the infrared image to be detected in the same shooting scene into the multispectral target detection model for detection to obtain the target position and target category.

2. The lightweight aviation multispectral target detection method according to claim 1, characterized in that: The multi-scale extraction network includes a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a seventh feature extraction module, an eighth feature extraction module, a ninth feature extraction module, a tenth feature extraction module, a first efficient cross-attention fusion module, a second efficient cross-attention fusion module, a third efficient cross-attention fusion module, a fourth efficient cross-attention fusion module and an SPPF module; wherein the infrared image is processed by the first feature extraction module and the second feature extraction module to obtain a feature map A1; the visible light image is processed by the third feature extraction module and the fourth feature extraction module to obtain a feature map A2; the feature map A1 and the feature map A2 are input into the first efficient cross-attention fusion module for processing, and the feature map A3 and the feature map A4 are obtained correspondingly; the feature map A3 is input into the fifth feature extraction module for processing to obtain a feature map A5, input feature map A4 to the sixth feature extraction module for processing to obtain feature map A6, input feature map A5 and feature map A6 to the second efficient cross-attention fusion module for processing, and obtain feature map A7 and feature map A8 accordingly; input feature map A7 to the seventh feature extraction module for processing to obtain feature map A9, input feature map A8 to the eighth feature extraction module for processing to obtain feature map A10, input feature map A9 and feature map A10 to the third efficient cross-attention fusion module for processing, and obtain feature map A11 and feature map A12 accordingly; input feature map A11 to the ninth feature extraction module for processing to obtain feature map A13, input feature map A12 to the tenth feature extraction module for processing to obtain feature map A14, input feature map A13 and feature map A14 to the fourth efficient cross-attention fusion module for processing, and obtain feature map A15 and feature map A16 accordingly; Add feature map A3 and feature map A4 to obtain the first fusion feature; add feature map A7 and feature map A8 to obtain the second fusion feature; add feature map A11 and feature map A12 to obtain the third fusion feature; add feature map A15 and feature map A16 to obtain feature map A17; input feature map A17 into the SPPF module for processing to obtain the fourth fusion feature.

3. The lightweight aviation multispectral target detection method according to claim 2, characterized in that: The first feature extraction module and the third feature extraction module have the same network structure, and the first feature extraction module includes a first CBS module and a second CBS module having the same network structure; The network structures of the second feature extraction module, the fourth feature extraction module, the fifth feature extraction module, the sixth feature extraction module, the seventh feature extraction module, the eighth feature extraction module, the ninth feature extraction module and the tenth feature extraction module are the same, and the second feature extraction module includes a third CBS module and a first C2F module connected in sequence.

4. The lightweight aviation multispectral target detection method according to claim 3 is characterized in that: The network structures of the first efficient cross-attention fusion module, the second efficient cross-attention fusion module, the third efficient cross-attention fusion module and the fourth efficient cross-attention fusion module are the same. The first efficient cross-attention fusion module includes a fourth CBS module, a fifth CBS module, a first downsampling layer, a second downsampling layer, a third downsampling layer, a first Conv module, a second Conv module, a third Conv module, a fourth Conv module, a fifth Conv module, a sixth Conv module, a visible light cross-attention mechanism, an infrared cross-attention mechanism, a first upsampling layer, a second upsampling layer, and a first MLP module; wherein, the feature map A1 is processed by the first downsampling layer and the first Conv module to obtain the feature map B1; the feature map A2 is processed by the second downsampling layer and the second Conv module to obtain the feature map B2; the feature map A1 and the feature map A2 are spliced ​​to obtain the feature map B3, and the feature map B3 is processed by the fourth CBS module, the third downsampling layer and the third Conv module. Processing to obtain feature map B4; splitting feature map B4 to obtain feature map B5 and feature map B6; processing feature map B1 and feature map B5 through the visible light cross attention mechanism and the first upsampling layer to obtain a visible light attention matrix; processing feature map B2 and feature map B6 through the infrared cross attention mechanism and the second upsampling layer to obtain an infrared attention matrix; inputting feature map A1 into the fourth Conv module for processing to obtain feature map B7; inputting feature map A2 into the fifth Conv module for processing to obtain feature map B8; adding feature map B7 and the visible light attention matrix to obtain feature map B9; adding feature map B8 and the infrared attention matrix to obtain feature map B10; concatenating feature map B9 and feature map B10 to obtain feature map B11; inputting feature map B11 into the sixth Conv module for processing to obtain feature map B12; concatenating feature map A1 and feature map A2 in the channel dimension and performing CBS processing to obtain feature map F f , the feature map F f Input the feature map B12 to the fifth CBS module for processing to obtain feature map B13, add the feature map B12 and the feature map B13 to obtain feature map B14, input the feature map B14 to the first MLP module for processing to obtain feature map B15, split the feature map B15 into feature map B17 and feature map B18, and add the feature map B17 to the feature map A1 to obtain feature map A3, add the feature map B18 to the feature map A2 to obtain feature map A4, add the feature map A3 to the feature map A4 to obtain the first fusion feature.

5. The lightweight aviation multispectral target detection method according to claim 4, characterized in that: The network structures of the first CBS module, the second CBS module, the third CBS module, the fourth CBS module and the fifth CBS module are the same. The first CBS module includes a 2d convolution layer, a batch normalization layer and a SiLU activation function connected in sequence, and the convolution kernel of the 2d convolution layer is 3×3 and the step size is 2; The network structures of the first downsampling layer, the second downsampling layer and the third downsampling layer are the same, wherein the first downsampling layer includes a maximum pooling layer and an average pooling layer, and after the feature map A1 is processed by the maximum pooling layer and the average pooling layer, the processing results of the two are multiplied by their respective assigned weights and then added, and the added result is input into the first Conv module for processing, and the weight assigned to the branch where the maximum pooling layer is located is 0.5, and the weight assigned to the branch where the average pooling layer is located is 0.5; The feature scaling factor of the downsampling operation in the first efficient cross-attention fusion module is set to 8, the feature scaling factor of the downsampling operation in the second efficient cross-attention fusion module is set to 4, the feature scaling factor of the downsampling operation in the third efficient cross-attention fusion module is set to 2, and the feature scaling factor of the downsampling operation in the fourth efficient cross-attention fusion module is set to 1.

6. The lightweight aviation multispectral target detection method according to claim 4, characterized in that: The neck network includes a second C2F module, a third C2F module, a fourth C2F module, a fifth C2F module, a sixth C2F module and a seventh C2F module. The fourth fusion feature is upsampled and concatenated with the third fusion feature to obtain a feature map A18. The feature map A18 is input to the second C2F module for processing to obtain a feature map A19. The feature map A19 is upsampled and concatenated with the second fusion feature to obtain a feature map A20. The feature map A20 is input to the third C2F module for processing to obtain a feature map A21. The feature map A21 is upsampled and concatenated with the first fusion feature to obtain a feature map A22. 22 is input to the fourth C2F module for processing to obtain a feature map A23; the feature map A23 is downsampled and concatenated with the feature map A21 to obtain a feature map A24, the feature map A24 is input to the fifth C2F module for processing to obtain a scale feature map P3, the scale feature map P3 is downsampled and concatenated with the feature map A19 to obtain a feature map A25, the feature map A25 is input to the sixth C2F module for processing to obtain a scale feature map P4, the scale feature map P4 is downsampled and concatenated with the feature map A17 to obtain a feature map A26, the feature map A26 is input to the seventh C2F module for processing to obtain a scale feature map P5.

7. The lightweight aviation multispectral target detection method according to claim 6, characterized in that: The head network includes a SIEM module, a first detection head, a second detection head and a third detection head. The scale feature map P3, the scale feature map P4 and the scale feature map P5 are input into the SIEM module for processing to obtain a small-scale feature map. The small-scale feature map is input into the first detection head for processing, the scale feature map P4 is input into the second detection head for processing, and the scale feature map P5 is input into the third detection head for processing. The target position and target category are obtained according to the processing results of the first detection head, the second detection head and the third detection head.

8. The lightweight aviation multispectral target detection method according to claim 7, characterized in that: The first detection head corresponds to the P3 detection layer, the second detection head corresponds to the P4 detection layer, and the third detection head corresponds to the P5 detection layer.

9. The lightweight aviation multispectral target detection method according to claim 7, characterized in that: The SIEM module includes a fourth downsampling layer, a fifth downsampling layer, a sixth CBS module, a seventh CBS module, a seventh Conv module, an attention mechanism, a third upsampling layer, an eighth Conv module, a ninth Conv module and a second MLP module, wherein the scale feature map P3 is input into the fourth downsampling layer for processing to obtain a feature map C1, the scale feature map P4 is input into the fifth downsampling layer for processing to obtain a feature map C2, the scale feature map P5 is input into the sixth CBS module for processing to obtain a feature map C3, the feature maps C1, C2 and C3 are spliced ​​to obtain a feature map C4, and the feature map C4 is processed by the seventh CBS module, the seventh Conv module and the attention mechanism to obtain an attention matrix, and the attention matrix is ​​input into the third upsampling layer for upsampling to obtain a feature map C5. The scale feature map P3 is input into the eighth Conv module for processing to obtain a feature map C6. The feature map C5 is added to the feature map C6 to obtain a feature map C7. The feature map C7 is input into the ninth Conv module for processing to obtain a feature map C8. The feature map C8 is added to the scale feature map P3 to obtain a feature map C9. The feature map C9 is input into the second MLP module for processing to obtain a small-scale feature map. .

10. The lightweight aviation multi-spectral target detection method according to claim 9, characterized in that: The network structure of the second MLP module is the same as that of the first MLP module. The first MLP module includes an eighth CBS module and a ninth CBS module. After the feature map B14 is input into the MLP module, it is processed by the eighth CBS module and the ninth CBS module to obtain the feature map D1. After adding the feature map D1 to the feature map B14, the feature map B15 is obtained. The channel expansion rate between the eighth CBS module and the ninth CBS module is set to 2.

Citation Information

Patent Citations

  • Hyperspectral and laser radar data fusion classification method based on AGLT network

    CN117475216A

  • Substation power equipment detection method and system based on infrared and visible light fusion

    CN117557775A

  • Unmanned aerial vehicle small target detection method based on multispectral interactive attention fusion

    CN117830878A

  • Multi-spectral feature fusion multi-scale target detection method based on universal attention mechanism

    CN119091122A

  • Multi-modal space-time fusion target detection method and device and medium

    CN119107638A

Cited By

  • Unmanned aerial vehicle road surface type identification method and device based on dual-spectrum cross attention

    CN121982593A