A Pixel-Level Multispectral Fusion Small Target Detection Method and System Based on Mask Enhancement
Through the pixel-level multi-spectral fusion method based on mask enhancement, the problems of large amount of parameters and poor fusion effect in edge devices are solved, and high-quality multi-spectral small object detection is achieved, which is suitable for devices with resource limitations.
Patent Information
- Application Number
- CN202411785162.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing multispectral modal fusion method has problems such as large amount of parameters, redundant calculations and poor fusion effect when deploying edge devices, resulting in poor detection performance.
Using a pixel-level multispectral fusion method based on mask enhancement, the mask is generated by using two 3×3 convolutional layers and ReLU activation functions, the original feature map is directly processed, and the mask value is restricted from 0 to 1 using the Sigmoid function, combined with the SE module to filter effective information, suppress redundant features, and finally input into the lightweight detection network for object detection.
It improves the accuracy of small object detection, reduces the amount of model parameters, is suitable for edge equipment deployment, and improves the fusion quality and detection efficiency.
Smart Images

Figure CN119810601B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image target detection, and particularly relates to a pixel-level multi-spectral fusion small target detection method and system based on mask enhancement. Background Art
[0002] In recent years, small target detection technology has been widely applied in civilian and military fields. With the development of multi-spectral modality devices, it has become a trend to use multi-spectral modality fusion technology to improve the detection accuracy of small targets, such as applying visible light spectrum, infrared spectrum, ultraviolet, etc. The information complementarity between multiple modality data can improve the detection accuracy. However, the existing multi-spectral modality fusion has problems such as a large number of parameters, computational redundancy, and poor fusion effect, resulting in poor detection performance when the model is deployed at the edge.
[0003] Multi-modal algorithms can be divided into pixel-level fusion, feature-level fusion, and decision-level fusion according to the level of feature fusion. Feature-level fusion has more parameters compared to pixel-level fusion. For example Figure 1 is the block diagram of the feature-level fusion method of the prior art. Feature-level fusion refers to the fusion at the feature representation level where data is converted into feature descriptors. This fusion method has two feature extraction backbone networks and three head networks, and the total number of model parameters reaches 122 million, with a real-time frame rate of only 6 FPS, which does not meet the requirements for edge device deployment.
[0004] This limits its application in resource-constrained mobile devices or real-time systems. Edge deployment mainly considers lightweight pixel-level fusion methods. However, the existing pixel-level fusion methods have the problem of low fusion quality. When different modality data are fused, they need to be converted into a unified feature space or representation form, which will result in the loss of some original feature information. In addition, there is duplicate information in different modalities, which will increase the computational cost and interfere with the model's extraction and understanding of key information, reducing the fusion effect.
[0005] Facing many challenges, traditional fusion methods cannot meet the requirements of real-time target detection. Developing a fusion method with high fusion quality and lightweight deployment has become an important research field. Summary of the Invention
[0006] Aiming at the problems of poor fusion quality, high number of parameters, and inapplicability to edge device deployment in the existing fusion technology, the purpose of the present invention is to overcome the above technical defects and provide a pixel-level multi-spectral fusion small target detection method and system based on mask enhancement.
[0007] In view of this, the present invention proposes a pixel-level multi-spectral fusion small target detection method based on mask enhancement, including:
[0008] Step (1): Obtain the remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of visible light spectrum and infrared spectrum;
[0009] Step (2): Split the multi-spectral modal image pair and input them into the feature fusion module simultaneously, respectively obtain their respective masks, then obtain the feature information after mask processing according to their respective masks, and perform feature fusion with the original split images to obtain the first fusion result of single-modal features, and then perform channel splicing and merging to obtain the result after the second fusion;
[0010] Step (3): Input the result after the second fusion into the SE module, screen out the effective information features, suppress the redundant features, and obtain the image after pixel-level fusion;
[0011] Step (4): Input the image after pixel-level fusion into the lightweight detection network to obtain the detection result.
[0012] Preferably, the feature fusion module in step (2) includes two mask generation modules with the same structure. Among them, the input of the first mask generation module is the feature map x of the visible light spectrum rgb_ori , and the output is the mask x of the visible light spectrum mask_rgb , the input of the second mask generation module is the feature map x of the infrared spectrum ir_ori , and the output is the mask x of the infrared spectrum mask_ir , each of the mask generation modules includes: two convolutional layers with a convolutional kernel size of 3, a ReLU activation function, and a Sigmoid activation function, where
[0013] The first convolutional layer is used for preliminary feature extraction of the input data;
[0014] The ReLU activation function is used to introduce non-linearity so that the network can better express features;
[0015] The second convolutional layer is used to transform the previously extracted features;
[0016] The Sigmoid activation function is used to map the value of the mask to between 0 and 1.
[0017] Preferably, the feature fusion module in step (2) includes two multiplication layers with the same structure, which are used to screen the pixel elements in the original image according to the mask, where
[0018] The mask x of the visible light spectrum mask_rgb is multiplied element-wise with the feature map x of the visible light spectrum rgb_ori , and the mask x of the infrared spectrum mask_ir is multiplied element-wise with the feature map x of the infrared spectrum ir_oriPerform element-wise multiplication to obtain the corresponding masked feature information x rgb_masked and x ir_masked .
[0019] Preferably, the feature fusion module in step 2) includes a first fusion network, and the first fusion network includes: two addition layers with the same structure and a convolutional layer with a convolutional kernel size of 3; wherein,
[0020] The first addition layer is used to perform element-wise addition of x rgb_masked and x rgb_ori ; the first convolutional layer is used to perform convolution on the output of the first addition layer to obtain the first fusion result x out_rgb of the visible light modality features;
[0021] The second addition layer is used to perform element-wise addition of x ir_masked and x ir_ori ; the second convolutional layer is used to perform convolution on the output of the second addition layer to obtain the first fusion result x out_ir of the infrared light modality features.
[0022] Preferably, the feature fusion module in step 2) includes channel splicing and merging, which is used to splice x out_rgb and x out_ir along the channel with dimension 1 to obtain the second fusion result x out_cat .
[0023] Preferably, the SE module in step 3) includes a global average pooling layer and a gate mechanism composed of two fully connected layers; wherein,
[0024] The global average pooling layer is used to compress the global spatial information of each channel;
[0025] The gate mechanism is used to provide a specific weight for each channel, and this weight is used to reflect the importance of the channel, screen out redundant information through weighting, and strengthen effective information to obtain the pixel-level fused image I f .
[0026] Preferably, the lightweight detection network in step 4) is the YOLO v 8n lightweight detection network, and the input is the fused image I f , and the output is an identification result map with confidence and class labels.
[0027] Preferably, the lightweight detection network includes a backbone network and a head network.
[0028] On the other hand, the present invention provides a pixel-level multi-spectral fusion small target detection system based on mask enhancement, including:
[0029] An acquisition module, configured to acquire remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of a visible light spectrum and an infrared spectrum;
[0030] A secondary fusion module, configured to split the multi-spectral modal image pair and input it into a feature fusion module simultaneously, respectively obtain their respective masks, then obtain the masked feature information according to their respective masks, and perform feature fusion with the original image after splitting to obtain the first fusion result of a single modal feature, and then perform channel splicing and merging to obtain the result after the second fusion;
[0031] A screening and suppression module, configured to input the result after the second fusion into an SE module, screen out effective information features, suppress redundant features, and obtain an image after pixel-level fusion;
[0032] A detection and output module, configured to input the image after pixel-level fusion into a lightweight detection network to obtain a detection result.
[0033] Compared with the prior art, the advantages of the present invention are as follows:
[0034] 1. Traditional pixel-level fusion methods use 1×1 convolutions to generate masks. Simple 1×1 convolutions are not sufficient to capture complex spatial relationships and will lose some original spatial structure information, thus affecting the quality of the masks. In contrast, our method uses two 3×3 convolutional layers and a ReLU activation function in the mask generation module of the feature fusion module to generate masks, which can learn richer spatial features.
[0035] 2. In terms of the mask application method, traditional methods multiply the original map features by 0.5 respectively and then apply the generated masks. This scaling may affect the expression ability of the features and the accuracy of subsequent processing. In contrast, our method directly processes the original feature map without additional scaling steps, maintaining the original information of the features. In addition, our method uses the Sigmoid function in the feature fusion module to limit the mask values between 0 and 1, which helps to more finely control the importance of the features.
[0036] 3. There is some duplicate or similar information between different modalities. For example, background information in the image, repetitive textures, etc. If the data of these two modalities are simply spliced or fused, a large amount of redundant information will be introduced. These redundant information will not only increase the computational cost, but may also interfere with the model's extraction and understanding of key information, reducing the fusion effect. Our method performs feature processing on each modality respectively, screens out effective features, completes the first fusion, and uses the SE module to screen the feature results after the second fusion again, so as to obtain a high-quality fused image. Description of the Drawings
[0037] Figure 1 It is a block diagram of an existing feature-level fusion method with a huge number of parameters;
[0038] Figure 2 It is a block diagram of the pixel-level multi-spectral fusion small target detection method based on mask enhancement of the present invention;
[0039] Figure 3 It is a block diagram of a traditional pixel-level fusion method;
[0040] Figure 4 It is a block diagram of the pixel-level multi-spectral fusion method based on mask enhancement of the present invention;
[0041] Figure 5 It is a visualization comparison diagram of the fusion effect between the present invention and the traditional pixel-level fusion method. Specific implementation manners
[0042] The method of the present invention includes:
[0043] Obtain the remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of visible light spectrum and infrared spectrum;
[0044] Split the multi-spectral modal image pair and input them into the feature fusion module simultaneously, respectively obtain their respective masks, then obtain the feature information after mask processing according to their respective masks, and perform feature fusion with the original split images to obtain the first fusion result of single-modal features, and then perform channel splicing and merging to obtain the second fusion result;
[0045] Input the second fusion result into the SE module, screen out the effective information features, suppress the redundant features, and obtain the pixel-level fused image;
[0046] Input the pixel-level fused image into the lightweight detection network to obtain the detection result.
[0047] The technical solution of the present invention will be described in detail below with reference to the drawings and embodiments.
[0048] Embodiment 1
[0049] Refer to Figure 2 , Embodiment 1 of the present invention proposes a pixel-level multi-spectral fusion small target detection method based on mask enhancement, and its specific steps are as follows:
[0050] Step (1): Read the multi-spectral modal image pair;
[0051] Among them, the multi-spectral modality in step (1) refers to infrared and visible light spectra. The remote sensing images are collected by drones, including various shooting conditions and scenes, covering different perspectives, heights and lighting conditions.
[0052] Step (2): Split the image pairs and send them into the mask generation module respectively;
[0053] Among them, the mask generation module in step (2) uses two convolutional layers with a convolutional kernel size of 3, a ReLU activation function, and a Sigmoid activation function to generate masks. First, the first convolutional layer is used to perform preliminary feature extraction on the input data; secondly, the ReLU activation function is used to introduce non-linear characteristics, enabling the network to better express features; then, the second convolutional layer is used to further transform the previously extracted features; finally, the Sigmoid activation function is used to map the values of the masks to between 0 and 1 to more finely control the importance of features. The original feature map x rgb_ori and x ir_ori After being processed by the mask generation module, the generated masks x mask_rgb and x mask_ir .
[0054] x mask_rgb =σ Sig (Conv 3×3 (σ ReLU (Conv 3×3 (x rgb_ori ))))
[0055] x mask_ir =σ Sig (Conv 3×3 (σ ReLU (Conv 3×3 (x ir_ori ))))
[0056] Among them, σ Sig is the Sigmoid activation function; σ ReLU is the ReLU activation function; Conv 3×3 is the convolutional layer with a convolutional kernel size of 3.
[0057] Step (3): Apply the generated masks to the images input to S2 to obtain the masked feature information;
[0058] Among them, the specific implementation of mask application in step (3) is to multiply the generated mask x mask_rgb element-wise with the original feature x rgb_ori to filter the pixel elements in the original feature x rgb_ori according to the mask values. After being processed by this module, x rgb_masked is obtained. Similarly, x ir_masked can be obtained.
[0059] Step (4): Perform feature fusion on the masked feature information and the original information input in S2, and obtain the first fusion result of the single-modal feature through the feature fusion module;
[0060] Among them, the specific implementation of the feature fusion module in step (4) consists of 2 parts. First, x rgb_masked and x rgb_ori are added element by element to fuse the masked feature with the original feature information. Secondly, the added result is sent to a convolutional layer with a kernel size of 3 to obtain x out_rgb . Similarly, x out_ir can be obtained. That is, the first fusion result of each modal feature is obtained through the feature fusion module.
[0061] x out_rgb = Conv 3×3 (x rgb_ori + x rgb_ori × x mask_rgb )
[0062] x out_ir = Conv 3×3 (x ir_ori + x ir_ori × x mask_ir )
[0063] Step (5): Concatenate and merge the results of different modal feature processing obtained in S4 to obtain the second fusion result;
[0064] Among them, the specific implementation of the second fusion in step (5) is to concatenate x out_rgb and x out_ir along the channel with dimension 1 to fuse the processing results of the two different modal features to obtain x out_cat .
[0065]
[0066] Step (6): Send the result of S5 into the SE module to screen out the effective information features and suppress the redundant features to obtain the finally pixel-level fused image;
[0067] The SE module in step (6) is used to strengthen the effective features in the feature map of x out_tat fused in step (5). The SE module first compresses the global spatial information of each channel through the global average pooling layer, and then provides a specific weight for each channel through the gate mechanism composed of two fully connected layers. This weight is used to reflect the importance of the channel. After that, the redundant information is screened out by weighting and the effective information is strengthened. Finally, the fused image I f is obtained for the subsequent object detection task.
[0068] I f = SE(x out_cat )
[0069] Step (7): Feed the fused image obtained in S6 into the YOLOv8n lightweight detection network to obtain the detection result.
[0070] In step (7), the YOLOv8n lightweight detection network includes a YOLOv8 backbone network and a head network. Here, n indicates that the total depth and total width of the network are set to 0.33 and 0.25 respectively to ensure the light weight of the network. The input of this detection network is the fused image I f , and the output is an identification result map with confidence and class labels.
[0071] As Figure 3 shown is a block diagram of a pixel-level fusion method. Pixel-level fusion refers to performing fusion on the original data by combining data information from different sensors in the data preprocessing stage. This fusion method has a relatively small number of parameters, but the technology for mask generation and application is simple, which may lead to the loss of modal features. And during the final fusion, only simple splicing is used, resulting in poor fusion quality.
[0072] As Figure 4 shown is a block diagram of the pixel-level multi-spectral fusion method based on mask enhancement of the present invention, corresponding to steps (2) - (6).
[0073] As Figure 5 shown is a visualization comparison diagram of the fusion effect between the present invention and the traditional pixel-level fusion method. It shows that the mask-enhanced pixel-level multi-spectral fusion method designed by the present invention can generate a higher-quality fusion feature map.
[0074] As shown in Table 1 is the comparison of the detection results between the present invention and the traditional pixel-level fusion method. By comparing the precision index mAP, it shows that the small target detection method of pixel-level multi-spectral fusion based on mask enhancement designed by the present invention can improve the detection and recognition accuracy of small targets.
[0075] Table 1 Comparison of Pixel-Level Fusion Detection Results
[0076]
[0077] As shown in Table 2 is the comparison of the total number of parameters of the method of the present invention with that of the MGMF feature-level fusion method. By comparing the number of parameters, it shows that the small target detection method of pixel-level multi-spectral fusion based on mask enhancement designed by the present invention is lighter and more suitable for deployment on edge devices.
[0078] Table 2 Comparison of the Total Number of Parameters of the Method of the Present Invention with Figure 1 the Number of Parameters of the Feature-Level Fusion Method
[0079]
[0080] Example 2
[0081] Example 2 of the present invention proposes a pixel-level multi-spectral fusion small target detection system based on mask enhancement, which is implemented based on the method of Example 1. The system includes:
[0082] An acquisition module, configured to acquire remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of visible light spectrum and infrared spectrum;
[0083] A secondary fusion module, configured to split the multi-spectral modal image pair and input it into the feature fusion module at the same time, respectively obtain their respective masks, then obtain the feature information after mask processing according to their respective masks, and perform feature fusion with the original image after splitting to obtain the first fusion result of a single modal feature, and then perform channel splicing and merging to obtain the result after the second fusion;
[0084] A screening and suppression module, configured to input the result after the second fusion into the SE module, screen out effective information features, suppress redundant features, and obtain an image after pixel-level fusion;
[0085] A detection and output module, configured to input the image after pixel-level fusion into a lightweight detection network to obtain a detection result.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that any modification or equivalent replacement of the technical solutions of the present invention does not depart from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A pixel-level multi-spectral fusion small target detection method based on mask enhancement, comprising: Step (1): Obtain the remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of visible light spectrum and infrared spectrum; Step (2): Split the multi-spectral modal image pair and input it into the feature fusion module simultaneously, respectively obtain their respective masks, then obtain the feature information after mask processing according to their respective masks, and perform feature fusion with the original image after splitting to obtain the first fusion result of the single-modal feature, and then perform channel splicing and merging to obtain the result after the second fusion; Step (3): Input the result after the second fusion into the SE module, screen out the effective information features, suppress the redundant features, and obtain the pixel-level fused image; Step (4): Input the pixel-level fused image into the lightweight detection network to obtain the detection result; The feature fusion module in step 2) includes two mask generation modules with the same structure. Among them, the input of the first mask generation module is the feature map x of the visible light spectrum rgb_ori , and the output is the mask x of the visible light spectrum mask_rgb . The input of the second mask generation module is the feature map x of the infrared spectrum ir_ori , and the output is the mask x of the infrared spectrum mask_ir . Each of the mask generation modules includes: two convolutional layers with a convolutional kernel size of 3, a ReLU activation function, and a Sigmoid activation function. Among them, The first convolutional layer is used for preliminary feature extraction of the input data; The ReLU activation function is used to introduce non-linearity so that the network can better express features; The second convolutional layer is used to transform the features extracted previously; The Sigmoid activation function is used to map the value of the mask to between 0 and 1.
2. The method for detecting small targets in pixel-level multi-spectral fusion based on mask enhancement according to claim 1, wherein The feature fusion module in step (2) includes two multiplication layers with the same structure, which are used to screen the pixel elements in the original image according to the mask. Among them, Mask x of the visible light spectrum mask_rgb Perform element-wise multiplication with the feature map x of the visible light spectrum rgb_ori Perform element-wise multiplication with the mask x of the infrared spectrum mask_ir Perform element-wise multiplication with the feature map x of the infrared spectrum ir_ori Perform element-wise multiplication respectively to obtain the corresponding masked feature information rgb_masked and ir_masked .
3. The method for detecting small targets in pixel-level multi-spectral fusion based on mask enhancement according to claim 2, wherein The feature fusion module in step (2) includes a first fusion network, and the first fusion network includes: two addition layers with the same structure and a convolutional layer with a convolution kernel size of 3; among them, The first addition layer is used to add x rgb_masked element-wise with x rgb_ori ; The first convolutional layer is used to perform convolution on the output of the first addition layer to obtain the first fusion result x out_rgb of the visible light modality features; The second addition layer is used to add x ir_masked element-wise with x ir_ori ; The second convolutional layer is used to perform convolution on the output of the second addition layer to obtain the first fusion result x out_ir of the infrared light modality features.
4. The pixel-level multi-spectral fusion small target detection method based on mask enhancement according to claim 3, wherein The feature fusion module in step 2) includes channel splicing and merging, which is used for x out_rgb and x out_ir to be spliced along the channels with dimension 1 to obtain the result x out_cat after the second fusion.
5. The method for detecting small targets in pixel-level multi-spectral fusion based on mask enhancement according to claim 4, wherein The SE module in step (3) includes a global average pooling layer and a gate mechanism composed of two fully connected layers; among them, The global average pooling layer is used to compress the global spatial information of each channel; The described gate mechanism is used to provide a specific weight for each channel. This weight is used to reflect the importance of the channel, sift out redundant information through weighting, and strengthen effective information to obtain the pixel-level fused image I f .
6. The method for detecting small targets in pixel-level multi-spectral fusion based on mask enhancement according to claim 5, wherein The lightweight detection network in step 4) is the YOLOv8n lightweight detection network, and the input is the fused image I f , and the output is the recognition result map with confidence and class labels.
7. The method for detecting small targets in pixel-level multi-spectral fusion based on mask enhancement according to claim 6, wherein The lightweight detection network includes a backbone network and a head network.
8. A system for a mask-enhanced pixel-level multi-spectral fusion small target detection method according to claim 1, characterized in that, Comprising: An acquisition module, which is used to obtain the remote sensing image data to be processed, where the remote sensing image data includes a multi-spectral image pair of visible light spectrum and infrared spectrum; A secondary fusion module, which is used to split the multi-spectral modal image pair and input it into the feature fusion module simultaneously, respectively obtain their respective masks, then obtain the feature information after mask processing according to their respective masks, and perform feature fusion with the original image after splitting to obtain the first fusion result of the single-modal feature, and then perform channel splicing and merging to obtain the result after the second fusion; A screening and suppression module, which is used to input the result after the second fusion into the SE module, screen out the effective information features, suppress the redundant features, and obtain the pixel-level fused image; And A detection output module, which is used to input the pixel-level fused image into the lightweight detection network to obtain the detection result.
Citation Information
Patent Citations
Infrared and visible light image fusion method and system and medium
CN114529794A
Pixel-level real-time multispectral image fusion method based on adaptive weight and target perception
CN116596822A