Remote sensing image target detection method based on depth separable convolution and multi-scale fusion
Through the method of depth separation convolution and multi-scale fusion, the problem of insufficient feature extraction efficiency and accuracy in remote sensing image object detection is solved, and efficient and accurate object detection is achieved in resource-constrained environments.
Patent Information
- Application Number
- CN202510477358.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-15
AI Technical Summary
The existing remote sensing image object detection methods have problems with insufficient computing efficiency and accuracy when extracting features. Especially when processing high-resolution images, target feature information is easily lost, and the performance of traditional convolution modules in resource-constrained environments is limited.
The method of fusion of depth separation convolution and multi-scale is adopted. Multi-scale features are extracted in parallel through 3x3 standard convolution, 3x3 depth separation convolution, and 5x5 depth separation convolution, and combined batch normalization and SiLU activation function, and finally channel compression is performed through 1x1 standard convolution to achieve feature extraction.
While reducing the computational volume and model parameters, the accuracy and robustness of remote sensing image object detection is significantly improved, especially in complex backgrounds and multi-scale object recognition, suitable for resource-constrained environments.
Smart Images

Figure CN120495620A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image-based target detection technology, in particular to a remote sensing image target detection technology. Background Art
[0002] Remote sensing imagery is widely used in fields such as environmental monitoring, agriculture, and urban planning. Due to their high dimensionality and complexity, remote sensing images pose numerous challenges to traditional image processing methods, particularly in terms of the efficiency and accuracy of image feature extraction. Therefore, extracting effective features from remote sensing images without incurring excessive computational overhead has become a key research topic.
[0003] In existing object detection networks, downsampling is often used to reduce image size and increase the depth of feature maps. While this downsampling operation can accelerate computation and improve efficiency, it also leads to a gradual loss of high-frequency details of the target. This is particularly true in high-resolution images such as remote sensing images, where many of the target's features are ignored or distorted during the downsampling process. Consequently, traditional convolutional modules have limited ability to extract target features in remote sensing environments, particularly when processing small, distant targets or those with rich details.
[0004] Depthwise separable convolutions were first proposed by MobileNet, primarily to reduce computational effort and parameter requirements. By separating spatial and channel-wise convolutions, depthwise convolutions optimize computation and are suitable for environments with limited computing resources. However, while depthwise convolutions improve computational efficiency, they are limited in that they cannot effectively capture complex cross-channel relationships in images.
[0005] Standard convolution can operate across all channels through the same convolution kernel, which can effectively capture the interaction information between channels, but the computation is large, especially when the number of image channels is large.
[0006] Currently, there are many solutions that combine depthwise separable convolution with standard convolution. These solutions aim to balance computational efficiency and model expressiveness, leveraging the computational efficiency of depthwise convolution while retaining the ability of standard convolution to capture complex relationships between channels. However, properly switching between these two convolution methods at different levels remains a technical challenge.
[0007] At the same time, multi-scale convolution methods have been widely used in target detection and image segmentation tasks, but they usually rely on multi-level convolution structures or multi-stage processing to extract features of different scales. This means that each layer needs to process information of different scales, resulting in increased computational complexity and memory consumption, especially when processing high-resolution images, where the computational overhead increases significantly. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a remote sensing image target detection method that can balance the performance of depthwise separable convolution and standard convolution to perform multi-scale feature extraction, thereby improving the performance of target detection while avoiding adding too many model parameters.
[0009] The technical problem adopted by the present invention to solve the above technical problems is a remote sensing image feature extraction method based on depthwise separable convolution and multi-scale fusion, comprising the following steps:
[0010] Multi-scale feature extraction steps: Perform 3x3 standard convolution on the input remote sensing image data to extract preliminary features. The preliminary features are then processed in parallel through 3x3 depthwise separable convolution, 3x3 standard convolution, and 5x5 depthwise separable convolution to obtain three scale features.
[0011] Feature fusion step: The three scale features are spliced with the input remote sensing image data to obtain spliced features, and the spliced features are batch normalized and then passed through the activation function;
[0012] Target detection step: The activated features are subjected to 1x1 standard convolution to complete channel compression to obtain the final remote sensing image features, and the remote sensing image features are input into the target detection module to complete target detection.
[0013] The present invention can significantly reduce the amount of computation and the number of model parameters by adopting depthwise separable convolution and multi-scale fused convolution. This structure effectively reduces the redundant calculations in traditional convolution operations. The specific selection of convolution form and convolution kernel size balances the ability to capture efficient calculations and complex relationships between channels, which can better extract rich features in remote sensing images and enhance the network's ability to recognize various targets. It improves the ability to detect targets while effectively fusing feature information from different scales without significantly increasing the amount of computation and parameters.
[0014] It is particularly suitable for application in resource-constrained environments (such as mobile devices and edge devices); at the same time, it maintains high-performance feature extraction capabilities, especially in the detection of multi-scale targets.
[0015] Specifically, let the input remote sensing image data be I, with a size of H×W×C, where H is the height of the image, W is the width of the image, and C is the number of channels of the image;
[0016] Through a 3x3 standard convolution layer Conv 3x3 Process the input image to obtain the preliminary feature map F1:
[0017]
[0018] Where W1(i,j) is the weight of the 3x3 convolution kernel, x,y are the horizontal and vertical coordinates of the input image, and i,j are the horizontal and vertical offset indexes of the convolution kernel;
[0019] Specifically, the preliminary feature map F1 is divided into three branches to extract features at three scales:
[0020] Branch 1: Through 3x3 depth-separable convolution DWConv 3x3 Extract features to get F DW3x3 :
[0021] F DW3x3 =DWConv 3x3 (F1)=Conv depthwise,3x3 (F1) Conv pointwise,1x1 (F1)
[0022] Among them, Conv depthwise,3x3 Represents 3x3 depth convolution, Conv pointwise,1x1 Represents 1x1 point-by-point convolution;
[0023] Branch 2: Extract features through 3x3 standard convolution Conv3x3 to obtain F Conv3x3 :
[0024] F Conv3x3 =Conv 3x3 (F1)
[0025] Branch 3: Through 5x5 depth-separable convolution DWConv 5x5 Extract features to get F DW5x5 :
[0026] F DW5x5 =DWConv 5x5 (F1)=Conv depthwise,5x5 (F1) Conv pointwise,1x1 (F1).
[0027] Specifically, the three scale features are spliced with the input remote sensing image data I to obtain the splicing feature F concat The specific expression is:
[0028] F concat =concat(F DW3x3 ,F Conv3x3 ,F DW5x5 ,I)
[0029] Among them, F DW3x3 、F Conv3x3 、F DW5x5These are the three scale features output by 3x3 depth-separable convolution, 3x3 standard convolution, and 5x5 depth-separable convolution. The concatenation operation is performed along the channel dimension to obtain a feature F containing multiple scale information. concat , the size is H×W×(C+K1+K2+K3), where K1, K2 and K3 are the number of channels of the three scale features respectively;
[0030] Specifically, the splicing feature F concat Perform batch normalization BN processing to make the batch normalized feature F BN has a mean of 0 and a variance of 1:
[0031]
[0032] Where μ and σ are the mean and standard deviation of batch normalization, γ and β are the learned scaling and offset coefficients, respectively.
[0033] Specifically, the SiLU activation function is used to normalize the batch feature F BN Perform nonlinear transformation to obtain the activated feature F SiLU :
[0034] F SiLU =F BN ·σ'(F BN )
[0035] The SiLU activation function σ' is defined as:
[0036]
[0037] Among them, x is the independent variable of the function and e is a natural constant.
[0038] Specifically, the activated feature F SiLU Through 1x1 standard convolution Conv 1x1 Perform channel compression to obtain the final remote sensing image feature F final :
[0039]
[0040] Among them, W 1x1 (i) is the weight of the 1x1 convolution kernel. Remote sensing image feature F final The dimensions are H×W×C final , C final It's F final The number of channels is used as the depth feature of the remote sensing image.
[0041] This method aims to efficiently extract deep features from remote sensing images through depthwise separable convolution and multi-scale feature fusion. This method is particularly suitable for environments with limited computing resources, such as mobile devices or edge computing devices, and can significantly reduce the amount of computation and model size while ensuring effective feature extraction.
[0042] The beneficial effect of this invention is that it can effectively preserve target features at different scales and in complex backgrounds in remote sensing images, significantly improving detection accuracy. This solves the problem of target feature loss during the downsampling process used in traditional feature extraction, especially when dealing with small targets at long distances and targets with rich details, thereby improving the robustness and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a flowchart of an implementation method for remote sensing image feature extraction based on depthwise separable convolution and multi-scale fusion;
[0044] Figure 2 This is a performance comparison chart of the traditional method and the improved method in terms of mAP@50 indicators on the RSOD, DIOR, and VHR-10 remote sensing datasets. DETAILED DESCRIPTION
[0045] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0046] like Figure 1 As shown in the figure, the steps of remote sensing image feature extraction based on depthwise separable convolution and multi-scale fusion before target detection are as follows:
[0047] Step 1: Receive remote sensing image data as input, extract preliminary features through 3x3 convolution operation, and divide the features into three branches, respectively, through 3x3 depth-wise separable convolution DWConv3x3, 3x3 standard convolution Conv3x3 and 5x5 depth-wise separable convolution DWConv5x5 for multi-scale feature extraction. The specific process includes:
[0048] Assume that the input remote sensing image data is I, with a size of H×W×C, where H is the height of the image, W is the width of the image, and C is the number of channels of the image;
[0049] First, the input image is processed through a 3x3 convolution layer to obtain the preliminary feature map F1. The convolution operation is:
[0050]
[0051] Where W1(i,j) is the weight of the 3x3 convolution kernel, and x,y are the coordinates of the input image;
[0052] Next, the obtained preliminary feature map F1 is divided into three branches for multi-scale feature extraction, which is performed through the following three branches of convolution operations:
[0053] Branch 1: Extract features through 3x3 depth-separable convolution DWConv3x3 to obtain F DW3x3 :
[0054] F DW3x3 =DWConv 3x3 (F1)=Conv depthwise,3x3 (F1) Conv pointwise,1x1 (F1)
[0055] Branch 2: Extract features through 3x3 standard convolution Conv3x3 to obtain F Conv3x3 :
[0056] F Conv3x3 =Conv 3x3 (F1)
[0057] Branch 3: Extract features through 5x5 depth-separable convolution DWConv5x5 to obtain F DW5x5 :
[0058] F DW5x5 =DWConv 5x5 (F1)=Conv depthwise,5x5 (F1) Conv pointwise,1x1 (F1)
[0059] Step 2: Concatenate the output features of the three branches with the input features, then perform batch normalization (BN) on the concatenated features to achieve feature standardization, and then perform nonlinear transformation using the SiLU activation function. The specific process includes:
[0060] First, the three feature maps F obtained in step 1 DW3x3 、F Conv3x3 、F DW5x5 It is concatenated with the original input feature map I to obtain a new feature map F concat , the calculation formula is:
[0061] F concat =concat(F DW3x3 ,F Conv3x3 ,F DW5x5 ,I)
[0062] The splicing operation is performed along the channel dimension to obtain a feature map containing multiple scale information with a size of H×W×(C+K1+K2+K3), where C is the number of channels of the input image, and K1, K2, and K3 are the number of output channels of the three branches respectively; Next, the spliced feature map F concatPerform batch normalization (BN) processing to make the mean of the feature map 0 and the variance 1. The specific formula is:
[0063]
[0064] Where μ and σ are F concat The mean and standard deviation of , γ and β are the learned scaling coefficients and offset coefficients;
[0065] Finally, the normalized feature map F is activated by the SiLU function. BN Perform nonlinear transformation to obtain the feature map F SiLU , the specific formula is:
[0066] F SiLU =F BN ·σ'(F BN )
[0067] Among them, the SiLU activation function is defined as:
[0068]
[0069] Step 3: Perform channel compression on the activated features through a 1x1 convolution operation to finally obtain the remote sensing image features. The specific process is as follows:
[0070] Feature F SiLU Channel compression is performed through 1x1 convolution to obtain the final feature F final :
[0071]
[0072] Among them, W 1x1 (i) is the weight of the 1x1 convolution kernel. 1x1 convolution compresses the number of channels through linear combination, and the final feature map size is H×W×C final , where C final is the number of output channels after 1x1 convolution; the final output feature map F final , as the depth feature of remote sensing images.
[0073] Afterwards, the remote sensing image feature F final Input the target detection module to complete target detection.
[0074] exist Figure 2The paper compares the proposed method with traditional methods in terms of mAP@50, evaluated on three different standard datasets. This method utilizes a remote sensing image feature extraction method based on multi-scale convolution and depthwise separable convolution, significantly improving the model's detection accuracy and efficiency by integrating multi-scale information and reducing computational effort. This method is particularly effective in recognizing complex backgrounds and multi-scale targets in remote sensing images, maintaining high accuracy while improving processing speed even with limited computing resources.
Claims
1. A remote sensing image target detection method based on depthwise separable convolution and multi-scale fusion, characterized by: The following steps are involved: Multi-scale feature extraction steps: Perform 3x3 standard convolution on the input remote sensing image data to extract preliminary features. The preliminary features are then processed in parallel through 3x3 depthwise separable convolution, 3x3 standard convolution, and 5x5 depthwise separable convolution to obtain three scale features. Feature fusion step: The three scale features are spliced with the input remote sensing image data to obtain spliced features, and the spliced features are batch normalized and then passed through the activation function; Target detection step: The activated features are subjected to 1x1 standard convolution to complete channel compression to obtain the final remote sensing image features, and the remote sensing image features are input into the target detection module to complete target detection.
2. The method according to claim 1, characterized in that Through a 3x3 standard convolution layer Conv 3x3 The input remote sensing image data I is processed to obtain the preliminary feature map F1: Among them, W1(i,j) is the weight of the 3x3 convolution kernel, x,y are the horizontal and vertical coordinates of the input image, and i,j are the horizontal and vertical offset indexes of the convolution kernel.
3. The method according to claim 2, characterized in that The preliminary feature map F1 is divided into three The branches perform feature extraction at three scales: Branch 1: Through 3x3 depth-separable convolution DWConv 3x3 Extract features to get F DW3x3 : F DW3x3 =DWConv 3x3 (F1)=Conv depthwise,3x3 (F1)·Conv pointwise,1x1 (F1) Among them, Conv depthwise,3x3 Represents 3x3 depth convolution, Conv pointwise,1x1 Represents 1x1 point-by-point convolution; Branch 2: Extract features through 3x3 standard convolution Conv3x3 to obtain F Conv3x3 : F Conv3x3 =Conv 3x3 (F1) Branch 3: Through 5x5 depth-separable convolution DWConv 5x5 Extract features to get F DW5x5 : F DW5x5 =DWConv 5x5 (F1)=Conv depthwise,5x5 (F1)·Conv pointwise,1x1 (F1)。 4. The method according to claim 3, characterized in that The three scale features are spliced with the input remote sensing image data I to obtain the splicing feature F concat The specific expression is: F concat =concat(F DW3x3 ,F Conv3x3 ,F DW5x5 ,I) Among them, F DW3x3 、F Conv3x3 、F DW5x5 They are 3x3 depth separable convolution, 3x3 standard The three scale features of the quasi-convolution and 5x5 depth-separable convolution outputs are concatenated along the channel dimension. Row, splicing feature F concat The size is H×W×(C+K1+K2+K3), where H, W and C are the height, width and number of channels of the input remote sensing image data, K1, K2 and K3 are the The number of channels of the three scale features is described.
5. The method according to claim 4, characterized in that: Splicing feature F concat Execute batch reduction The batch normalization feature F is obtained by BN processing. BN : Here, μ and σ are the mean and standard deviation of batch normalization, and γ and β are the learned scaling and offset coefficients, respectively.
6. The method according to claim 5, characterized in that The feature F after batch normalization is activated by SiLU function BN Perform nonlinear transformation to obtain the activated feature F SiLU : F SiLU =F BN ·σ'(F BN ) The SiLU activation function σ' is defined as: Among them, x is the independent variable of the function and e is a natural constant.
7. The method according to claim 6, characterized in that Activated Feature F SiLU Pass Through 1x1 standard convolution Conv 1x1 Perform channel compression to obtain the final remote sensing image feature F final : Among them, W 1x1 (i) is the weight of the 1x1 convolution kernel, the remote sensing image feature F final The dimensions are H×W×C final , C final It's F final The number of channels is used as the depth feature of the remote sensing image.