Remote sensing image semantic segmentation method based on fusion network and frequency domain guidance

By fusing the network with frequency domain guidance, the problem of balancing local details and global context in the semantic segmentation of remote sensing images is solved, and efficient remote sensing image segmentation accuracy and consistency are achieved, especially in multi-scale object segmentation.

CN120689755APending Publication Date: 2025-09-23XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510851796.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies have difficulty in balancing the local details of high-resolution representation and the global context of deep semantic features in semantic segmentation of remote sensing images, resulting in insufficient segmentation accuracy and consistency.

Method used

A method based on fusion network and frequency domain guidance is adopted. Multi-scale local features are extracted through the encoder, and adjacent resolution features are merged using multiple identical fusion networks. Frequency domain guided decoding is performed through the decoder. Frequency domain filters and strip fusion blocks are combined for feature fusion and decoding to achieve a balance between spatial details and global context.

Benefits of technology

The accuracy and consistency of semantic segmentation of remote sensing images are improved, especially the segmentation performance of multi-scale objects is significantly improved while maintaining computational efficiency and low computational cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689755A_ABST
    Figure CN120689755A_ABST
Patent Text Reader

Abstract

The invention relates to the field of remote sensing semantic segmentation, and discloses a remote sensing image semantic segmentation method based on fusion network and frequency domain guidance, which comprises the following steps: S1, extracting multi-scale local features of a remote sensing image through an encoder; s2, combining features with adjacent resolutions in the multi-scale local features through a plurality of same fusion networks to obtain a plurality of fusion features; and S3, performing frequency domain guided decoding on the multi-scale local features and the plurality of fusion features through a decoder to obtain semantic segmentation features. According to the method, the multi-scale local features and the fusion features in the remote sensing image are acquired to perform decoding operation, so that both acquisition of spatial details and efficient modeling of global context can be taken into account.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing semantic segmentation, and in particular to a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance. Background Art

[0002] Rapid advances in sensor and aerospace technologies have greatly facilitated the acquisition of high-resolution satellite and aerial remote sensing imagery. These images provide detailed representations of various geographic features, including urban structures, farmlands, forested areas, and water bodies. As a fundamental technique for segmenting surface images into meaningful objects or semantic categories, remote sensing image segmentation plays a key role in applications such as geographic information systems (GIS), agricultural planning, environmental monitoring, and disaster management.

[0003] For position-sensitive vision tasks such as semantic segmentation and object detection, high-resolution representations are crucial for capturing edge, texture, and detail information for accurate localization. However, while shallow semantic features retain higher resolution and rich position / detail information, they exhibit limited semantic content and increased noise due to the small number of convolutional layers. In contrast, deep semantic features capture more abstract and semantically rich information through an expanded receptive field, but at the expense of resolution and insufficient preservation of local details. Summary of the Invention

[0004] The purpose of the present invention is to provide a remote sensing image semantic segmentation method based on fusion network and frequency domain guidance, aiming to solve the semantic segmentation problem of remote sensing images.

[0005] The present invention provides a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance, comprising: S1, extracting multi-scale local features of remote sensing images through encoder; S2, multiple fusion features are obtained by merging features of adjacent resolutions in multi-scale local features through multiple identical fusion networks; S3. The decoder decodes the multi-scale local features and multiple fusion features to obtain semantic segmentation features.

[0006] The present invention obtains multi-scale local features and multiple fusion features in remote sensing images for decoding operations, so that both obtaining spatial details and efficiently modeling global context can be taken into account.

[0007] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it is implemented in accordance with the contents of the specification, and in order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0009] Figure 1 Flowchart of a remote sensing image semantic segmentation method based on fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 2 Schematic diagram of the network structure of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 3 1 is a schematic diagram of a single fusion network of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 4 1 is a schematic diagram of a frequency domain filter of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 5 1 is a global branch diagram of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 6 1 is a schematic diagram of local branches of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 7 1 is a schematic diagram of a strip fusion block of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention; Figure 8 3. It is a schematic diagram of model comparison of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance on the LoveDA dataset according to an embodiment of the present invention; Figure 9 This is a schematic diagram of model comparison on the Potsdam dataset of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0011] Method Example According to an embodiment of the present invention, a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance is provided. Figure 1 This is a flowchart of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention. Figure 1 As shown, specifically including: S1, extracting multi-scale local features of remote sensing images through encoder; S2, multiple fusion features are obtained by merging features of adjacent resolutions in multi-scale local features through multiple identical fusion networks; S3. The decoder performs frequency domain-guided decoding on multi-scale local features and multiple fusion features to obtain semantic segmentation features.

[0012] Figure 2 : is a schematic diagram of the network structure of a remote sensing image semantic segmentation method based on a fusion network and frequency domain guidance according to an embodiment of the present invention. From left to right are the encoder, multiple fusion networks and decoders connected in sequence. The encoder is continuous downsampling. The decoder specifically includes: three identical band fusion blocks and four identical frequency domain filters with the same function and structure. The first frequency domain filter is connected to the first band fusion block, the second band fusion block is connected to the second frequency domain filter, the third frequency domain filter is connected to the third band fusion block, and the fourth frequency domain filter is connected to the third band fusion block.

[0013] The specific embodiments of the present invention are as follows: S1, extracting multi-scale local features of remote sensing images through encoder; The S1 specifically includes: performing four-layer downsampling on the remote sensing image through an encoder to generate four-layer multi-scale local features.

[0014] In the embodiment of the present invention, the number of downsampling channels C is doubled, and the width W and height H are halved.

[0015] S2, multiple fusion features are obtained by merging features of adjacent resolutions in multi-scale local features through multiple identical fusion networks; The S2 specifically includes: inputting the first layer multi-scale local features and the second layer multi-scale local features into the first fusion network to obtain the first fusion features; Inputting the first-layer multi-scale local features and the second-layer multi-scale local features into the first fusion network to obtain the first fusion features specifically includes: selecting the first-layer multi-scale local features input into the first fusion network as main features, adding the main features to other features input into the first fusion network and then inputting the enhanced features into BasicBlock to obtain enhanced fusion features, inputting the enhanced fusion features into the convolution layer and the normalization layer for processing to obtain a single-channel feature map, inputting the single-channel feature map into the activation function to generate an attention coefficient, and multiplying the attention coefficient by the main features to output the first fusion features.

[0016] The remaining main features are the second-layer multi-scale local features and the third-layer multi-scale local features; the acquisition process of the second fusion features and the third fusion features is similar to that of the first fusion features.

[0017] Inputting the first layer multi-scale local features, the second layer multi-scale local features and the third layer multi-scale local features into the second fusion network to obtain the second fusion features; The second layer multi-scale local features, the third layer multi-scale local features and the fourth layer multi-scale local features are input into the third fusion network to obtain the third fusion features.

[0018] In an embodiment of the present invention, multiple fusion networks are established to avoid feature loss and eliminate redundant features caused by frequent upsampling and downsampling. The present invention only fuses adjacent features and uses channel gating signals to control and optimize the information flow within the fused feature map. This method gradually combines key features with the attention coefficients learned by the network, enhancing the activation of relevant features while suppressing irrelevant ones. This prevents the loss of edge features and the discontinuity caused by frequent resampling, thereby enhancing feature fusion. This method enables the present invention to retain more useful information and improve segmentation accuracy.

[0019] like Figure 3 As shown, first adjust all fusion branch features to be consistent with the main feature F main The spatial dimensions of are the same, and then summed and fused features are enhanced by BasicBlock. The specific formula is as follows: ; Here, Represents fused features, and BasicBlock refers to a basic residual convolution module in ResNet. The generated fused features are first processed using 1×1 convolution and batch normalization (BN) layers to obtain a single-channel feature map. Subsequently, the single-channel feature map generates the attention coefficient through the Sigmoid activation function. Finally, F main Multiplying with the attention coefficient generates the final output: ; Among them, δ k×k () indicates that the convolution kernel size is k×k, k is 1, and the number of channels changes from C to 1. BN refers to batch normalization (BatchNorm2d), and Sigmoid represents the Sigmoid activation function.

[0020] S3. The decoder decodes the multi-scale local features and multiple fusion features to obtain semantic segmentation features.

[0021] The S3 specifically includes: inputting the fourth layer multi-scale local feature into the fourth frequency domain filter to obtain the fourth filter feature, inputting the fourth filter feature and the third fusion feature into the third strip fusion block to obtain the third strip fusion feature, inputting the third strip fusion feature into the third frequency domain filter to obtain the third filter feature, inputting the third filter feature and the second fusion feature into the second strip fusion block to obtain the second strip fusion feature, inputting the second strip fusion feature into the second frequency domain filter to obtain the second filter feature, inputting the second filter feature and the first fusion feature into the first strip fusion block to obtain the first strip fusion feature, and inputting the first strip fusion feature into the first frequency domain filter to obtain the semantic segmentation feature.

[0022] In the embodiment of the present invention, the fourth filter feature, the third filter feature, and the second filter feature are all upsampled by 2 times and input into the corresponding strip fusion block, and the semantic segmentation feature is upsampled by 4 times and output; The step of inputting the fourth fusion feature into the fourth frequency domain filter to obtain the fourth filter feature specifically includes: splitting the fourth fusion feature into local channel features and global channel features using the Split function, inputting the local channel features into the local branch to obtain local features, inputting the global channel features into the global branch to obtain global features, convolving the local features and the global features, inputting the convolution layer and layer normalization to obtain normalized features, inputting the normalized features into the multilayer perceptron to obtain multilayer perception features, and outputting the fourth filter feature after adding the multilayer perception features to the input fourth fusion feature.

[0023] In an embodiment of the present invention, the present invention adopts a more efficient frequency domain mechanism to model global information and uses convolution to enhance local information.

[0024] Frequency domain filters such as Figure 4 As shown, the formula used for a single frequency domain filter is as follows: ; F out Represents the filtering features, uses 1x1 convolution to adjust the number of channels to C, LN represents layer normalization, MLP represents multi-layer perceptron, A represents global features, V represents local features, and X represents the features of the input frequency domain filter.

[0025] The inputting of global channel features into the global branch to obtain global features specifically includes: inputting the input global channel features into deep convolution to obtain Q and K, converting Q and K into frequency domain features and multiplying them to obtain multiplied features, modulating the multiplied features to obtain different channel spectrum distributions, and performing inverse Fourier transform on the channel spectrum distribution to generate global features.

[0026] In the embodiment of the present invention, the global branch is as follows Figure 5As shown in Figure 1, the global branch is used as a frequency domain guidance method. Q and K are respectively passed through the first linear layer and the second linear layer and then converted into the frequency domain. The correlation between Q and K is estimated in the frequency domain by converting them into frequency domain features and multiplying them. The formula is as follows: F ; Among them, FFT is fast Fourier transform, Q represents query vector, and K represents key vector.

[0027] Use 1×1 partial convolution and BN to modulate the spectral distribution of different channels, and generate the global feature A through inverse Fourier transform. The formula is as follows: ; Among them, IFFT is the inverse fast Fourier transform, PConv() represents partial convolution, uses a convolution kernel size of 1x1, and BN refers to batch normalization (BatchNorm2d).

[0028] Inputting the local channel features into the local branch to obtain the local features specifically includes: inputting the local channel features into a multi-level depth-separable convolution to obtain the local features.

[0029] Local branches such as Figure 6 As shown, the formula used for multi-level depth-separable convolution is as follows: ; Among them, BN stands for batch normalization, DWC 3×3()表示卷积核为3×3的深度可分离卷积,x表示输入特征, , represents a 1x1 convolutional layer (Conv).

[0030] ; DWC 5×5()表示卷积核为5×5的深度可分离卷积,DWC7×7 () represents a depth-wise separable convolution with a convolution kernel of 3×3, v represents the local channel feature, and V represents the local feature.

[0031] The inputting the fourth filter feature and the third fusion feature into the third strip fusion block to obtain the third strip fusion feature specifically includes: summing the fourth filter feature and the third fusion feature and inputting them into the first convolution layer for alignment, splitting them into two identical feature parts using the Split function after alignment, inputting the two feature parts into different strip convolutions to obtain convolution feature one and convolution feature two respectively, summing the fourth filter feature, the third fusion feature, the convolution feature one and the convolution feature two and inputting them into the second convolution layer for fusion to obtain the third strip fusion feature.

[0032] In the embodiment of the present invention, during the feature fusion process, simply adding two features in pairs usually leads to information loss because the receptive fields of the features do not match. The present invention uses strip convolution to align the receptive fields of different features, thereby enhancing the effect of feature fusion.

[0033] Specifically, the banded fusion blocks such as Figure 7 As shown; For the shallow features of the input and deep features F h , the present invention first performs element-wise summation to integrate their information, and then performs 1×1 convolution to align channel correlation; The formula is as follows: ; Then, the fused features are equally divided into two equal parts along the channel dimension, and each part undergoes a different strip convolution operation to strengthen the contextual attention of the central area. The formula used is as follows: ; Finally, the feature flow is optimized by residual summation and 1x1 convolution to generate fusion features, DWC 5×1()表示卷积核为5×1的深度可分离卷积,DWC1×5 () indicates a depth-wise separable convolution with a convolution kernel of 1×5, DWC 7×1()表示卷积核为7×1的深度可分离卷积,DWC1×7()表示卷积核为1×7的深度可分离卷积; The formula is as follows: , F fuse Indicates fusion features.

[0034] In the embodiment of the present invention, an experiment is set up, the experimental setting parameters use the AdamW algorithm, and a cosine learning rate change strategy is adopted as the optimizer, and the basic learning rate is 6e-4.

[0035] The model was trained on an NVIDIA 3090 24GB graphics card. Images were randomly cropped into 512×512 patches. For the Vaihinge and Potsdam datasets, the training epoch was set to 105. For the LoveDA dataset, the training epoch was 60. During training, augmentation techniques such as random scaling, random vertical flipping, random horizontal flipping, and random rotation were used.

[0036] Evaluation Metrics: This paper uses mean Intersection over Union (mIoU) as the primary evaluation metric for the LoveDA dataset. For the Vaihingen and Potsdam datasets, this paper uses mIoU, overall accuracy (OA), and mean F1 score (mF1) as evaluation metrics.

[0037] Module ablation experiment When only the frequency domain filter is enabled, mIoU increases to 84.23% (+0.41%), indicating that the frequency domain filter enhances semantic consistency by filtering global information in the frequency domain, significantly improving the model's ability to capture global semantic clues.

[0038] After adding the fusion network, mIoU further improves, demonstrating that the fusion network uses high-resolution features to guide low-resolution feature learning and refine local details. Although the gain of using SFB alone is limited, its combination with the fusion network optimizes the alignment of multi-scale local features, ultimately achieving the best performance when all components are activated.

[0039] Comparison with SOTA models We compare the effectiveness of our model with recent SOTA methods using four widely used open datasets to verify the effectiveness of our model.

[0040] 1) LoveDA dataset results: The proposed method shows significant advantages over all the compared methods, with an mIoU of 55.65%, which is nearly 0.93% higher than the second best method ConvLSR (using the same ConvNeXt-Small backbone), showing the best overall segmentation performance. Although the model of the present invention only requires 54.82% of its FLOPs and 68.02% of its parameters, its mIoU is 3.16% higher than it. Although the scores of roads and water bodies are relatively low but still competitive, the mIoU values ​​of agriculture and background reach 62.29% and 56.71% respectively, far exceeding other methods. Figure 8 As shown, GroundTruth is the real annotation, UperNet, DeeplabV3+, Hrnet, DCSwin and SFFNet are semantic segmentation methods in the existing technology, which segment the buildings, roads, water, wasteland, forest, farmland and background in the picture. The fusion method of the present invention better retains the details and performs well in maintaining the internal continuity of objects of different scales.

[0041] Results on the Vaihingen and Potsdam datasets: On the Potsdam dataset, our method achieved 93.28% mF1, 91.75% OA, and 87.61% mIoU, comprehensively surpassing existing state-of-the-art methods. Compared to traditional high-resolution networks such as HRNet (85.93% mIoU), our method improved mIoU by 1.68% through high-resolution feature fusion.

[0042] like Figure 9 As shown, GroundTruth is the real annotation, DeeplabV3+, MANet, FT-UnetFormer, SegFormer and SFFNet are semantic segmentation methods in the existing technology, which segment impervious surfaces, low vegetation, buildings, trees, cars and backgrounds. This method shows superior performance in building and vegetation edge refinement and regional internal consistency, thanks to its multi-scale local feature alignment mechanism and global-local context modeling. On the Waihingen and Potsdam datasets, our method outperforms ConvLSR by 0.23%, 0.23%, and 0.36% in mF1, OA, and mIoU, respectively, while requiring fewer parameters and lowering computational cost. This advantage stems from the lower computational complexity and global modeling capabilities of frequency-domain filters. Our method demonstrates particularly strong performance in fine-grained segmentation of low vegetation and car categories, outperforming the second-best method, ConvLSR, by 0.75% and 0.3%, respectively.

[0043] Beneficial effects: This paper proposes a new network architecture for semantic segmentation of remote sensing images, with the following main contributions: 1) Fusion Network: A gated feature fusion mechanism that selectively integrates features from adjacent resolutions via channel-wise attention, maximizing high-resolution guidance while eliminating cross-scale interaction redundancy.

[0044] 2) Frequency Domain Filter: An efficient global context modeling module that combines deep convolution to enhance local features with a frequency domain attention mechanism implemented via Fast Fourier Transform (FFT). This reduces computation while maintaining accuracy, achieving self-attention-level performance.

[0045] 3) Strip Fusion Block: A cross-scale feature alignment module that uses strip convolution to unify the receptive fields between network layers, ensuring spatial consistency of feature fusion and improving segmentation accuracy, especially for linear structures.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. These modifications or replacements of the technical solutions of the embodiments of the present invention do not cause the essence of the corresponding technical solutions to deviate from the scope of this solution.

Claims

1. A remote sensing image semantic segmentation method based on fusion network and frequency domain guidance, characterized in that: include, S1, extracting multi-scale local features of remote sensing images through encoder; S2, multiple fusion features are obtained by merging features of adjacent resolutions in multi-scale local features through multiple identical fusion networks; S3. The decoder performs frequency domain-guided decoding on multi-scale local features and multiple fusion features to obtain semantic segmentation features.

2. The method according to claim 1, characterized in that The S1 specifically includes: performing four-layer downsampling on the remote sensing image through an encoder to generate four-layer multi-scale local features.

3. The method according to claim 2, characterized in that The S2 specifically includes: inputting the first layer multi-scale local features and the second layer multi-scale local features into the first fusion network to obtain the first fusion features; Inputting the first layer multi-scale local features, the second layer multi-scale local features and the third layer multi-scale local features into the second fusion network to obtain the second fusion features; The second layer multi-scale local features, the third layer multi-scale local features and the fourth layer multi-scale local features are input into the third fusion network to obtain the third fusion features.

4. The method according to claim 3, characterized in that The step of inputting the first-layer multi-scale local features and the second-layer multi-scale local features into the first fusion network to obtain the first fusion features specifically includes: selecting the first-layer multi-scale local features input into the first fusion network as main features, adding the main features to other features input into the first fusion network, and then inputting the main features into the basic residual convolution module for enhancement to obtain enhanced fusion features, inputting the enhanced fusion features into the convolution layer and the normalization layer for processing to obtain a single-channel feature map, inputting the single-channel feature map into the activation function to generate an attention coefficient, and multiplying the attention coefficient by the main features to output the first fusion features.

5. The method according to claim 4, characterized in that The decoder specifically includes: three identical strip fusion blocks and four identical frequency domain filters, the first frequency domain filter is connected to the first strip fusion block, the second strip fusion block is connected to the second frequency domain filter, the third frequency domain filter is connected to the third strip fusion block, and the fourth frequency domain filter is connected to the third strip fusion block.

6. The method according to claim 5, characterized in that The S3 specifically includes: inputting the fourth layer multi-scale local feature into the fourth frequency domain filter to obtain the fourth filter feature, inputting the fourth filter feature and the third fusion feature into the third strip fusion block to obtain the third strip fusion feature, inputting the third strip fusion feature into the third frequency domain filter to obtain the third filter feature, inputting the third filter feature and the second fusion feature into the second strip fusion block to obtain the second strip fusion feature, inputting the second strip fusion feature into the second frequency domain filter to obtain the second filter feature, inputting the second filter feature and the first fusion feature into the first strip fusion block to obtain the first strip fusion feature, and inputting the first strip fusion feature into the first frequency domain filter to obtain the semantic segmentation feature.

7. The method according to claim 6, characterized in that The step of inputting the fourth fusion feature into the fourth frequency domain filter to obtain the fourth filter feature specifically includes: splitting the fourth fusion feature into two identical channel features using the Split function, recorded as global channel features and local channel features, inputting the local channel features into the local branch to obtain local features, inputting the global channel features into the global branch for frequency domain guidance to obtain global features, convolving the local features and the global features, inputting the convolution layer and layer normalization to obtain normalized features, inputting the normalized features into the multilayer perceptron to obtain multilayer perception features, and outputting the fourth filter feature after adding the multilayer perception features to the input fourth fusion feature.

8. The method according to claim 7, characterized in that The inputting of global channel features into the global branch to obtain global features specifically includes: inputting the input global channel features into deep convolution to obtain Q and K, converting Q and K into frequency domain features and multiplying them to obtain multiplied features, modulating the multiplied features to obtain different channel spectrum distributions, and performing inverse Fourier transform on the channel spectrum distribution to generate global features.

9. The method according to claim 7, characterized in that Inputting the local channel features into the local branch to obtain the local features specifically includes: inputting the local channel features into a multi-level depth-separable convolution to obtain the local features.

10. The method according to claim 6, characterized in that Inputting the fourth filter feature and the third fusion feature into the third strip fusion block to obtain the third strip fusion feature specifically includes: summing the fourth filter feature and the third fusion feature and then inputting them into the first convolution layer for alignment, and after alignment, using the Split function to split them into two identical features, and inputting the two features into different strip convolutions to obtain convolution feature one and convolution feature two respectively, and summing the fourth filter feature, the third fusion feature, the convolution feature one and the convolution feature two and inputting them into the second convolution layer for fusion to obtain the third strip fusion feature.