Remote sensing image segmentation method based on double-domain feature enhancement and adaptive scale fusion

By adopting dual-domain feature enhancement and adaptive scale fusion technology in remote sensing image segmentation, the shortcomings of traditional methods in capturing complex features and global information are solved, and higher semantic segmentation accuracy and robustness are achieved.

CN120219734APending Publication Date: 2025-06-27ANHUI AGRICULTURAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510189501.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional remote sensing image segmentation methods are difficult to effectively capture complex features of images, resulting in inaccurate identification of details, and are mostly limited to spatial domain processing, and insufficient capture of global information.

Method used

Using a remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion, a feature extraction network including dual-domain feature enhancement module and an adaptive scale fusion module are constructed, multi-scale feature extraction and fusion, and the fused feature map is obtained for semantic segmentation.

Benefits of technology

Effectively capturing the global information and detailed features of the image improves the accuracy and robustness of semantic segmentation, especially in complex structures and small object recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219734A_ABST
    Figure CN120219734A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on double-domain feature enhancement and adaptive scale fusion. The method comprises the following steps: S1, acquiring remote sensing image data; s2, constructing a feature extraction network comprising a double-domain feature enhancement module, and performing multi-scale feature extraction on the input remote sensing image to obtain feature maps of different levels; s3, constructing an adaptive scale fusion module, and performing feature fusion on the extracted multi-scale feature map to obtain a fused feature map; and S4, further processing based on the fused feature map to generate a final semantic segmentation result map. The invention also discloses a remote sensing image segmentation system based on double-domain feature enhancement and adaptive scale fusion. According to the method, the feature extraction network and the adaptive scale feature fusion network are linked, fusion of different semantic information feature maps is realized, and the detail analysis capability of complex remote sensing images is improved while high precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion. Background Art

[0002] In recent years, with the rapid development of remote sensing technology, the quality and coverage of remote sensing images have been significantly improved. These images contain rich semantic features and complex spatial distributions. These high-quality images are widely used in remote sensing semantic segmentation tasks, that is, accurately segmenting and classifying each pixel in a remote sensing image to support land cover classification, urban planning and management, environmental monitoring, geological exploration, etc. However, traditional semantic segmentation methods usually rely on handcrafted features or simple feature extraction techniques, making it difficult to effectively capture complex features in images, resulting in inaccurate recognition of details. At the same time, most of these methods are limited to spatial domain processing and also have deficiencies in capturing global information.

[0003] Therefore, there is an urgent need to provide a new remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion to solve the above problems. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion, which can effectively capture the global information and detail features of images.

[0005] To solve the above technical problem, one technical solution adopted by the present invention is: to provide a remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion, including the following steps:

[0006] S1: Obtain remote sensing image data;

[0007] S2: Construct a feature extraction network including a dual-domain feature enhancement module, perform multi-scale feature extraction on the input remote sensing image, and obtain feature maps at different levels;

[0008] S3: Construct an adaptive scale fusion module, fuse multi-scale information through multi-scale convolution and combined attention mechanism, perform feature fusion on the extracted multi-scale feature maps, and obtain a fused feature map;

[0009] S4: Perform further processing based on the fused feature map to generate a final semantic segmentation result map.

[0010] In a preferred embodiment of the present invention, in step S2, the feature extraction network is composed of a ResNet-50 network combined with a dual-domain feature enhancement module. ResNet-50, as the backbone network, is composed of four stages of Resblocks. After each stage of Resblocks, a dual-domain feature fusion module is designed to perform frequency-domain hybrid enhancement on the feature maps extracted by the residual network blocks in the Resblocks, outputting four-scale feature maps, and the size and number of channels of each feature map are halved in turn.

[0011] In a preferred embodiment of the present invention, the dual-domain feature enhancement module is composed of a frequency-domain branch and a spatial-domain branch, which combines the extracted frequency-domain and spatial-domain information to obtain a comprehensive output feature, enhancing the diversity and richness of feature representation.

[0012] Furthermore, the frequency-domain branch is composed of a low-pass filter, a 1×1 convolution operator, and a ReLU activation function. The low-pass filter performs low-pass filtering on the frequency-domain features to suppress high-frequency information and enhance the global information in the features. Subsequently, basic convolution operations and a ReLU activation function are performed on the features. Convolution is used to fuse different frequency features, and the ReLU function enables the network to learn and express more complex features.

[0013] Furthermore, the specific construction steps of the spatial-domain branch include:

[0014] First, multi-level features are extracted through convolution operations of different scales; specifically, convolution kernels k(m)×k(m) of different sizes are applied in parallel to the features X of the input remote sensing image to generate multiple feature maps. The specific results are as follows:

[0015] F (m,1) =Conv k(m)×k(m) (X), m = 1, 3, 5

[0016] where m ∈ {1, 3, 5}, and m represents the scale size of the convolution kernel, which is used to capture semantic information of different scales;

[0017] Secondly, after obtaining the preliminary features, further convolution operations of the same scale are performed on them. The specific results are as follows, to obtain deeper features;

[0018] F (m,2) =Conv k(m)×k(m) (F (m,1) )

[0019] Finally, after adding the deep feature maps of each scale element-wise by channel, they are fused through 1×1 convolution to obtain the final spatial branch feature X c , and the specific results are as follows:

[0020]

[0021] In a preferred embodiment of the present invention, it is assumed that the outputs of the feature extraction network are the first feature map, the second feature map, the third feature map, and the fourth feature map respectively. The adaptive scale fusion module first fuses the fourth feature map with the third feature map to obtain a fifth feature map; then fuses the fifth feature map with the second feature map to obtain a sixth feature map; and finally fuses the sixth feature map with the first feature map to obtain a seventh feature map.

[0022] Furthermore, the adaptive scale fusion module includes three branches: an upper branch, a middle branch, and a lower branch;

[0023] Assume that the scale feature of the third feature map is f m , and the size is R ∈ c×h×w . For the upper branch, global average pooling and global max pooling are performed on f m to generate two global description features, which are concatenated in the channel dimension. The feature map is converted into a feature map with a size of R ∈ c×1×1 through a 1×1 convolutional layer. Subsequently, it passes through a BN layer, a ReLU layer, and a 3×3 convolutional layer, and an attention weight vector is generated through a Sigmoid layer. The weight vector and the feature f m are subjected to a Hadamard product, and the weighted feature is output. The result is as follows:

[0024]

[0025] where σ represents the sigmoid function, Cat represents the concatenation operation, is the Hadamard product operation, W S represents the ReLU activation function, Conv1 is the 1×1 convolutional layer, Conv3 is the 3×3 convolutional layer, GAP(·) represents global average pooling, and GVP(·) represents global max pooling;

[0026] Assume that the scale map feature of the fourth feature map is f s , and the size is . For the middle branch, the feature f s is upsampled and aligned with f m . Subsequently, the two are added element-wise to obtain a feature map with a size of R ∈ c×h×w . Then, the feature is converted into a feature map with a size of through a 3×3 convolutional layer and a ReLU layer. Three convolutional branches with different convolutional kernel sizes (3, 5, 7) are designed, and the results of the three branches are added element-wise to obtain the final feature with a size of R ∈ c×h×w . The result is as follows:

[0027]

[0028] Among them, up represents the upsampling operation, and W S represents the ReLU activation function, and Conv m represents the m×m convolutional layer, and Conv1 is the 1×1 convolutional layer;

[0029] In the lower branch, the 3×3 convolutional layer is used to convert f s into a feature map with a size of Subsequently, it sequentially passes through the BN layer, the ReLU layer, and the 5×5 convolutional layer, and then passes through the Sigmoid layer to generate the attention weight map; f s is subjected to the Hadamard product with the attention weight map to obtain the enhanced feature map, and the upsampling operation is performed on the enhanced feature map to obtain the feature The result is as follows:

[0030]

[0031] Among them, σ represents the Sigmoid function, up represents the upsampling operation, and W S represents the ReLU activation function, Conv5 is the 5×5 convolutional layer, Conv3 is the 3×3 convolutional layer, is the Hadamard product operation.

[0032] Furthermore, the adaptive scale fusion module further includes generating an enhanced feature map for the original feature f m using the 3×3 convolutional layer, the BN layer, and the ReLU layer generating a feature map for the original feature f s using the 1×1 convolutional layer, the BN layer, and the ReLU layer Combining the enhanced feature and by element-wise addition to obtain the final output feature map R(f) of the adaptive scale fusion module. The result is as follows:

[0033]

[0034] Among them, represents the element-wise addition operation.

[0035] In a preferred embodiment of the present invention, in step S4, the final feature map output by the adaptive scale fusion module passes through the classification head to generate the final semantic segmentation result map.

[0036] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a remote sensing image segmentation system based on dual-domain feature enhancement and adaptive scale fusion, which adopts the remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion described in any one of the above, including:

[0037] A remote sensing image acquisition module, used to acquire remote sensing image data;

[0038] A multi-scale feature map acquisition module, used to construct a feature extraction network including a dual-domain feature enhancement module, perform multi-scale feature extraction on the remote sensing image input by the remote sensing image acquisition module, and obtain feature maps at different levels;

[0039] A feature fusion module, used to construct an adaptive scale fusion module, fuse multi-scale information through multi-scale convolution and combined attention mechanism, perform feature fusion on the extracted multi-scale feature maps, and obtain a fused feature map;

[0040] A semantic segmentation result map output module, used to perform further processing based on the fused feature map of the feature fusion module to generate a final semantic segmentation result map.

[0041] The beneficial effects of the present invention are:

[0042] (1) Aiming at the problem of insufficient ability of the model to handle edge blurring in the remote sensing image segmentation task, the present invention proposes an efficient and innovative solution. By designing a new semantic segmentation model based on a dual-domain feature enhancement module, combining the feature information in the frequency domain and the spatial domain, it effectively captures the global information and detailed features of the image. Among them, the frequency domain captures the global structure, and the spatial domain strengthens the local details, making the model perform better in the segmentation and recognition of complex structures; compared with the traditional convolution block, this module more accurately separates different elements in the image, reduces edge blurring, and thus significantly improves the overall recognition accuracy;

[0043] (2) Aiming at the problem of neglecting small target classification in remote sensing image segmentation, the present invention constructs an adaptive scale fusion module to integrate multi-scale features. Through a multi-scale and multi-level feature fusion method, the model can learn more general feature patterns and can capture more comprehensive information, including global context and local details. This fusion method can effectively enhance the robustness of the model to targets at different scales and different perspectives, and further improve the classification ability of small targets;

[0044] (3) By linking the feature extraction network and the adaptive scale feature fusion network, the present invention realizes the fusion of feature maps of different semantic information. While ensuring high accuracy, it improves the detail analysis ability of complex remote sensing images and adapts to the large-scale data processing requirements under diverse scenarios. Description of the Drawings

[0045] Figure 1 is the flowchart of the remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion of the present invention;

[0046] Figure 2 is the architecture diagram of the remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion;

[0047] Figure 3 is the architecture diagram of the dual-domain feature enhancement module;

[0048] Figure 4 is the architecture diagram of the adaptive scale fusion module;

[0049] Figure 5 is the structure diagram of the remote sensing image segmentation system based on dual-domain feature enhancement and adaptive scale fusion. Detailed implementation manners

[0050] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0051] Please refer to Figure 1 and Figure 2 , the embodiments of the present invention include:

[0052] A remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion, comprising the following steps:

[0053] S1: Obtain remote sensing image data;

[0054] In this example, the publicly available ISPRS Potsdam and ISPRS Vaihingen remote sensing image datasets are used as the training and test datasets, and the training set and test set are divided according to a ratio of 7:3.

[0055] S2: Construct a feature extraction network including a dual-domain feature enhancement module, perform multi-scale feature extraction on the input remote sensing image, and obtain feature maps at different levels;

[0056] Among them, the feature extraction network is composed of a ResNet-50 network combined with a dual-domain feature enhancement module to perform multi-scale feature extraction on the input remote sensing image. Specifically, ResNet-50, as the backbone network, consists of four stages of Resblocks. After each stage of Resblocks, a dual-domain feature fusion module is designed to perform frequency-domain hybrid enhancement on the feature maps extracted by the residual network blocks in the Resblocks, outputting four-scale feature maps, and the size and number of channels of each feature map are halved in turn.

[0057] The dual-domain feature enhancement module refers to Figure 3 , and is composed of a frequency-domain branch and a spatial-domain branch. The frequency-domain branch consists of a low-pass filter, a 1×1 convolutional operator, and a ReLU activation function. In the frequency-domain branch, the features are converted from the spatial domain to the frequency domain form. The low-pass filter performs low-pass filtering on the frequency-domain features to suppress high-frequency information and enhance the global information in the features. Subsequently, basic convolution operations and the ReLU activation function are applied to the features. Convolution is used to fuse different frequency features to ensure that the features in each channel can be effectively integrated in the frequency domain. This operation helps to reduce feature redundancy and extract the most useful frequency-domain features; the ReLU function enables the network to learn and express more complex features.

[0058] The specific results are as follows:

[0059] X f = F -1 [W Low (W Linear (W R (W Linear (F(X)))))]

[0060] where F(·) represents the Fourier transform, F -1 (·) represents the inverse Fourier transform, W Linear represents the 1×1 convolution operation, W R represents the ReLU activation function, W Low represents the low-pass filter, X f represents the features finally output by the frequency-domain branch.

[0061] The dual-domain feature enhancement module also designs a spatial-domain branch for combining spatial-domain information. The specific construction steps are as follows:

[0062] First, multi-level features are extracted through convolution operations of different scales; specifically, convolution kernels of different sizes k(m)×k(m) are applied in parallel to the input features X to generate multiple feature maps. The specific results are as follows:

[0063] F (m,1) = Conv k(m)×k(m) (X), m = 1, 3, 5

[0064] where m ∈ {1, 3, 5}, and m represents the scale size of the convolution kernel for capturing semantic information of different scales.

[0065] After obtaining the preliminary features, further convolution operations of the same scale are performed on them. The specific results are as follows, and deeper features are obtained.

[0066] F (m,2)= Conv k(m)×k(m) (F (m,1) )

[0067] To comprehensively integrate the feature information at different scales, after adding the deep feature maps at each scale element-wise by channel, they are fused through 1×1 convolution to obtain the final spatial branch feature X c , and the specific results are as follows:

[0068]

[0069] Finally, the features extracted from the frequency domain branch and the spatial domain branch are added element-wise to obtain the final comprehensive output feature representation R(x); the results are as follows:

[0070]

[0071] S3: Construct an adaptive scale fusion module to fuse multi-scale information through multi-scale convolution and combined attention mechanism, and perform feature fusion on the extracted multi-scale feature maps to obtain the fused feature maps;

[0072] Refer to Figure 4 , the adaptive scale fusion module includes upper, middle, and lower branches, aiming to capture useful features at different scales and perform effective fusion, which is used to fuse the four different-scale feature maps output by the feature extraction network, realizing the effective integration and enhancement of multi-scale features, and integrating channel and spatial information.

[0073] Furthermore, the output of the feature extraction network includes four different-scale feature maps from shallow to deep, representing different levels of information in the image: shallow features mainly capture local detail information such as edges and textures, while deep features are more focused on global semantic information and abstract representations. To make full use of this different-level feature information and combine the advantages of multi-scale features, the adaptive scale fusion module performs adaptive fusion on them. Combining Figure 2 , assuming that the outputs of the feature extraction network are the first feature map, the second feature map, the third feature map, and the fourth feature map respectively, the adaptive scale fusion module first fuses the fourth feature map with the third feature map to obtain the fifth feature map; then fuses the fifth feature map with the second feature map to obtain the sixth feature map; finally fuses the sixth feature map with the first feature map to obtain the seventh feature map.

[0074] The specific construction steps of the adaptive scale fusion module are as follows:

[0075] First, assume that the input of the module for a larger-scale feature (such as the third feature map) is f m , with a size of R ∈ c×h×w, for the upper branch, for f m perform global average pooling and global max pooling to generate two global descriptive features, and perform concatenation on the channel dimension. Convert the feature map into a feature map of size R ∈ c×1×1 through a 1×1 convolutional layer. Subsequently, pass through a BN layer, a ReLU layer, and a 3×3 convolutional layer in sequence, and generate an attention weight vector through a Sigmoid layer. Multiply the weight vector with the feature f m using the Hadamard product to output the weighted feature The results are shown as follows:

[0076]

[0077] where σ represents the sigmoid function, Cat represents the concatenation operation, is the Hadamard product operation, W S represents the ReLU activation function, Conv1 is the 1×1 convolutional layer, Conv3 is the 3×3 convolutional layer, GAP(·) represents global average pooling, and GVP(·) represents global max pooling.

[0078] Furthermore, assume that the input small-scale map feature (e.g., the fourth feature map) is f s , with a size of For the middle branch, upsample the feature f s to align it with f m . Subsequently, add the two element-wise to obtain a feature map of size R ∈ c×h×w . Then, convert the feature to a feature map of size through a 3×3 convolutional layer and a ReLU layer. To capture multi-scale information, three convolutional branches with different convolutional kernel sizes (3, 5, 7) are designed, and the results of the three-way branches are added element-wise to obtain the final feature of size R ∈ c×h×w The results are shown as follows:

[0079]

[0080] where up (upsample) represents the upsampling operation, W S represents the ReLU activation function, Conv m represents the m×m convolutional layer, and Conv1 is the 1×1 convolutional layer.

[0081] Furthermore, in the lower branch, use a 3×3 convolutional layer to convert f s into a feature map of size . Subsequently, pass through a BN layer, a ReLU layer, and a 5×5 convolutional layer in sequence, and then pass through a Sigmoid layer to generate an attention weight map; for f s ​Perform Hadamard product with the attention weight map to obtain the enhanced feature map, and perform upsampling operation on the enhanced feature map to obtain the feature The results are as follows:

[0082]

[0083] where σ represents the Sigmoid function, up represents the upsampling operation, and W S represents the ReLU activation function, Conv5 is a 5×5 convolutional layer, Conv3 is a 3×3 convolutional layer, is the Hadamard product operation.

[0084] Furthermore, for the original feature f m Use a 3×3 convolutional layer, BN layer, and ReLU layer to generate an enhanced feature map For the original feature f s Use a 1×1 convolutional layer, BN layer, and ReLU layer to generate a feature map Use element-wise addition to combine the enhanced feature and to obtain the final output feature map R(f) of the module. The results are as follows:

[0085]

[0086] where represents the element-wise addition operation.

[0087] In the above process, performing conventional convolution, batch normalization, and ReLU activation function operations on the two original features is to further extract the effective information in the original features, making the features more compact during fusion. Finally, adding the five enhanced features element-wise, this step fuses local details and global context information, realizing the unified processing of cross-scale features, thereby improving the overall performance of semantic segmentation, especially the segmentation accuracy in complex scenes.

[0088] S4: Further process based on the fused feature map to generate the final semantic segmentation result map.

[0089] The final feature map output by the adaptive scale fusion module passes through the classification head to generate the final semantic segmentation result map. The classification head includes a ReLU layer, a BN layer, and multiple convolutional layers, which are used to generate the class probability of each pixel point to obtain an accurate semantic segmentation result map.

[0090] Therefore, the present invention proposes a new semantic segmentation model that combines the advantages of frequency domain information and convolutional local features to improve the segmentation quality of various remote sensing images. At the same time, considering that the local and global information of images may not be fully utilized, in order to fully utilize the features obtained at different stages, the present invention also designs an adaptive scale fusion module. This module uses an attention mechanism to adaptively fuse the semantic information between different scale features, enabling the model to effectively classify various targets. The adaptive scale fusion module focuses on the semantic segmentation of remote sensing images. By fusing multi-scale information through multi-scale convolution and combining the attention mechanism, and at the same time performing different operations on the original features and fusing them, the model realizes the adaptive processing of the input data from multiple aspects such as channels, space, and feature fusion, so as to better adapt to the remote sensing image segmentation tasks with different scenarios and data distributions, and can better capture the complex features and detailed information in complex scenes.

[0091] In addition, the present invention also includes the training and testing of the model. Specifically, after building a remote sensing image semantic segmentation model based on dual-domain feature enhancement and dual attention feature fusion, the training and evaluation of the model are carried out. The training is performed on a GeForce RTX 3080Ti GPU, equipped with an Intel i7 processor, 32G of memory, and running the Windows system. During the training process, the Adam optimizer is used for gradient update, and parameters such as the learning rate, weight decay, and momentum are set. The loss function uses cross-entropy loss to guide the optimization process of the model parameters and improve the generalization ability of the model. The training set is trained for 300 rounds, and the training results of the model are observed.

[0092] Refer to Figure 5 In addition, in the example of the present invention, a remote sensing image segmentation system based on dual-domain feature enhancement and adaptive scale fusion is also provided, including a remote sensing image acquisition module, a multi-scale feature map acquisition module, a feature fusion module, and a semantic segmentation result map output module.

[0093] The remote sensing image acquisition module is used to acquire remote sensing image data;

[0094] The multi-scale feature map acquisition module is used to construct a feature extraction network including a dual-domain feature enhancement module, perform multi-scale feature extraction on the remote sensing images input by the remote sensing image acquisition module, and obtain feature maps at different levels;

[0095] The feature fusion module is used to construct an adaptive scale fusion module, fuse multi-scale information through multi-scale convolution and combining the attention mechanism, perform feature fusion on the extracted multi-scale feature maps, and obtain the fused feature maps;

[0096] The semantic segmentation result map output module is used to further process the feature map fused by the feature fusion module to generate a final semantic segmentation result map.

[0097] A remote sensing image segmentation system based on dual-domain feature enhancement and adaptive scale fusion in this example can execute a remote sensing image segmentation method provided by the present invention, can execute any combination of implementation steps of the method example, and has the corresponding functions and beneficial effects of the method.

[0098] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion, characterized in that: The following steps are involved: S1: Acquire remote sensing image data; S2: Construct a feature extraction network containing a dual-domain feature enhancement module to extract multi-scale features from the input remote sensing image and obtain different levels of S3: construct an adaptive scale fusion module, fuse multi-scale information through multi-scale convolution and combined attention mechanism, perform feature fusion on the extracted multi-scale feature map, and obtain a fused feature map; S4: Further processing is performed based on the fused feature map to generate the final semantic segmentation result map.

2. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 1, characterized in that: In step S2, the feature extraction network is composed of a ResNet-50 network combined with a dual-domain feature enhancement module. ResNet-50 is used as the backbone network and consists of four stages of Resblocks. A dual-domain feature fusion module is designed after each stage of Resblocks to perform frequency domain mixed enhancement on the feature maps extracted by the residual network blocks in Resblocks, and output feature maps of four scales, in which the size and number of channels of each feature map are halved successively.

3. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 1, characterized in that: The dual-domain feature enhancement module is composed of a frequency domain branch and a space domain branch, and combines the extracted frequency domain and space domain information to obtain comprehensive output features, thereby improving the diversity and richness of feature representation.

4. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 3, characterized in that: The frequency domain branch is composed of a low-pass filter, a 1×1 convolution operator and a ReLU activation function. The low-pass filter performs low-pass filtering on the frequency domain features to suppress high-frequency information and enhance the global information in the features. The features are then subjected to basic convolution operations and ReLU activation functions. Convolution is used to fuse different frequency features. The ReLU function enables the network to learn and express more complex features.

5. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 3, characterized in that: The specific steps of constructing the spatial domain branch include: First, multi-level features are extracted through convolution operations of different scales. Specifically, convolution kernels of different sizes k(m)×k(m) are applied in parallel to the feature X of the input remote sensing image to generate multiple feature maps. The specific results are as follows: F (m,1) =Conv k(m)×k(m) (X),m=1,3,5 Among them, m∈{1,3,5}, m represents the scale of the convolution kernel, which is used to capture semantic information of different scales; Secondly, after obtaining the preliminary features, we further perform convolution operations on them at the same scale. The specific results are as follows, and deeper features are obtained; F (m,2) =Conv k(m)×k(m) (F (m,1) ) Finally, the deep feature maps of each scale are added element by element according to the channel and fused through 1×1 convolution to obtain the final spatial branch feature X c , the specific results are as follows:

6. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 1, characterized in that: Assuming that the outputs of the feature extraction network are respectively a first feature map, a second feature map, a third feature map, and a fourth feature map, the adaptive scale fusion module first fuses the fourth feature map with the third feature map to obtain a fifth feature map; The fifth feature map is then fused with the second feature map to obtain a sixth feature map; finally, the sixth feature map is fused with the first feature map to obtain a seventh feature map.

7. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 6, characterized in that: The adaptive scale fusion module includes three branches: an upper branch, a middle branch and a lower branch; Assume that the scale feature of the third feature map is f m , size is R∈ c×h×w , for the upper branch, for f m Perform global average pooling and global maximum pooling to generate two global description features, and perform splicing on the channel dimension. The feature map is converted into a size of R∈ through a 1×1 convolution layer. c×1×1 The feature map is then passed through the BN layer, ReLU layer and 3×3 convolution layer in sequence, and the attention weight vector is generated through the Sigmoid layer. The weight vector is combined with the feature f m Perform Hadamard product and output weighted features The result is as follows: Where σ represents the sigmoid function, Cat represents the concatenation operation, is the Hadamard product operation, W S represents the ReLU activation function, Conv1 is a 1×1 convolutional layer, Conv3 is a 3×3 convolutional layer, GAP(·) represents global average pooling, and GVP(·) represents global maximum pooling; Assume that the fourth feature map scale map feature is f s , the size is For the middle branch, for the feature f s Upsampling, with f m Align, then add the two element by element to get a size of R∈ c×h×w The feature map is then transformed into a size of We design three convolution branches with different convolution kernel sizes (3, 5, 7) and add the three branch results element by element to obtain a convolution branch of size R∈ c×h×w The final feature The result is as follows: Where up represents the upsampling operation, W S ReLU activation function, Conv m represents an m×m convolutional layer, Conv1 is a 1×1 convolutional layer; In the lower branch, a 3×3 convolutional layer is used to transform f s Convert to size The feature map is then passed through the BN layer, ReLU layer and 5×5 convolution layer in sequence, and then through the Sigmoid layer to generate the attention weight map; f s Perform Hadamard product with the attention weight map to obtain the enhanced feature map, and perform upsampling operation on the enhanced feature map to obtain the feature The result is as follows: Where σ represents the Sigmoid function, up represents the upsampling operation, and W S ReLU activation function, Conv5 is a 5×5 convolution layer, Conv3 is a 3×3 convolution layer, is the Hadamard product operation.

8. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 7, characterized in that: The adaptive scale fusion module also includes the original feature f m Use 3×3 convolutional layer, BN layer, and ReLU layer to generate enhanced feature maps For the original feature f s Use 1×1 convolution layer, BN layer, and ReLU layer to generate feature maps Use element-wise addition to enhance the features and Combined, the final output feature map R(f) of the adaptive scale fusion module is obtained, and the result is as follows: in, Represents an element-by-element addition operation.

9. The remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion according to claim 1, characterized in that: In step S4, the final feature map output by the adaptive scale fusion module passes through the classification head to generate the final semantic segmentation result map.

10. A remote sensing image segmentation system based on dual-domain feature enhancement and adaptive scale fusion, using the remote sensing image segmentation method based on dual-domain feature enhancement and adaptive scale fusion as claimed in any one of claims 1 to 9, characterized in that: include: A remote sensing image acquisition module is used to acquire remote sensing image data; A multi-scale feature map acquisition module is used to construct a feature extraction network including a dual-domain feature enhancement module, perform multi-scale feature extraction on the remote sensing image input by the remote sensing image acquisition module, and obtain feature maps at different levels; A feature fusion module is used to construct an adaptive scale fusion module, fuse multi-scale information through multi-scale convolution and combined attention mechanism, perform feature fusion on the extracted multi-scale feature map, and obtain a fused feature map; The semantic segmentation result map output module is used to further process the feature map fused by the feature fusion module to generate a final semantic segmentation result map.

Citation Information

Cited By

  • Ankle joint ultrasonic image enhancement method

    CN121437300A

  • Ankle joint ultrasound image enhancement method

    CN121437300B

  • Remote sensing image segmentation method and system based on fine screening double-domain attention mechanism

    CN121582565A