Remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment

Through the wavelet-Mamba global feature enhancement module and the fast Fourier detail adjustment unit, the problem of balancing global semantic modeling and local detail perception in remote sensing image segmentation is solved, and high-precision segmentation of complex remote sensing images is achieved, especially fine-grained recognition of buildings, roads, water bodies and other land objects.

CN120765933AActive Publication Date: 2025-10-10耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202510868053.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-10
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing remote sensing image segmentation methods have difficulty balancing global semantic modeling and local detail perception when processing high-resolution remote sensing images, especially in complex scenes where the accuracy of small target recognition and fuzzy boundary segmentation is insufficient.

Method used

The wavelet-Mamba global feature enhancement module and the fast Fourier detail adjustment unit are used. Through the multi-level information enhancement mechanism in the frequency domain and spatial domain, the discrete wavelet transform and the fast Fourier transform are combined to model the structure, edge and texture information of the image respectively, and feature reconstruction is achieved through frequency-space interaction modeling.

Benefits of technology

It significantly improves the ability to express the structure, texture and edge information of objects in remote sensing images, and achieves high-precision segmentation of small targets and fuzzy boundaries in complex scenes, especially the fine-grained segmentation accuracy of key objects such as buildings, roads, and water bodies in high-resolution remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765933A_ABST
    Figure CN120765933A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment, and belongs to the technical field of remote sensing image processing. The method comprises the following steps: constructing an image segmentation model comprising a wavelet-Mama global feature enhancement module, a fast Fourier detail adjustment unit and a decoding and segmentation prediction module; performing remote sensing image segmentation training on the built image segmentation model; and performing image segmentation on the target remote sensing image by using the trained image segmentation model. According to the invention, through the wavelet-Mama global feature enhancement module and the fast Fourier detail adjustment unit, the expression ability of surface feature structures, textures and edge information in remote sensing images can be effectively improved, and high-precision segmentation of small targets and fuzzy boundaries in complex scenes is realized. The method is especially suitable for accurate recognition of buildings, roads, water bodies and other targets under high-resolution remote sensing images, and has high practical value and popularization prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment. Background Art

[0002] With the rapid development of remote sensing technology and the widespread adoption of high-resolution imaging systems, remote sensing image segmentation, as a key step in extracting ground object information, plays a vital role in urban planning, land use monitoring, disaster assessment, agricultural monitoring, and other fields. The goal of remote sensing image segmentation is to divide image pixels into semantically meaningful regions, thereby enabling the automatic identification and extraction of ground objects such as buildings, water bodies, roads, and farmland. Due to the high resolution, complex content, blurred ground object boundaries, and dramatic scale variations of remote sensing images, traditional image segmentation methods face significant challenges in this field.

[0003] Currently, mainstream remote sensing image segmentation methods are mostly based on deep convolutional neural networks (CNNs) or Transformer architectures, leveraging their powerful representational capabilities to perform semantic modeling and pixel classification. However, while CNNs excel at extracting local spatial features, they still struggle to capture large-scale structures, long-range dependencies, and the continuity of object boundaries in remote sensing images. While Transformers possess global modeling capabilities, their high computational overhead and limited ability to capture spatial details in remote sensing image processing limit their practical application in high-resolution remote sensing image segmentation.

[0004] Furthermore, existing methods are prone to problems such as blurred boundaries, missed detection of small objects, and missing texture information when dealing with fine-grained features in remote sensing images (such as narrow roads, rivers, canals, and roof ridges). This is especially true in areas with complex scenes and small inter-class differences, making it difficult to further improve segmentation accuracy. Therefore, balancing global semantic modeling and local detail perception in remote sensing images has become a key issue that urgently needs to be addressed in remote sensing image segmentation technology. Summary of the Invention

[0005] In view of this, the present invention provides a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment, which aims to integrate the multi-level information enhancement mechanism of frequency domain and spatial domain. Through the wavelet-Mamba global feature enhancement module and the fast Fourier detail adjustment unit, it can effectively improve the expression ability of the ground structure, texture and edge information in remote sensing images, and achieve high-precision segmentation of small targets and fuzzy boundaries in complex scenes.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, an embodiment of the present invention provides a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment, the method mainly comprising the following steps:

[0008] S1. Build an image segmentation model including a wavelet-Mamba global feature enhancement module, a fast Fourier detail adjustment unit, and a decoding and segmentation prediction module;

[0009] S2. Performing remote sensing image segmentation training on the constructed image segmentation model;

[0010] S3. Use the trained image segmentation model to perform image segmentation on the target remote sensing image.

[0011] Furthermore, in the wavelet-Mamba global feature enhancement module, the discrete wavelet transform is combined to perform multi-frequency domain decomposition of the input features to model the image structure, edge, and texture information respectively. The channel-level Mamba mechanism is also introduced to implement global dependency modeling. The specific process includes:

[0012] Normalize the input features to obtain normalized features

[0013] The normalized features are transformed using two-dimensional discrete wavelet transform Decomposed into four frequency domain subbands, the expression is:

[0014]

[0015] in, It is the low-frequency-low-frequency subband, which contains the overall structure and texture information; is the low-frequency-high-frequency subband, representing the horizontal edge; is the high-low frequency subband, representing the vertical edge; is the high-frequency-high-frequency subband, representing the diagonal edge; DWT represents two-dimensional discrete wavelet transform;

[0016] The low-frequency-low-frequency subband is enhanced by applying a convolution layer, an activation function, a convolution layer, an activation function, a normalization layer, a Mamba module, and a convolution layer; the low-frequency-high-frequency subband, the high-frequency-low-frequency subband, and the high-frequency-high-frequency subband are enhanced only by the convolution layer;

[0017] The four sub-bands are fused and restored to the enhanced features through two-dimensional discrete wavelet inverse transform.

[0018] The enhanced features With normalized features The output features are obtained by fusion through residual connection.

[0019] Furthermore, the Mamba module is used to perform global dependency modeling on feature maps in the channel dimension. Key operations include:

[0020] Normalization: used to stabilize training and prevent gradient disappearance or explosion;

[0021] State-space modeling: using state-space equations to enhance channel information flow;

[0022] Linear mapping and gating mechanism: Enhance the selective control of features, making the inter-class differences in remote sensing images more obvious.

[0023] Furthermore, in the image segmentation model, a multi-level wavelet-Mamba global feature enhancement module is stacked, and after multiple downsampling and enhancement, deep semantic features are extracted.

[0024] Furthermore, in the fast Fourier detail adjustment unit, a frequency domain amplitude-phase dual-channel enhancement strategy is adopted, dynamic convolution and dilated convolution are introduced to enhance the structure and edges respectively, and feature reconstruction is achieved through frequency-space interaction modeling. The specific process includes:

[0025] Deep semantic features As input, a two-dimensional fast Fourier transform is applied to map it into the frequency domain to obtain the amplitude spectrum and phase spectrum;

[0026] The amplitude spectrum is enhanced by applying full-dimensional dynamic convolution, activation functions, and convolutional layers, while the phase spectrum is enhanced by applying convolutional layers, activation functions, and dilated convolutional layers.

[0027] The enhanced amplitude spectrum is fused with the phase spectrum, and the inverse fast Fourier transform is applied to restore it to the spatial domain to obtain the enhanced detail feature map:

[0028] Apply a depthwise separable convolutional layer to the enhanced detail feature map to obtain detail adjustment features:

[0029] Combine detail adjustment features with deep semantic features Perform addition fusion to obtain fusion features Fusion Features After two branches, the upper branch contains linear layer, convolution layer, and Sigmoid activation function to obtain features The lower branch obtains features through the linear layer and ReLU activation function The two branches are then fused and passed through a linear layer to obtain the output features of frequency domain detail adjustment

[0030] Furthermore, in the decoding and segmentation prediction module, the output features of the frequency domain details adjustment are The upsampling operation is performed through a series of transposed convolution or bilinear interpolation and convolution to restore the spatial resolution, and the semantic information is restored through the decoder to generate the final remote sensing image segmentation result.

[0031] Furthermore, when the constructed image segmentation model is trained for remote sensing image segmentation, a joint loss function is used for training optimization. The joint loss function includes cross entropy loss and boundary perception loss, and its expression is:

[0032]

[0033] Where, represents the joint loss; λ1, λ2 represent the weight coefficients; represents the cross entropy loss; represents the boundary-aware loss; Represents the true label of the (i, j)th pixel in class c; Represents the corresponding predicted probability value; H, W represent the height and width of the image respectively, and C represents the number of channels; represents a binary edge map; Represents the prediction boundary map.

[0034] In a second aspect, the present invention also provides an electronic device comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment.

[0035] Compared with the prior art, the present invention has at least the following beneficial effects:

[0036] 1) The present invention provides a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment. By applying the wavelet-Mamba global feature enhancement module and the fast Fourier detail adjustment unit, etc., the expression ability of the structure, texture and edge information of the ground objects in the remote sensing image can be effectively improved, and high-precision segmentation of small targets and fuzzy boundaries in complex scenes can be achieved.

[0037] 2) The wavelet-Mamba global feature enhancement module proposed in this paper combines discrete wavelet transform with multi-frequency domain decomposition of features, which can model the structure, edge and texture information of the image respectively. It also introduces the channel-level Mamba mechanism to realize global dependency modeling, significantly improving the ability to express complex landforms and spatial hierarchical relationships in remote sensing images, and effectively compensating for the shortcomings of traditional convolutional structures in extracting long-range dependencies and detailed features.

[0038] 3) The fast Fourier detail adjustment unit constructed in the present invention adopts a frequency domain amplitude-phase dual-channel enhancement strategy, introduces dynamic convolution and dilated convolution to enhance the structure and edges respectively, and realizes feature reconstruction through frequency-space interaction modeling, thereby effectively improving the segmentation model's ability to recognize small targets and fuzzy edges, and enhancing the fine-grained segmentation accuracy of key landforms such as buildings, roads, and water bodies in remote sensing images.

[0039] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0040] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0043] Figure 1 A flowchart of a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment is provided in an embodiment of the present invention.

[0044] Figure 2 A schematic diagram of the working principles of the feature extraction and wavelet Mamba global enhancement module provided in an embodiment of the present invention.

[0045] Figure 3 A schematic diagram of the working principle of the fast Fourier detail adjustment unit provided in an embodiment of the present invention.

[0046] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0048] In the description of the present application, it should be noted that in some of the processes described in the specification and drawings, a plurality of operations appear in a specific order, but it should be clearly understood that these operations can be executed or in parallel without following the order in which they appear in this text. In addition, various serial numbers are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0049] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0050] Referring to Figure 1 As shown, the present application provides a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment, which mainly includes the following steps:

[0051] S1, build an image segmentation model including a wavelet-Mamba global feature enhancement module, a fast Fourier detail adjustment unit, a decoding and segmentation prediction module;

[0052] S2, train the built image segmentation model for remote sensing image segmentation;

[0053] S3, use the trained image segmentation model to perform image segmentation on the target remote sensing image.

[0054] The working principle and specific implementation of the method of the present application will be described in detail below Figure 2-Figure 3 As shown, the working principle and specific implementation of the method of the present application will be described in detail below

[0055] I. Feature extraction and wavelet Mamba global feature enhancement:

[0056] The module structure and workflow are as shown in Figure 2 Specifically, it includes:

[0057] 1.1: Remote sensing image preprocessing and feature extraction:

[0058] First, the remote sensing image is subjected to a series of preprocessing steps (including noise suppression, geometric correction, image registration, etc.) to obtain an image I input to the neural network. In order to effectively extract multi-scale spatial features, the image I is processed in turn through two deep separable convolution layers: a 3x3 deep separable convolution layer is used to capture local spatial structure information, and another 1x1 deep separable convolution layer is used for channel compression and fusion. The output features of the two convolution layers are denoted as:

[0059]

[0060] DSConv stands for Depthwise Separable Convolutional Layer. The two layers are fused element-wise to enhance feature representation capability:

[0061]

[0062] The fused features are then spatially downsampled (e.g., using maximum pooling or convolution with a stride of 2) to obtain a higher level of semantic representation and obtain a preliminary feature map.

[0063] 1.2: Wavelet-Mamba global feature enhancement module:

[0064] In order to improve the model's ability to perceive global details such as texture and edges in remote sensing images, a global feature enhancement module based on wavelet transform is introduced, and the Mamba mechanism is combined for channel enhancement.

[0065] First, the input features Perform normalization to obtain features

[0066] Then, the feature map is transformed into Decomposed into four frequency domain subbands. The present invention adopts the Haar wavelet which is simple in calculation and suitable for high-resolution image processing, and its low-pass filter L and high-pass filter H are defined as follows:

[0067]

[0068] The input feature map is decomposed into four sub-bands through horizontal and vertical filtering operations:

[0069] Low-Low, which contains overall structure and texture information;

[0070] Low-High, representing horizontal edges;

[0071] High-Low, representing vertical edges;

[0072] High-High frequency (High-High) represents diagonal edges.

[0073] The wavelet transform process can be expressed as:

[0074]

[0075] Next, the low-frequency-low-frequency subband The application convolution layer, activation function, convolution layer, activation function, normalization layer, channel Mamba module (detailed in section 1.3 below) and convolution layer; while the low-frequency-high-frequency subband, high-frequency-low-frequency subband, high-frequency-high-frequency subband only through the convolution layer.

[0076] Then, the four subbands are fused and restored to the enhanced feature map

[0077]

[0078] Finally, the enhanced feature and the original normalized feature are fused through residual connection to obtain the final output.

[0079] 1.3: Channel-level Mamba module:

[0080] In order to further enhance the semantic modeling capability in remote sensing images, a channel-level Mamba module is introduced to model the global dependence of the feature map in the channel dimension. Traditional convolution mainly focuses on local regions, while the Mamba module, as an efficient modeling method similar to Transformer, can effectively capture long-range dependencies and is suitable for identification and segmentation of complex object boundaries in remote sensing images.

[0081] Suppose the input is the low-frequency subband feature map after wavelet processing The present application first rearranges and normalizes the channel dimension, and then passes through the Mamba module, which is expressed as follows:

[0082]

[0083] The module mainly includes the following key operations:

[0084] Normalization layer: used to stabilize training and prevent gradient vanishing or explosion;

[0085] State space modeling: use state space equation to enhance channel information flow;

[0086] Linear mapping and gating mechanism: enhance the selective control of features, so that the inter-class differences in remote sensing images are more obvious.

[0087] The final output feature is further processed by a convolution layer Conv and a normalization layer Norm to be fused with other subbands, expressed as:

[0088] 1.4: Stack multi-level wavelet-Mamba feature enhancement module:

[0089] To gradually extract hierarchical semantic information from remote sensing images, this method cascades the aforementioned wavelet-Mamba global feature enhancement modules, and through continuous downsampling, develops deep multi-scale feature abstraction capabilities.

[0090] Specifically, the feature map The first level module is input, and the output is used as the input for the next level module. Each level module contains a downsampling and a set of wavelet-Mamba enhancement processes. This process is iterated 4 times:

[0091]

[0092] in, WaveletMamba(·) represents the complete enhancement module structure. After four downsampling and enhancement modules, the final deep semantic features (multi-scale semantic features) extracted are:

[0093]

[0094] This feature will serve as the input of the subsequent fast Fourier detail adjustment unit for accurate segmentation of ground object categories in remote sensing images.

[0095] 2. Fast Fourier detail adjustment unit:

[0096] In order to further improve the ability to express the edges and structural details of objects in remote sensing images, a fast Fourier detail adjustment unit based on frequency domain processing is proposed. This unit effectively improves the model's detail perception ability in high-resolution remote sensing image segmentation by combining frequency domain enhancement with spatial reconstruction. The structure and workflow of the fast Fourier detail adjustment unit are as follows: Figure 3 As shown, the details are as follows:

[0097] 2.1: Fast Fourier transform to obtain frequency domain information:

[0098] The deep semantic features finally output in step 1 above As input, a two-dimensional Fast Fourier Transform (FFT) is applied to map it to the frequency domain to obtain the amplitude spectrum (Amplitude) and phase spectrum (Phase):

[0099]

[0100] In the Fast Fourier Transform (FFT), an image is transformed from the spatial domain (usually row and column pixel coordinates (x, y)) to the frequency domain. The transformed image coordinates are the frequency coordinates (u, v): where u represents the horizontal frequency component and v represents the vertical frequency component. These two coordinates represent different frequency "components" in the frequency domain, that is, the periodic changing pattern in the image. FFT2(·) represents the two-dimensional Fast Fourier Transform. Represents the complex coefficients on the frequency coordinate (u, v); Represents Euler's formula, which represents the polar coordinate form of a complex number; Represents frequency domain amplitude information, the intensity of frequency components, and mainly includes the structure and texture of the image; It represents the frequency domain phase information and the position offset of the frequency components, mainly retaining the edge and position of the image.

[0101] 2.2: Amplitude channel enhancement path:

[0102] amplitude The branch first enhances its adaptive expression capabilities through a set of Omni-Dimensional Dynamic Convolution (ODConv) modules. Compared to conventional convolution, ODConv can dynamically adjust the convolution kernel parameters in the spatial dimension, channel dimension, and kernel size dimension, making it suitable for processing complex remote sensing image structures. The amplitude processing process is as follows:

[0103]

[0104] ReLU represents the ReLU activation function, and Conv represents the convolutional layer. This path enhances image texture and structural information, improving the model's ability to perceive the terrain's morphology.

[0105] 2.3: Phase channel enhancement path:

[0106] Phase The branch is used to preserve image edges and object location information. This path is designed as follows: first passing through a standard convolution layer, then entering an activation function and a dilated convolution layer (DConv) to expand the receptive field, thereby obtaining richer edge context. The phase processing process is as follows:

[0107]

[0108] Dilated convolution can effectively capture sparse but important boundary features in remote sensing images, such as roads, water edges, building outlines, etc.

[0109] 2.4: Frequency domain fusion and spatial domain reconstruction:

[0110] The enhanced amplitude spectrum and phase spectrum After fusion, the inverse fast Fourier transform (IFFT) is applied to restore it to the spatial domain to obtain the enhanced detail feature map:

[0111]

[0112] Subsequently, to further improve the spatial feature consistency and edge continuity, a 3×3 depthwise separable convolutional layer (DWConv) is applied to obtain the final detail adjustment features:

[0113]

[0114] in, Represents the final detail adjustment feature.

[0115] 2.5: and Perform addition fusion to obtain features feature After two branches, the upper branch contains linear layer, convolution layer, and Sigmoid activation function. The lower branch is obtained by the linear layer and ReLU activation function The two branches are fused and passed through a linear layer to obtain the output features of frequency domain detail adjustment.

[0116] This invention effectively addresses the difficulty in identifying small targets and blurred edges in remote sensing images through dual-channel amplitude-phase enhancement and interactive reconstruction in the frequency-space domain. It provides high-precision, fine-grained feature support for subsequent segmentation tasks. This module effectively compensates for the limited ability of conventional spatial convolution to depict local details. By modeling and reconstructing frequency-domain information, it significantly improves the segmentation accuracy of complex terrain areas in remote sensing images. It is particularly suitable for identifying fine-grained targets such as buildings, roads, and water systems in high-resolution remote sensing images.

[0117] 3. Decoding and segmentation prediction module:

[0118] After frequency domain detail enhancement, the output features Finally, a lightweight decoder structure is constructed to restore high-level semantic features to spatial resolution and generate the final remote sensing image segmentation results. This module includes sub-processes such as step-by-step upsampling, edge-guided fusion, semantic restoration, and supervised optimization, aiming to achieve accurate segmentation of ground objects in remote sensing images.

[0119] 3.1: Feature upsampling and edge information guidance:

[0120] In order to gradually restore the spatial detail information, the features First, the up-sampling operation is performed by a series of transposed convolution or bilinear interpolation and convolution, gradually recovering to the original image resolution. Let the output of the up-sampling layer be Considering that the edge information in the remote sensing image is crucial for small target and complex structure recognition, an auxiliary edge guiding branch is introduced. The branch uses an edge detection module (such as Sobel convolution or Laplacian convolution) to process the low-level features to obtain an edge response map, which is spliced and fused with the up-sampled features to obtain The fusion operation helps to enhance the positioning accuracy of the segmentation boundary, especially in areas such as building edges and road intersections.

[0121] 3.2: Semantic prediction and multi-scale fusion:

[0122] Fusion features After entering the decoder main body, the semantic information is restored through a number of convolution layers (such as 3x3 standard convolution + BN + ReLU) to obtain a class prediction probability map for each pixel:

[0123]

[0124] wherein, is the prediction output of the final segmentation result.

[0125] 3.3: Constructing a joint loss function for training optimization:

[0126] In order to fully consider the problems of class imbalance and boundary ambiguity in remote sensing images, a combined loss function is designed in the present invention for end-to-end optimization, including cross-entropy loss and boundary-aware loss:

[0127] ① Cross-entropy loss:

[0128] Used for pixel-level classification optimization, defined as follows:

[0129]

[0130] wherein, is the true label of the (i, j)th pixel in the cth class; is the corresponding prediction probability value.

[0131] ② Boundary-aware loss:

[0132] Guides the model to pay attention to the boundaries of ground objects and improves edge accuracy. The binary edge map and the predicted boundary map between the binary cross-entropy (BCE) loss:

[0133]

[0134] ③Total loss function:

[0135] Combining the two sub-losses, the total loss function is constructed as:

[0136]

[0137] Among them, λ1 and λ2 are weight coefficients that balance the contribution of the main segmentation task and the boundary enhancement task.

[0138] Finally, the trained image segmentation model can be used to efficiently and accurately segment the target remote sensing image.

[0139] 3.4: Segmentation result output:

[0140] Finally, in the inference stage, the category corresponding to the maximum predicted probability is taken as the final segmentation category label to generate a pixel-level remote sensing object segmentation map.

[0141]

[0142] Among them, H and W are the height and width of the input image.

[0143] This module achieves high-precision, fine-grained segmentation of complex features (such as buildings, roads, water bodies, and farmland) in remote sensing images through the combined effects of semantic restoration, edge guidance, and loss constraints. Combined with the Wavelet-Mamba global feature enhancement module in step one and the Fast Fourier Detail Adjustment Unit (frequency domain enhancement) in step two, it effectively balances global semantic understanding with local detail recognition, significantly improving the overall performance of remote sensing image segmentation.

[0144] From the description of the above embodiments, those skilled in the art can know that the present invention proposes a remote sensing image segmentation method based on wavelet-Mamba global feature enhancement and Fourier detail adjustment. This method comprehensively utilizes the advantages of frequency domain analysis and spatial modeling, optimizes the feature expression capability from the perspective of multi-scale, multi-channel, and multi-branch, and improves the accuracy of the model's analysis and recognition of complex land object structures in remote sensing images. Specifically, the present invention first designs a wavelet-Mamba global enhancement module, combines the two-dimensional discrete wavelet transform (DWT) with the lightweight Mamba mechanism, and realizes the joint modeling of edge, texture and semantic information in remote sensing images in the feature extraction stage. The different frequency sub-bands obtained by wavelet transform can respectively enhance the contour boundary and overall structural characteristics of the land object, while the Mamba module introduces long-distance dependency modeling capabilities in the channel dimension, further improving the discriminability and robustness of feature expression.

[0145] To further enhance the model's ability to perceive spatial details, the application also proposes a fast Fourier detail adjustment unit. Based on frequency domain information, the unit performs Fourier transform (FFT) on deep semantic features to obtain amplitude and phase spectra, and performs feature enhancement on the amplitude and phase paths through Omni-Dimensional Dynamic Convolution (ODConv) and dilated convolution (DConv), respectively, and then restores them to a detail enhancement map through inverse Fourier transform (IFFT). This process can effectively restore the edges of ground objects, fine-grained targets and texture structures, making up for the lack of expression ability of spatial convolution for small targets.

[0146] In the model decoding stage, the application introduces a lightweight semantic restoration structure and an edge-guided branch to realize the step-by-step restoration of high-level semantic features through multi-level upsampling and edge enhancement fusion, further improving the accuracy and boundary coherence of ground object class segmentation. In addition, to address the problems of class imbalance and boundary ambiguity in remote sensing images, the application also designs a joint optimization strategy combining pixel-level cross-entropy loss and boundary perception loss, effectively promoting the collaborative learning of the model between semantic expression and edge perception.

[0147] In summary, the application introduces innovative designs in key links of remote sensing image segmentation (feature extraction, detail enhancement, decoding prediction), builds a unified processing framework that takes into account spatial structure and frequency domain information, significantly improves the segmentation accuracy of fine-grained ground objects in remote sensing images, and is particularly suitable for accurate identification of buildings, roads, water bodies and other targets in high-resolution remote sensing images, with high practical value and promotion prospects.

[0148] In addition, as shown in Figure 4 The electronic device can include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and can also include a computer program stored in the memory 11 and executable on the processor 10, and the processor executes the computer program to implement a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to one of the above method embodiments.

[0149] In some embodiments, the processor 10 can be composed of integrated circuits, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple packaged integrated circuits with the same function or different functions, including one or more combinations of central processors, microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core of the electronic device, which connects all components of the electronic device through various interfaces and lines, and executes programs or modules stored in the memory 11 and calls data stored in the memory 11 to perform various functions and process data of the electronic device.

[0150] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, electronic devices, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0151] It should be noted that the word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer.

[0152] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0153] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment, characterized in that: The method comprises the following steps: S1. Build an image segmentation model including a wavelet-Mamba global feature enhancement module, a fast Fourier detail adjustment unit, and a decoding and segmentation prediction module; S2. Performing remote sensing image segmentation training on the constructed image segmentation model; S3. Use the trained image segmentation model to perform image segmentation on the target remote sensing image.

2. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 1, characterized in that: In the wavelet-Mamba global feature enhancement module, the discrete wavelet transform is combined to perform multi-frequency domain decomposition of the input features to model the image structure, edge, and texture information respectively. The channel-level Mamba mechanism is also introduced to implement global dependency modeling. The specific process includes: Normalize the input features to obtain normalized features The normalized features are transformed using two-dimensional discrete wavelet transform Decomposed into four frequency domain subbands, the expression is: in, It is the low-frequency-low-frequency subband, which contains the overall structure and texture information; is the low-frequency-high-frequency subband, representing the horizontal edge; is the high-low frequency subband, representing the vertical edge; is the high-frequency-high-frequency subband, representing the diagonal edge; DWT represents two-dimensional discrete wavelet transform; The low-frequency-low-frequency subband is enhanced by applying a convolution layer, an activation function, a convolution layer, an activation function, a normalization layer, a Mamba module, and a convolution layer; the low-frequency-high-frequency subband, the high-frequency-low-frequency subband, and the high-frequency-high-frequency subband are enhanced only by the convolution layer; The four sub-bands are fused and restored to the enhanced features through two-dimensional discrete wavelet inverse transform. The enhanced features With normalized features The output features are obtained by fusion through residual connection.

3. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 2, characterized in that: The Mamba module is used to perform global dependency modeling on feature maps in the channel dimension. Key operations include: Normalization: used to stabilize training and prevent gradient disappearance or explosion; State-space modeling: using state-space equations to enhance channel information flow; Linear mapping and gating mechanism: Enhance the selective control of features, making the inter-class differences in remote sensing images more obvious.

4. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 2, characterized in that: In the image segmentation model, a multi-level wavelet-Mamba global feature enhancement module is stacked, and after multiple downsampling and enhancement, deep semantic features are extracted.

5. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 4, characterized in that: In the fast Fourier detail adjustment unit, a frequency domain amplitude-phase dual-channel enhancement strategy is adopted. Dynamic convolution and dilated convolution are introduced to enhance the structure and edges respectively. Feature reconstruction is achieved through frequency-space interaction modeling. The specific process includes: Deep semantic features As input, a two-dimensional fast Fourier transform is applied to map it into the frequency domain to obtain the amplitude spectrum and phase spectrum; The amplitude spectrum is enhanced by applying full-dimensional dynamic convolution, activation functions, and convolutional layers, while the phase spectrum is enhanced by applying convolutional layers, activation functions, and dilated convolutional layers. The enhanced amplitude spectrum is fused with the phase spectrum, and the inverse fast Fourier transform is applied to restore it to the spatial domain to obtain the enhanced detail feature map: Apply a depthwise separable convolutional layer to the enhanced detail feature map to obtain detail adjustment features: Combine detail adjustment features with deep semantic features Perform addition fusion to obtain fusion features Fusion Features After two branches, the upper branch contains linear layer, convolution layer, and Sigmoid activation function to obtain features The lower branch obtains features through the linear layer and ReLU activation function The two branches are then fused and passed through a linear layer to obtain the output features of frequency domain detail adjustment 6. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 5, characterized in that: In the decoding and segmentation prediction module, the output features of the frequency domain details adjustment are The upsampling operation is performed through a series of transposed convolution or bilinear interpolation and convolution to restore the spatial resolution, and the semantic information is restored through the decoder to generate the final remote sensing image segmentation result.

7. The remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment according to claim 1, characterized in that: When the constructed image segmentation model is trained for remote sensing image segmentation, a joint loss function is used for training optimization. The joint loss function includes cross entropy loss and boundary perception loss, and its expression is: Where, represents the joint loss; λ1, λ2 represent the weight coefficients; represents the cross entropy loss; represents the boundary-aware loss; Represents the true label of the (i, j)th pixel in class c; Represents the corresponding predicted probability value; H, W represent the height and width of the image respectively, and C represents the number of channels; represents a binary edge map; Represents the prediction boundary map.

8. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement a remote sensing image segmentation method based on global feature enhancement and Fourier detail adjustment as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method combining Unet and Transform

    CN116091929A

  • Instance segmentation method and device for small road target in remote sensing image

    CN119027667A

  • Ultrahigh-resolution remote sensing image segmentation method based on frequency domain information fusion

    CN119559200A

  • Underwater image enhancement method based on frequency domain analysis and visual Mama

    CN119784598A

  • Real scene remote sensing image super-resolution reconstruction method and system based on progressive feature aggregation network, and storage medium

    CN120013762A

Cited By

  • Remote sensing image enhancement processing method and device for soil pollution observation

    CN120852261A

  • Lightweight target detection Transform model based on space-frequency domain joint modeling, method and application

    CN121280870A

  • Rock automatic extraction method and system based on deep learning and Mars rover camera image

    CN121685999A

  • Remote sensing image segmentation method based on light visual scanning and frequency domain discrimination feedforward

    CN121999216A