Remote sensing image segmentation method based on multilayer wavelet transform and dynamic memory network

Through the combination of multi-layer wavelet transformation and dynamic memory network, efficient feature extraction and global local feature fusion of remote sensing images are achieved, solving the problems of local details loss and insufficient global information in remote sensing image segmentation, and improving the accuracy and robustness of segmentation.

CN120260043AActive Publication Date: 2025-07-04耕宇牧星(北京)空间科技有限公司

Patent Information

Application Number
CN202510316817.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

When the existing remote sensing image segmentation method deals with complex backgrounds, low contrast and multi-scale targets, there are local details loss, insufficient global information, large model parameters and high computational complexity, making it difficult to achieve high-precision and robust segmentation.

Method used

The multi-layer wavelet transformation and dynamic memory network are used to decompose the image into low-frequency subbands and high-frequency subbands through multi-layer Haar wavelet transformation, feature extraction is performed by combining convolution operations, and a dynamic memory unit network is constructed to fusion globally and local features, and finally image segmentation is performed.

Benefits of technology

It improves the accuracy and robustness of remote sensing image segmentation, can effectively capture multi-scale features, retain detailed information of complex land objects, enhances the ability to distinguish complex land objects such as urban buildings and rural roads, and improves the accuracy and overall consistency of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260043A_ABST
    Figure CN120260043A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a remote sensing image segmentation method based on multilayer wavelet transform and a dynamic memory network, which comprises the following steps: decomposing an image into a low-frequency sub-band and a plurality of high-frequency sub-bands through multilayer Haar wavelet transform, and performing feature extraction on each frequency band in combination with convolution operation to obtain a final feature map; constructing a dynamic memory unit network, performing dynamic feature extraction on the feature map by adopting a query-key-value structure, generating a dynamic memory enhanced feature with global context information, performing global-local feature modeling on the feature map, and fusing the dynamic memory enhanced feature and the global-local feature to obtain a fused feature map; and segmenting the fused feature map to obtain a final segmentation result. According to the method, low-frequency smooth information and high-frequency detail information of the image can be accurately extracted, efficient fusion of global and local features can be realized, and the accuracy and robustness of remote sensing image segmentation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to a remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network. Background Art

[0002] With the rapid development of remote sensing technology, remote sensing images play an increasingly important role in fields such as land resource monitoring, urban planning, environmental protection, and disaster assessment. Due to the wide acquisition range and large amount of information, remote sensing images often contain rich ground object details and complex multi-scale features. Therefore, how to accurately extract target information from massive data has become a key issue. Traditional image segmentation methods often perform poorly when dealing with complex backgrounds, low contrast, and multi-scale targets. Although the segmentation methods based on deep learning have improved this situation to a certain extent in recent years, there are still deficiencies such as loss of local details, insufficient global information, large model parameter quantities, and high computational complexity.

[0003] Therefore, how to improve the accuracy and robustness of remote sensing image segmentation is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the present invention provides a remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network, which can accurately extract the low-frequency smooth information and high-frequency detail information of the image, and can achieve the efficient fusion of global and local features, thereby improving the accuracy and robustness of remote sensing image segmentation.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network, comprising the following steps:

[0007] S1. Feature extraction is performed on the input remote sensing image based on a pre-trained Mamba model, and the image is decomposed into a low-frequency sub-band and multiple high-frequency sub-bands through multi-layer Haar wavelet transform. Then, feature extraction is performed on each frequency band in combination with convolution operations to obtain the final feature map φ o ;

[0008] S2. A dynamic memory unit network is constructed, adopting a query-key-value structure, to perform dynamic feature extraction on the feature map φ o to generate a dynamic memory enhanced feature with global context information, and to perform global-local feature modeling on the feature map φ o and fuse the dynamic memory enhanced feature and the global-local feature to obtain a fused feature map φ + ;

[0009] S3. Perform operations on the fused feature map φ +Perform segmentation to obtain the final segmentation result.

[0010] Furthermore, S1 includes:

[0011] S11. Extract features from the input remote sensing image based on the pre-trained Mamba model to obtain the initial feature map φ;

[0012] S12. Perform a one-layer Haar wavelet transform on the initial feature map φ, and then perform convolution operations through depth convolution operations in combination with a low-pass filter and a high-pass filter to obtain a low-frequency sub-band and three high-frequency sub-bands;

[0013] S13. Perform multi-layer wavelet transforms on the low-frequency sub-band, and use the same convolution operation to extract low-frequency features and high-frequency features for each layer of transformation;

[0014] S14. After multi-layer wavelet transforms, use inverse wavelet transforms to reconstruct each sub-band back into the low-frequency and high-frequency components of the original image;

[0015] S15. Combine wavelet transforms and convolution operations to extract features from the transformation results of each layer to obtain the final feature map φ o .

[0016] Furthermore, in S12, a low-frequency sub-band and three high-frequency sub-bands are represented as:

[0017] [φ L , φ LH , φ HL , φ HH = Conv(f LL , f LH , f HL , f HH , φ)

[0018] where φ L represents the low-frequency sub-band, containing the smooth information of the image; φ LH represents the horizontal high-frequency sub-band, containing the horizontal edge information of the image; φ HL represents the vertical high-frequency sub-band, containing the vertical edge information of the image; φ HH represents the diagonal high-frequency sub-band, containing the diagonal edge information of the image; the resolution of each sub-band is half of the original image; f LL represents the low-pass filter, f LH represents the horizontal high-frequency filter, f HL represents the vertical high-frequency filter, f HH represents the diagonal high-frequency filter; Conv represents the convolution operation.

[0019] Furthermore, in S13, each layer of wavelet transform is represented as:

[0020]

[0021] Among them, represents the low-frequency subband of the k-th layer, and the feature information of the image is extracted layer by layer until the predetermined number of layers is reached; DWT k represents the wavelet transform operation of the k-th layer. The decomposition of each layer generates new subbands, including LL k , LH k , HL k and HH k , which respectively contain information of the image at different frequencies.

[0022] Furthermore, S2 includes:

[0023] S21. Construct key memory units and value memory units through two independent and parameter-sharing linear layers. The two memory units are used to extract and store deep feature information across regions and scales in the remote sensing image;

[0024] S22. Extract the query weight Q from the feature map φ o . Adopt a dynamic query mechanism to match the query weight Q with the stored features in the two memory units to generate dynamic memory-enhanced features with global context information;

[0025] S23. Use adaptive average pooling and depthwise separable convolution to model the local information of the feature map φ o to obtain local features;

[0026] S24. Calculate the learnable weight of the global features of the feature map φ o , and multiply the learnable weight by the local features to obtain global-local features;

[0027] S25. Perform weighted summation on the global-local features and the dynamic memory-enhanced features to obtain the fused feature map φ + .

[0028] Furthermore, in S22, the calculation formula of the query weight Q is:

[0029] Q = W q * Norm(Conv(φ o )) + b q

[0030] Among them, Conv(·) is responsible for local feature extraction; Norm represents the normalization operation; W q and b q are learnable parameters to ensure that the model can automatically focus on the key areas;

[0031] The dynamic memory-enhanced features are expressed as:

[0032]

[0033] Among them, φ MU represents the dynamic memory enhancement feature, softmax represents the normalized exponential function, Q represents the query weight, K represents the key, V represents the value, and d k represents the dimension of the key.

[0034] Furthermore, in S23, the local feature is represented as:

[0035] φ DW = DWConv(AAP(φ o ))

[0036] Among them, DWConv represents the depthwise separable convolution, AAP represents the adaptive average pooling layer, and φ DW represents the local feature.

[0037] Furthermore, in S24, the learnable weight is represented as:

[0038]

[0039] Among them, P represents the position embedding operation, which is used to supplement the geographical location information; FC represents the fully connected layer, and Sigmoid represents the Sigmoid activation function; As the learnable weight, it can dynamically adjust the contribution degree of the global feature;

[0040] The global-local feature is represented as:

[0041]

[0042] Among them, represents the matrix multiplication operation.

[0043] Furthermore, in S25, the fused feature map φ + is represented as:

[0044] φ + = λ1φ MU + λ2φ GL

[0045] Among them, λ1 and λ2 are learnable parameters, which can adaptively adjust the proportion of global and local information.

[0046] Furthermore, S3 includes:

[0047] S31. Perform dimensionality reduction and semantic feature extraction on the fused feature map φ + to obtain the dimension-reduced feature map F;

[0048] S32. Use bilinear interpolation upsampling operation to restore the downsampled feature map F to the original spatial size, and then score the category of each pixel in the upsampled feature.

[0049] S33. Perform softmax operation on the category scores of each pixel to calculate the probability of each pixel in each category, and perform argmax operation on the probability distribution of each pixel to determine its predicted category, obtaining the segmentation result.

[0050] S34. Use post-processing method to refine the segmentation boundary of the segmentation result and smooth the edge of the target area.

[0051] As can be seen from the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. When the present invention extracts features from remote sensing images, it first uses the pre-trained Mamba model to perform preliminary feature extraction on the remote sensing images to obtain a primary feature map with certain semantic information. Considering that remote sensing images contain both smooth low-frequency information (such as large areas of water bodies and farmland) and rich high-frequency details (such as roads and building edges), multi-layer Haar wavelet transform is used to decompose the image. By performing wavelet transform on the image, the image can be decomposed into a low-frequency sub-band and multiple high-frequency sub-bands, which can effectively capture the low-frequency smooth information and high-frequency detail features in the remote sensing image at different scales, thus providing a richer and more accurate feature expression for subsequent image segmentation. In addition, by combining convolution operations to further process each frequency band, the expression ability of the feature map can be greatly improved, so that the detail information of complex ground objects can be fully retained, greatly improving the accuracy and robustness of the segmentation result, and providing a more accurate feature basis for subsequent segmentation tasks. The strategy of combining multi-layer wavelet transform and convolution in the present invention enables the model to take into account both global and local information when facing large-scale changes and fine textures, and realizes the full utilization of multi-scale features.

[0053] 2. The present invention makes full use of the synergistic advantages of multi-scale features and dynamic memory mechanism. By constructing a dynamic memory unit network based on the query-key-value structure, effective fusion of global and local features is achieved through an adaptive dynamic query mechanism. It can not only capture long-range dependencies and cross-scale information in the image, but also adaptively adjust the contribution degrees of the global background and local details, significantly enhancing the discrimination ability for complex ground objects (such as urban buildings and rural roads). Compared with traditional methods, this innovative strategy makes up for the deficiencies of the prior art in capturing multi-frequency, multi-scale information and complex structure modeling in remote sensing images, so that the segmentation result has significant improvements in both detail restoration and overall coherence. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0055] Figure 1 It is a flowchart of the remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network provided by the present invention;

[0056] Figure 2 It is a flowchart of S1 of the present invention;

[0057] Figure 3 It is a flowchart of S2 of the present invention. Detailed implementation manners

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0059] As Figure 1 shown, the embodiments of the present invention disclose a remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network, including the following steps:

[0060] S1. Extract features from the input remote sensing image based on the pre-trained Mamba model, decompose the image into a low-frequency sub-band and multiple high-frequency sub-bands through multi-layer Haar wavelet transform, and then perform feature extraction on each frequency band in combination with convolution operations to obtain the final feature map φ o ;

[0061] S2. Construct a dynamic memory unit network, adopt a query-key-value structure, perform dynamic feature extraction on the feature map φ o to generate dynamic memory enhanced features with global context information, and perform global-local feature modeling on the feature map φ o to fuse the dynamic memory enhanced features and the global-local features to obtain the fused feature map φ + ;

[0062] S3. Segment the fused feature map φ + to obtain the final segmentation result.

[0063] Next, the above steps will be further described.

[0064] S1. Extract features from the input remote sensing image based on the pre-trained Mamba model, decompose the image into a low-frequency subband and multiple high-frequency subbands through multi-layer Haar wavelet transform, and then perform feature extraction on each frequency band by combining convolution operations to obtain the final feature map φ o ; As Figure 2 shown, S1 specifically includes:

[0065] S11. Preliminary feature extraction: Extract features from the input remote sensing image based on the pre-trained Mamba model to obtain the initial feature map φ, where H and W are the height and width of the image respectively, and C is the number of channels of the feature map.

[0066] S12. Haar wavelet transform: Perform one-layer Haar wavelet transform on the initial feature map , and then perform convolution operations through depth convolution operations combined with low-pass and high-pass filters to obtain a low-frequency subband and three high-frequency subbands.

[0067] Operation of one-layer wavelet transform: In each dimension (e.g., width or height), perform depth convolution using the following convolution kernels:

[0068] Low-pass filter f LL :

[0069] Horizontal high-frequency filter f LH :

[0070] Vertical high-frequency filter f HL :

[0071] Diagonal high-frequency filter f HH :

[0072] After performing the convolution operation, through downsampling with a stride of 2, the image is decomposed into four subbands, including a low-frequency subband and three high-frequency subbands. It can be expressed as:

[0073] [φ L , φ LH , φ HL , φ HH = Conv(f LL , f LH , f HL , f HH , φ)

[0074] where φ L represents the low-frequency subband, containing the smooth information of the image; φ LH represents the horizontal high-frequency subband, containing the horizontal edge information of the image; φHL represents the vertical high-frequency subband, which contains the vertical edge information of the image; φ HH represents the diagonal high-frequency subband, which contains the diagonal edge information of the image; the resolution of each subband is half of the original image; f LL represents the low-pass filter, f LH represents the horizontal high-frequency filter, f HL represents the vertical high-frequency filter, f HH represents the diagonal high-frequency filter; Conv represents the convolution operation.

[0075] S13. Perform multi-level wavelet transform on the low-frequency subband to obtain finer feature information. The same convolution operation is used for each layer of transformation to extract low-frequency features and high-frequency features. Each layer of wavelet transform is expressed as:

[0076]

[0077] where represents the low-frequency subband of the k-th layer, and the feature information of the image is extracted layer by layer until the predetermined number of layers is reached; DWT k represents the wavelet transform operation of the k-th layer. The decomposition of each layer generates new subbands, including LL k , LH k , HL k and HH k , which contain the information of the image at different frequencies respectively.

[0078] S14. Reconstruct the image using the inverse wavelet transform: After multi-level wavelet transform, the inverse wavelet transform (IWT) is used to reconstruct each subband back into the low-frequency component and high-frequency component of the original image; the inverse wavelet transform can be realized by transposed convolution:

[0079] φ = Trans([f LL , f LH , f HL , f HH , [φ L , φ LH , φ HL , φ HH )

[0080] This operation restores the four subbands to the form of the original image. Through the inverse wavelet transform, an image with different frequency information can be obtained, thus providing support for subsequent feature extraction.

[0081] S15. To improve the expression ability of the feature map and reduce the number of parameters, combine wavelet transform and convolution operation, perform feature extraction on the transformation results of each layer, and use a convolution kernel to further extract features to obtain the final feature map φ o . The formula is:

[0082] φ o = IWT(Conv(W, WT(φ)))

[0083] where W is a k×k convolutional kernel. Generally, depth convolution operations are considered, and the number of input channels is four times that of the image. φ o is the final output feature map. This method separates different frequency components through convolution operations, increases the receptive field, and at the same time maintains computational efficiency. The convolution operations for each frequency band help further extract detailed information.

[0084] S2. Construct a dynamic memory unit network, the structure of which is as Figure 3 shown. Adopt a query-key-value structure to perform dynamic feature extraction on the feature map φ o to generate dynamic memory-enhanced features with global context information, and perform global-local feature modeling on the feature map φ o to fuse the dynamic memory-enhanced features and the global-local features to obtain the fused feature map φ + . S2 specifically includes:

[0085] S21. Memory unit construction: Construct a key memory unit and a value memory unit through two independent and parameter-sharing linear layers. The two memory units are used to extract and store deep feature information across regions and scales in the remote sensing image; to enhance the long-term memory ability of the model, DMUNet uses the key memory unit and the value memory unit to store and retrieve deep feature information in the remote sensing image. The key memory unit and the value memory unit are implemented through two independent linear layers. The parameters of these linear layers are learnable and shared across the entire dataset, which ensures the consistency of the model among different samples and reduces the computational complexity at the same time. At the initialization of the model, the parameters of the key memory unit and the value memory unit are randomly generated, and their output dimensions (i.e., the memory unit size) are controlled by the hyperparameter memory_size, which is usually set to a relatively small value (such as 64 or 128) to prevent overfitting and improve computational efficiency. During the training process, the parameters of the key memory unit and the value memory unit are updated through backpropagation and an optimizer (such as Adam or SGD).

[0086] S22. Dynamic feature query: Adopt a query-key-value (QKV) structure to store and dynamically extract feature information, enabling the network to efficiently utilize deep feature information across scales and regions.

[0087] Extract the query weight Q from the feature map φ o to match the global features stored in the memory unit. The calculation formula for the query weight Q is:

[0088] Q = W q *Norm(Conv(φ o)) + b q

[0089] Among them, Conv(·) is responsible for local feature extraction; Norm represents the normalization operation; W q and b q are learnable parameters to ensure that the model can automatically focus on key regions.

[0090] Adopt a dynamic query mechanism to match the query weight Q with the stored features in the two memory units to generate dynamic memory-enhanced features with global context information; the dynamic memory-enhanced features are expressed as:

[0091]

[0092] Among them, φ MU represents the dynamic memory-enhanced feature, which can effectively supplement long-range context information and improve the segmentation accuracy of the target area in remote sensing images; softmax represents the normalization exponential function, Q represents the query weight, K represents the key, V represents the value, and d k represents the dimension of the key (Key), which is often used to scale the dot product attention mechanism to ensure numerical stability.

[0093] After that, combined with global-local feature modeling, the segmentation task of remote sensing images not only depends on global information modeling (such as background consistency, land cover classification), but also needs to capture local details (such as edges, texture features). For this reason, the present invention adopts full-dimensional dynamic convolution to improve the local feature expression ability, specifically as follows:

[0094] S23. Local feature extraction: Local information is crucial for remote sensing targets (such as roads, buildings, water bodies, etc.). The present invention uses adaptive average pooling and depthwise separable convolution to model the local information of the feature map φ o to obtain local features; the local features are expressed as:

[0095] φ DW = DWConv(AAP(φ o ))

[0096] Among them, DWConv represents depthwise separable convolution, AAP represents the adaptive average pooling layer, and φ DW represents local features. This convolution structure can reduce the computational complexity while retaining key details, making it easier for the model to distinguish small targets (such as ships, vehicles, etc.) in remote sensing images.

[0097] S24. Global feature extraction: To ensure that the model has global context information, the present invention introduces dynamic adaptive weight calculation:

[0098]

[0099] Among them, P represents the position embedding operation, which is used to supplement geographical location information; FC represents the fully connected layer, and Sigmoid represents the Sigmoid activation function; As learnable weights, they can dynamically adjust the contribution degree of global features.

[0100] After that, multiply the learnable weights by the local features to obtain the global-local features; the global-local features are expressed as:

[0101]

[0102] Among them, represents the matrix multiplication operation.

[0103] S25. Memory enhancement and global-local feature fusion: In remote sensing image segmentation, targets at different scales (such as urban buildings vs. rural roads) require different degrees of global background information and local details. Therefore, the present invention performs weighted fusion of memory-enhanced features, global features, and local features to obtain the fused feature map φ + . The fused feature map φ + is expressed as:

[0104] φ + =λ1φ MU +λ2φ GL

[0105] Among them, λ1 and λ2 are learnable parameters, which can adaptively adjust the proportion of global and local information, ensuring that the model maintains good generalization ability on different tasks and datasets. The final fused feature φ + has the following key advantages:

[0106] 1) Enhance the global perception ability (the dynamic memory unit stores cross-scale information and reduces class confusion).

[0107] 2) Strengthen the expression of detail features (the local convolution structure improves the edge detection ability).

[0108] 3) Improve the robustness (the multi-scale fusion strategy reduces the influence of illumination changes and noise interference on the segmentation effect).

[0109] After S2, the obtained fused feature φ + fully integrates global memory and local detail information and can accurately depict multi-scale ground objects in remote sensing images.

[0110] S3. Segment the fused feature map φ + to obtain the final segmentation result. S3 extracts from the fused feature φ +Generate the final pixel-level segmentation result. The whole process mainly includes feature map dimensionality reduction, upsampling prediction, normalization, and necessary post-processing optimization. Specifically, it includes:

[0111] S31. Feature map dimensionality reduction and semantic fusion: To convert the fused feature φ + into an intermediate representation suitable for the segmentation task, first perform dimensionality reduction and semantic information extraction on it. Specifically, use a 1×1 convolutional layer for channel fusion and further enhance the feature expression through a non-linear activation function:

[0112] F = f(Conv 1×1 (φ + ))

[0113] where Conv 1×1 represents the 1×1 convolution operation for dimensionality reduction and feature fusion; f(·) is the activation function (such as ReLU or GELU); F represents the dimensionality-reduced feature map. This dimensionality reduction process helps to extract high-level semantic information while reducing the computational complexity, laying the foundation for subsequent upsampling and segmentation prediction.

[0114] S32. Upsampling and segmentation prediction: Remote sensing images usually have a high resolution, so the dimensionality-reduced feature map F must be restored to the original spatial size to ensure the accuracy of the pixel-level segmentation result. Use the bilinear interpolation upsampling operation to expand F to the size matching the input image, and use the bilinear upsampling operation to make aligned with the height H and width W of the original image. Next, use another 1×1 convolutional layer to further process the upsampled feature to generate the class scores for each pixel point represents the original scores of each pixel over all classes, where C cls is the number of classes.

[0115] S33. Normalization and generation of segmentation result: To convert the original scores S of each pixel point class into an easily interpretable probability distribution, apply the softmax operation to S to calculate the probability of each pixel point over each class represents the pixel-level class probability distribution. Finally, by performing the argmax operation on the probability distribution of each pixel, determine its predicted class to obtain the final segmentation result is the final pixel-level segmentation map, and each pixel (i, j) is assigned a specific class label.

[0116] S34. Refine the segmentation boundaries of the segmentation result through post - processing and smooth the edges of the target area. For remote sensing image segmentation tasks, especially when dealing with the boundaries of complex ground objects, there may be noise or uneven edges in the initial segmentation result. Therefore, the following post - processing techniques can be used to further optimize the segmentation result: Conditional Random Field (CRF), which refines the segmentation boundaries and improves the overall segmentation coherence by leveraging the color and position similarities between pixels. Morphological operations, such as erosion and dilation operations, help to eliminate small noise and smooth the edges of the target area. These post - processing steps can be selectively used according to the actual application requirements to obtain a higher - quality segmentation map.

[0117] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method section.

[0118] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network, characterized in that It includes the following steps: S1. Extract features from the input remote sensing image based on the pre-trained Mamba model, decompose the image into a low-frequency sub-band and multiple high-frequency sub-bands through multi-layer Haar wavelet transform, and then perform feature extraction on each frequency band in combination with convolution operations to obtain the final feature map φ o ; S2. Construct a dynamic memory cell network, adopt a query-key-value structure, and perform dynamic feature extraction on the feature map φ o to generate dynamic memory-enhanced features with global context information, and perform global-local feature modeling on the feature map φ o Then fuse the dynamic memory-enhanced features and the global-local features to obtain a fused feature map φ + ; S3. Segment the fused feature map φ + to obtain the final segmentation result.

2. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 1, characterized in that, S1 includes: S11. Based on the pre-trained Mamba model, extract features from the input remote sensing image to obtain the initial feature map φ; S12. Perform a one-layer Haar wavelet transform on the initial feature map φ, and then perform a convolution operation through a depth convolution operation in combination with a low-pass filter and a high-pass filter to obtain a low-frequency sub-band and three high-frequency sub-bands; S13. Perform multi-layer wavelet transforms on the low-frequency sub-band, and use the same convolution operation for each layer of transformation to extract low-frequency features and high-frequency features; S14. After multi-layer wavelet transforms, use the inverse wavelet transform to reconstruct each sub-band back into the low-frequency component and high-frequency component of the original image; S15. Combine wavelet transform and convolution operation to extract features from the transformation results of each layer, and obtain the final feature map φ o .

3. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 2, wherein In S12, a low-frequency sub-band and three high-frequency sub-bands are expressed as: [φ L ,φ LH ,φ HL ,φ HH = Conv([f LL ,f LH ,f HL ,f HH ,φ) Among them, φ L represents the low-frequency subband, which contains the smooth information of the image; φ LH represents the horizontal high-frequency subband, which contains the horizontal edge information of the image; φ HL represents the vertical high-frequency subband, which contains the vertical edge information of the image; φ HH represents the diagonal high-frequency subband, which contains the diagonal edge information of the image; the resolution of each subband is half of the original image; f LL represents the low-pass filter, f LH represents the horizontal high-frequency filter, f HL represents the vertical high-frequency filter, f HH represents the diagonal high-frequency filter; Conv represents the convolution operation.

4. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 2, wherein In S13, each layer of wavelet transform is expressed as: Among them, represents the low-frequency sub-band of the k-th layer, and the feature information of the image is extracted layer by layer until the predetermined number of layers is reached; DWT k represents the wavelet transform operation of the k-th layer. The decomposition of each layer generates new sub-bands, including LL k , LH k , HL k and HH k , which respectively contain the information of the image at different frequencies.

5. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 1, wherein S2 It includes: S21. Construct a key memory unit and a value memory unit through two independent and parameter-sharing linear layers. The two memory units are used to extract and store the deep feature information across regions and scales in the remote sensing image; S22. Extract the query weight Q from the feature map φ o and adopt a dynamic query mechanism to match the query weight Q with the stored features in the two memory units to generate dynamic memory-enhanced features with global context information; S23. Model the local information of the feature map φ using adaptive average pooling and depthwise separable convolution to obtain local features; o ​ S24. Calculate the feature map φ o The learnable weights of the global features are multiplied by the local features to obtain the global-local features; S25. Perform a weighted summation of the global-local features and the dynamic memory-enhanced features to obtain the fused feature map φ + .

6. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 5, wherein In S22, the calculation formula for the query weight Q is: Q = Q q *Norm(Conv(φ o )) + b q Among them, Conv(·) is responsible for local feature extraction; Norm represents the normalization operation; W q and b q are learnable parameters to ensure that the model can automatically focus on key regions; The dynamic memory enhanced feature representation is: Among them, φ MU represents the dynamic memory enhancement feature, softmax represents the normalized exponential function, Q represents the query weight, K represents the key, V represents the value, and d k represents the dimension of the key.

7. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 5, characterized in that In S23, the local feature representation is: φ DW = DWConv(AAP(φ o )) Among them, DWConv represents depthwise separable convolution, AAP represents adaptive average pooling layer, and φ DW represents local features.

8. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 7, wherein, In S24, the learnable weight is expressed as: W = Sigmoid(FC(Conv(Conv(φ o )) + P)) Where, P represents the position embedding operation, which is used to supplement the geographical location information; FC represents the fully connected layer, and Sigmoid represents the Sigmoid activation function; W is the learnable weight, which can dynamically adjust the contribution degree of the global feature; The global-local feature representation is: Among them, represents a matrix multiplication operation.

9. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 8, characterized in that, In S25, the fused feature map φ + is expressed as: φ + = λ1φ MU + λ2φ GL Where, λ1 and λ2 are learnable parameters, which can adaptively adjust the proportion of global and local information.

10. The remote sensing image segmentation method based on multi-layer wavelet transform and dynamic memory network according to claim 1, characterized in that S3 It includes: S31. Downsample the fused feature map φ + and perform semantic feature extraction to obtain the downsampled feature map F; S32. Use the bilinear interpolation upsampling operation to restore the downsampled feature map F to the original spatial size, and then score the category of each pixel point in the upsampled feature; S33. Perform the softmax operation on the category scores of each pixel point, calculate the probability of each pixel point in each category, perform the argmax operation on the probability distribution of each pixel point, determine its predicted category, and obtain the segmentation result; S34. Use the post-processing method to refine the segmentation boundary of the segmentation result and smooth the edge of the target area.

Citation Information

Patent Citations

  • Video instance segmentation method of dynamic convolution solution of lightweight attention mechanism

    CN117115703A

  • Remote sensing image accurate segmentation method

    CN118504427A

  • High-resolution remote sensing image semantic segmentation method based on multi-scale depth supervision

    CN119559403A

  • Double-path medical image segmentation method and device based on multistage comprehensive attention

    CN119579887A

  • Medical image segmentation method based on combination of wavelet transform and SAM of bridging mode

    CN119888230A

Cited By

  • Remote sensing image semantic change detection method based on time sequence remote sensing Mama

    CN120580599A

  • Remote sensing image semantic change detection method based on time-series remote sensing mamba

    CN120580599B

  • Medical image segmentation method and system based on frequency context feature mixing

    CN121095560A

  • A Medical Image Segmentation Method and System Based on Frequency Context Feature Hybridization

    CN121095560B