A U-Net Remote Sensing Image Road Segmentation Method Based on Adaptive Dual Filters
By introducing adaptive dual filter technology into U-Net networks, the feature extraction and fusion strategy is optimized, and the problem of insufficient feature fusion in complex scenarios in traditional remote sensing image segmentation methods is solved, achieving higher segmentation accuracy and boundary positioning accuracy.
Patent Information
- Application Number
- CN202411900676.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-23
AI Technical Summary
Traditional remote sensing image segmentation methods have problems such as low spatial resolution and inaccurate boundary positioning due to insufficient feature fusion in complex scenarios.
U-Net network based on adaptive dual filters is adopted, and two fusion layers are set up in the upsampling module of the decoder to generate an adaptive low-pass filter and a high-pass filter, and the extracted feature map is filtered and fused, and the feature extraction and fusion strategy is optimized.
It significantly enhances the ability to identify road boundaries, improves segmentation accuracy and accuracy, reduces the misclassification rate, and enhances the robustness and generalization ability of the model.
Smart Images

Figure CN119850949B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and more specifically, to a U-Net remote sensing image road segmentation method based on an adaptive dual filter. Background Art
[0002] With the rapid development of remote sensing technology, the ability to obtain high-resolution remote sensing images has been significantly enhanced, which provides rich data support for fields such as urban planning, agricultural monitoring, and environmental protection. However, the segmentation task of remote sensing images still faces many challenges, especially in complex scenes where the boundaries of targets such as roads, buildings, and vegetation are blurred. Traditional segmentation methods often rely on pixel-level feature extraction and are easily affected by factors such as illumination changes, perspective changes, and scene complexity, resulting in insufficient segmentation accuracy.
[0003] Existing deep learning models, such as U-Net, although showing excellent performance in image processing in certain specific fields, still need further improvement in the application of remote sensing image processing. The encoder-decoder structure of U-Net can effectively integrate low-level detail information and high-level semantic information, which makes it perform well in certain tasks. However, during the feature fusion process, it often faces interference from within-class inconsistency and between-class similarity. This interference can lead to blurred object boundaries and inaccurate localization in the processing results, especially in remote sensing images with complex backgrounds, where the problem is more prominent. To solve these problems, researchers have begun to focus on the extraction and fusion of frequency-domain features. In the frequency domain, the high-frequency information of an image corresponds to details and edges, while the low-frequency information represents the overall structure of the image. By performing feature processing in the frequency domain, it is possible to more effectively capture and enhance the boundary information of the target, improving the segmentation effect. Based on this, frequency-domain-based feature processing techniques have gradually been introduced into the image segmentation task. However, existing methods often lack adaptive processing of frequency-domain features and are difficult to adjust the feature fusion strategy according to different image contents.
[0004] Therefore, how to adaptively combine spatial features to achieve image segmentation and improve the segmentation accuracy of remote sensing images is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a U-Net remote sensing image road segmentation method based on an adaptive dual filter, aiming to solve the problems of low spatial resolution and inaccurate boundary localization caused by insufficient feature fusion in traditional remote sensing image segmentation. By introducing a dual-filter frequency-domain perception fusion technology, it is possible to effectively enhance feature consistency and sharpen object boundaries, improving the segmentation accuracy, and further supporting applications such as high-precision agricultural monitoring, urban planning, and environmental monitoring.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A U-Net remote sensing image road segmentation method based on an adaptive dual filter, comprising the following steps:
[0008] Step 1: Collect remote sensing images and perform preprocessing to construct a training sample set;
[0009] Step 2: Construct an adaptive dual-filter U-Net network, including an encoder and a decoder. Two fusion layers are set in the upsampling module of the decoder. Each fusion layer generates an adaptive low-pass filter and an adaptive high-pass filter to perform filter fusion on the extracted feature maps;
[0010] Step 3: Use the training sample set to train and optimize the adaptive dual-filter U-Net network to obtain a remote sensing image segmentation model;
[0011] Step 4: Collect the remote sensing image to be segmented, input it into the remote sensing image segmentation model, and obtain the road segmentation result.
[0012] Preferably, the preprocessing includes image annotation, and each pixel in the remote sensing image is labeled with a pixel-level label for its category.
[0013] Preferably, the encoder extracts low-level features in the remote sensing images of the training sample set; the low-level features include edge information, texture information, etc.; the encoder includes 4 downsampling layers. The remote sensing image is used as the input. Each downsampling layer outputs a feature map, and the output of the previous downsampling layer is used as the input of the next downsampling layer to obtain the first input feature map, the second input feature map, the third input feature map, and the fourth input feature map respectively, and the sizes of the 4 feature maps decrease in sequence. The initial image I is processed through a series of convolutional layers. Each convolutional layer applies a set of convolutional kernels (filters) to extract the feature map of the image; a simple CNN structure can be selected to implement the feature map extraction. This structure includes multiple convolutional layers, activation functions (ReLU), and pooling layers (max pooling). After being processed by these convolutional layers, a raw feature map X (the first input feature map) is obtained, and its size is C×H×W, where C is the number of feature channels, and H and W are the height and width of the feature map.
[0014] Preferably, the decoder includes an upsampling module, a skip connection layer, a convolution module, and a segmentation module connected in sequence; the upsampling module includes 4 upsampling layers and two fusion layers. Using the fourth input feature map as the input, each upsampling layer outputs a feature map, and the output of the previous upsampling layer is used as the input of the next upsampling layer to obtain the fourth output feature map, the third output feature map, the second output feature map, and the first output feature map respectively, and the sizes of the 4 feature maps increase sequentially; the first fusion layer is arranged in parallel with the third upsampling layer to fuse the third output feature map and output the first fused feature map; the second fusion layer is arranged in parallel with the fourth upsampling layer to fuse the second output feature map and output the second fused feature; the fourth output feature map, the first fused feature map, the third fused feature map, and the first output feature map are merged with the corresponding input feature maps in the skip connection layer and then input to the convolution module for fusion to obtain the third fused feature map, and the third fused feature map is input to the segmentation module for road classification to obtain the road segmentation result.
[0015] Preferably, each fusion layer includes an expansion module, 2 convolutional layers, a filter generation module, a low-pass filtering module, a high-pass filtering module, and a feature fusion module;
[0016] The expansion module expands the size of the input output feature map by 2 times to obtain an expanded output feature map;
[0017] The first convolutional layer performs a convolutional operation on the input output feature map to obtain a first convolutional feature map, and the second convolutional layer performs a convolutional operation on the expanded output feature map to obtain a second convolutional feature map;
[0018] The filter generation module generates an adaptive low-pass filter and an adaptive high-pass filter respectively according to the second convolutional feature map;
[0019] The low-pass filtering module uses the adaptive low-pass filter to filter the first convolutional feature map to obtain a low-pass filtered feature map, and performs pixel scrambling fusion with the first convolutional feature map to obtain a low-frequency feature map;
[0020] The high-pass filtering module uses the adaptive high-pass filter to filter the second convolutional feature map to obtain a high-pass filtered feature map, and performs pixel-by-pixel addition fusion with the second convolutional feature map to obtain a high-frequency feature map;
[0021] The feature fusion module fuses the low-frequency feature map and the high-frequency feature map to obtain a fused feature map, including the first fused feature map or the second fused feature map.
[0022] Preferably, the filter generation module includes a convolutional layer, a pixel rearrangement layer, a softmax layer, and a filter inversion layer connected in sequence; the second convolutional feature map is convolved by the convolutional layer to obtain a third convolutional feature map; the pixel rearrangement layer rearranges the pixels of the third convolutional feature map to obtain a rearranged feature map; the rearranged feature map is activated by the softmax layer to obtain a low-pass filter; the filter inversion layer inverses the low-pass filter according to the input identity kernel, and the identity kernel is a 3×3 convolutional kernel, which actually convolves the low-pass filter with the 3×3 convolutional kernel to achieve inversion and obtain a high-pass filter. The convolutional layer uses a 3×3 convolutional layer, and the softmax layer uses the Softmax function.
[0023] Preferably, the specific process of generating an adaptive low-pass filter and performing low-pass filtering through the second convolutional layer, the filter generation module, and the low-pass filtering module is as follows:
[0024] Step 311: The second convolutional layer uses a 1×1 convolutional layer to perform a convolutional operation on the extended output feature map to obtain a second convolutional feature map X;
[0025] Step 312: The filter generation module sequentially performs convolution, pixel rearrangement, and activation operations on the second convolutional feature map to generate an adaptive low-pass filter Expressed as:
[0026]
[0027] Among them, i and j respectively represent the row and column coordinates in the feature map; p and q respectively represent the relative positions in the low-pass filter kernel; Ω represents a 3×3 filter kernel; PixelShuffle represents pixel disordering to achieve pixel rearrangement; Softmax represents the Softmax function to achieve activation operation; the specific process is as follows:
[0028] Step 3121: Perform a convolutional operation on the second convolutional feature map to obtain a third convolutional feature map
[0029] Step 3122: Use pixel disordering operation to reshape the third convolutional feature map in the way of pixel de-mixing, reduce the height and width by half respectively, and expand the channels by 4 times at the same time;
[0030] Step 3123: Apply the Softmax function to each filter kernel separately to ensure that the weights of each filter kernel still satisfy the properties of the probability distribution after applying the Softmax function, divide the channels into 4 groups, and each group has an adaptive low-pass filter with spatial transformation, expressed as Among them, g ∈ 1, 2, 3, 4 represents the group;
[0031] Step 313: The low-pass filtering module smooths and upsamples the first convolutional feature map using an adaptive low-pass filter; the filtered low-pass filtering feature ψ is rearranged using a sub-pixel upsampling technique with a low-pass filter to achieve upsampling and obtain a 2x upsampled feature value. l+1 Specifically: Specifically:
[0032] Step 3131: The first convolutional feature map is filtered using an adaptive low-pass filter to obtain 4 sets of low-pass filtering features. g = 1, 2, 3, 4, which is expressed as:
[0033]
[0034] Step 3132: The 4 sets of low-pass filtering features are pixel-shuffled and rearranged into a 2x upsampled feature value. It is expressed as:
[0035]
[0036] Among them, PixelShuffle() represents pixel shuffling, which is an operation that converts channel information into spatial information.
[0037] Preferably, the process of the filtering inversion layer inverting the low-pass filter to generate an adaptive high-pass filter, and the high-pass filtering module filtering using the adaptive high-pass filter to enhance the detailed boundary information lost during the downsampling process is as follows:
[0038] Step 321: Invert the low-pass filter to generate an adaptive high-pass filter. It is expressed as:
[0039]
[0040] Among them, E represents the unit kernel; to ensure that the finally generated kernel is high-pass, first obtain the low-pass kernel using per-kernel softmax, and then subtract these kernels from the unit kernel E, and the weight of the unit kernel E is [[0, 0, 0], [0, 1, 0], [0, 0, 0]];
[0041] Step 322: After filtering the second convolutional feature map using the adaptive high-pass filter, perform residual addition with the second convolutional feature map to obtain an enhanced result and generate a high-frequency feature map. It is expressed as:
[0042]
[0043] represents the feature value at the (i, j) position in the l-th layer of high-frequency feature map.
[0044] Preferably, the feature fusion module doubles the upsampled feature values and the high-frequency feature map are added pixel by pixel to obtain a fused feature map; specifically:
[0045] For the doubled upsampled feature values and the high-frequency feature map channel compression is respectively performed using a 1×1 convolutional layer, the compressed features are added pixel by pixel, and then the fused feature map is obtained through upsampling, expressed as:
[0046]
[0047] where, represents the compressed feature after fusion, where r is the channel reduction rate, used to reduce the subsequent calculation cost of the three; represents the upsampling operation, applying nearest neighbor interpolation to upsample the low-resolution feature map to a high resolution; Conv 1×1 represents the 1×1 convolutional operation.
[0048] Preferably, the segmentation module uses the activation function Softmax to perform road target segmentation.
[0049] The technical effects of the above technical solution are as follows: An adaptive low-pass filter generator is applied to predict a spatially varying low-pass filter, adaptively use the spatially transformed low-pass filter to smooth the high-level features, resample the nearby low-frequency features with consistent categories to replace the inconsistent features in the high-level features, and enhance the high-frequency boundary details of the low-level features, thereby solving the problems of category inconsistency and boundary displacement, reducing the high-frequency components inside the objects in the remote sensing image, and reducing the intra-class inconsistency during the upsampling process; use an adaptive high-pass filter generator to predict a high-pass filter, extract high-frequency edge and detail features from the low-level features, and enhance the high-frequency information lost during the downsampling process to achieve more accurate boundary delineation; further optimize the feature fusion strategy through a dual filter. The adaptive low-pass filter generator and the self-use high-pass filter generator separate the information into low-frequency and high-frequency parts in the subsequent feature processing through parameter learning or fixed filter design. During the training process of the U-Net network, different types of features are filtered by the two generators, and the generated low-pass and high-pass filters act together on feature extraction and separation, increasing the intra-class feature similarity and improving the inter-class feature discrimination, enabling the network to more specifically learn different category information during the training process, improving the discriminant ability and separation ability of the model, thereby reducing the misclassification rate and improving the classification effect.
[0050] Preferably, the process of constructing an optimization target to achieve model training is as follows:
[0051] S331: Fused feature map Y l After passing through the skip connection layer, the convolutional module, and the segmentation head of the segmentation module in sequence, the predicted segmentation result is obtained The training sample set includes the manually annotated true segmentation result y i ;
[0052] S332: According to the predicted segmentation result and the true segmentation result y i Calculate the loss value, optimize the network, with the goal of making the predicted segmentation result consistent with the true segmentation result. Determine the termination of model training according to the optimization goal to obtain the remote sensing image segmentation model
[0053] Preferably, in step 3, a multi-loss function is used to optimize the network. The multi-loss function L is expressed as:
[0054] L = L ce + λL b
[0055]
[0056]
[0057] Among them, L ce represents the cross-entropy loss; L b represents the boundary loss; λ represents the model parameter
[0058] Preferably, in step 3, the Adam optimizer is used to optimize the network. The parameter update rule of the Adam optimizer is:
[0059]
[0060] Among them, θ t is the model parameter at time step t; α is the learning rate; β is the exponential decay rate for calculating the gradient, and β is the exponential decay rate for calculating the square of the gradient; g t is the gradient at time step t; m t is the cumulative average of the square of the gradient at time step t
[0061] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for segmenting roads in remote sensing images based on an adaptive dual filter. By combining an adaptive low-pass filter and a high-pass filter, and integrating a feature fusion layer in the upsampling path, the feature extraction and fusion strategy is optimized, enhancing the model's ability to capture details and improving the accuracy of target boundary localization. This method not only improves the segmentation accuracy but also effectively addresses the deficiencies of traditional methods in complex scenarios, providing a more accurate solution for the application of remote sensing images, with higher accuracy and better boundary alignment, meeting the application requirements of remote sensing image segmentation in multiple fields. Specifically:
[0062] 1) By introducing an adaptive dual filter, the present invention optimizes the feature extraction and fusion strategy, significantly enhancing the ability to identify road boundaries. Especially in complex backgrounds and variable lighting conditions, it can achieve more refined and accurate road segmentation, effectively improving the accuracy and precision of road segmentation in remote sensing images;
[0063] 2) Through the adaptive processing of frequency domain features, the present invention can maintain stable segmentation performance under different conditions, effectively reducing misclassification caused by within-class inconsistency and boundary displacement, enhancing the robustness and generalization ability of the model, and making the model have better adaptability and reliability in a variable environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0065] Figure 1 It is a schematic diagram of the adaptive dual-filter U-Net network structure provided by the present invention;
[0066] Figure 2 It is a schematic diagram of the filter generation module structure provided by the present invention;
[0067] Figure 3 It is a schematic diagram of the fusion layer structure provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0069] An embodiment of the present invention discloses a U-Net remote sensing image road segmentation method based on an adaptive dual filter, including the following steps:
[0070] S1: Collect remote sensing images and perform preprocessing to construct a training sample set;
[0071] S2: Construct an adaptive dual-filter U-Net network, including an encoder and a decoder. Two fusion layers are set in the upsampling module of the decoder. Each fusion layer generates an adaptive low-pass filter and an adaptive high-pass filter to filter and fuse the extracted feature maps;
[0072] S3: Use the training sample set to train and optimize the adaptive dual-filter U-Net network to obtain a remote sensing image segmentation model;
[0073] S4: Collect the remote sensing image to be segmented, input it into the remote sensing image segmentation model, and obtain the road segmentation result.
[0074] Further, the preprocessing includes image annotation, which labels the category of each pixel in the remote sensing image with a pixel-level label.
[0075] Further, the encoder extracts low-level features in the remote sensing images of the training sample set; the low-level features include edge information, texture information, etc.; the encoder includes 4 downsampling layers. With the remote sensing image as the input, each downsampling layer outputs a feature map, and the output of the previous downsampling layer is used as the input of the next downsampling layer to obtain the first input feature map, the second input feature map, the third input feature map, and the fourth input feature map respectively, and the sizes of the 4 feature maps decrease in turn.
[0076] Further, the decoder includes an upsampling module, a skip connection layer, a convolution module, and a segmentation module connected in sequence; the upsampling module includes 4 upsampling layers and two fusion layers. With the fourth input feature map as the input, each upsampling layer outputs a feature map, and the output of the previous upsampling layer is used as the input of the next upsampling layer to obtain the fourth output feature map, the third output feature map, the second output feature map, and the first output feature map respectively, and the sizes of the 4 feature maps increase in turn; the first fusion layer is arranged in parallel with the third upsampling layer to fuse the third output feature map and output the first fusion feature map; the second fusion layer is arranged in parallel with the fourth upsampling layer to fuse the second output feature map and output the second fusion feature; the fourth output feature map, the first fusion feature map, the third fusion feature map, and the first output feature map are merged with the corresponding input feature maps in the skip connection layer and then input into the convolution module for fusion to obtain the third fusion feature map, and the third fusion feature map is input into the segmentation module for road classification to obtain the road segmentation result. The processing flow of the upsampling layer is as Figure 1 shown.
[0077] Furthermore, each fusion layer includes an expansion module, two convolutional layers, a filter generation module, a low-pass filtering module, a high-pass filtering module, and a feature fusion module; as Figure 3 shown;
[0078] The expansion module expands the size of the input output feature map by 2 times to obtain an expanded output feature map;
[0079] The first convolutional layer performs a convolutional operation on the input output feature map to obtain a first convolutional feature map, and the second convolutional layer performs a convolutional operation on the expanded output feature map to obtain a second convolutional feature map;
[0080] The filter generation module generates an adaptive low-pass filter and an adaptive high-pass filter respectively according to the second convolutional feature map;
[0081] The low-pass filtering module filters the first convolutional feature map with the adaptive low-pass filter to obtain a low-pass filtered feature map, and performs pixel scrambling fusion with the first convolutional feature map to obtain a low-frequency feature map;
[0082] The high-pass filtering module filters the second convolutional feature map with the adaptive high-pass filter to obtain a high-pass filtered feature map, and performs pixel-by-pixel addition fusion with the second convolutional feature map to obtain a high-frequency feature map;
[0083] The feature fusion module fuses the low-frequency feature map and the high-frequency feature map to obtain a fused feature map, including a first fused feature map or a second fused feature map.
[0084] Furthermore, the filter generation module includes a convolutional layer, a pixel reordering layer, a softmax layer, and a filter inversion layer connected in sequence, as Figure 2 shown; The second convolutional feature map obtains a third convolutional feature map through the convolutional operation of the convolutional layer; the pixel reordering layer performs pixel rearrangement on the third convolutional feature map to obtain a rearranged feature map; the rearranged feature map is reshaped or activated by the softmax layer to obtain a low-pass filter; the filter inversion layer inverses the low-pass filter according to the input identity kernel, and the identity kernel is a 3×3 convolutional kernel, and its essence is to perform convolution on the low-pass filter with the 3×3 convolutional kernel to achieve inversion and obtain a high-pass filter. The convolutional layer uses a 3×3 convolutional layer, and the softmax layer uses the Softmax function.
[0085] Furthermore, the filter generation module predicts a spatially varying low-pass filter to reduce intra-class inconsistency during the upsampling process; The specific process of generating and performing low-pass filtering with the adaptive low-pass filter through the second convolutional layer, the filter generation module, and the low-pass filtering module is as follows:
[0086] S311: The second convolutional layer uses a 1×1 convolutional layer to perform a convolution operation on the extended output feature map to obtain a second convolutional feature map X;
[0087] S312: The filter generation module sequentially performs convolution, pixel rearrangement, reshaping, or activation operations on the second convolutional feature map to generate an adaptive low-pass filter Expressed as:
[0088]
[0089] Among them, represents the kernel size of the low-pass filter. After passing through the softmax layer, a smoothed low-pass filter is obtained i and j respectively represent the row and column coordinates in the feature map; p and q respectively represent the relative positions in the low-pass filter kernel; Ω represents a 3×3 filter kernel; PixelShuffle represents pixel shuffling to achieve pixel rearrangement; Softmax represents the Softmax function to achieve the activation operation; The specific process is as follows:
[0090] Step 3121: Perform a convolution operation on the second convolutional feature map to obtain a third convolutional feature map
[0091] Step 3122: Use the pixel shuffling operation to reshape the third convolutional feature map W l in the way of pixel de-mixing, reducing the height and width by half respectively, and at the same time expanding the number of channels by 4 times;
[0092] Step 3123: Apply the Softmax function to each filter kernel separately to ensure that the weights of each filter kernel still satisfy the properties of the probability distribution after applying the Softmax function. Divide the channels into 4 groups, and each group has an adaptive low-pass filter with spatial transformation, expressed as where g ∈ {1, 2, 3, 4} represents the group;
[0093] S313: The low-pass filtering module uses the adaptive low-pass filter to smooth and upsample the first convolutional feature map; Use the sub-pixel upsampling technique with the low-pass filter to rearrange the filtered low-pass filtering feature ψ l+1 to achieve upsampling and obtain a 2-fold upsampled feature value Specifically:
[0094] Step 3131: Filter the first convolutional feature map with the adaptive low-pass filter to obtain 4 groups of low-pass filtering features g = 1, 2, 3, 4, expressed as:
[0095]
[0096] S3132: Rearrange the 4 groups of low-pass filter features into 2x upsampled feature values to obtain a low-frequency feature map Denoted as:
[0097]
[0098] where PixelShuffle() represents pixel shuffling, which is an operation that converts channel information into spatial information.
[0099] Furthermore, the filter inversion layer inverts the low-pass filter to generate an adaptive high-pass filter, and the high-pass filtering module uses the adaptive high-pass filter for filtering to extract high-frequency detail information to enhance the detailed boundary information lost during the downsampling process, achieving more accurate boundary delineation; the specific process is as follows:
[0100] S321: Invert the low-pass filter to generate an adaptive high-pass filter Denoted as:
[0101]
[0102] where E represents the identity kernel; to ensure that the finally generated kernel is high-pass, first obtain the low-pass kernel using per-kernel softmax, and then subtract these kernels from the identity kernel E, where the weights of the identity kernel E are [[0,0,0],[0,1,0],[0,0,0]];
[0103] S322: After filtering the second convolutional feature map using the adaptive high-pass filter, add it to the second convolutional feature map residually to obtain a high-frequency feature map Denoted as:
[0104]
[0105] represents the feature value at the (i,j) position in the high-frequency feature map of the l-th layer.
[0106] Furthermore, the feature fusion module adds the 2x upsampled feature values (low-frequency feature map) and the high-frequency feature map pixel by pixel to obtain a fused feature map; specifically:
[0107] For the 2x upsampled feature values and the high-frequency feature map perform channel compression using a 1x1 convolutional layer respectively, add the compressed features pixel by pixel, and then obtain the fused feature map through upsampling, denoted as:
[0108]
[0109] Among them, represents the fused compressed feature, where r is the channel reduction rate, which is used to reduce the subsequent calculation cost of the three; represents the upsampling operation. Applying the nearest neighbor interpolation, the low-resolution feature map is upsampled to a high resolution; Conv 1×1 represents the 1×1 convolution operation.
[0110] Furthermore, the segmentation module uses the activation function Softmax to perform road target segmentation.
[0111] Furthermore, the process of constructing the optimization objective to achieve model training is as follows:
[0112] S331: The fused feature map Y l passes through the skip connection layer, the convolution module, and the segmentation head of the segmentation module in sequence to obtain the predicted segmentation result The training sample set includes the manually annotated true segmentation result y i ;
[0113] S332: According to the predicted segmentation result and the true segmentation result y i calculate the loss value, optimize the network, take the consistency between the predicted segmentation result and the true segmentation result as the optimization objective, and judge to terminate the model training according to the optimization objective to obtain the remote sensing image segmentation model.
[0114] Furthermore, in S3, a multi-loss function is used to optimize the network to calculate the loss value. The loss value L is expressed as:
[0115] L = L ce + λL b
[0116]
[0117]
[0118] Among them, L ce represents the cross-entropy loss; L b represents the boundary loss; λ represents the model parameter; N represents the number of samples.
[0119] Furthermore, in S3, the Adam optimizer is used to optimize the network. The parameter update rule of the Adam optimizer is:
[0120]
[0121] Among them, θ t is the model parameter at time step t; α is the learning rate; β is the exponential decay rate for calculating the gradient, and β is the exponential decay rate for calculating the square of the gradient; g tis the gradient at time step t; m t is the cumulative average of the squared gradients at time step t.
[0122] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method part.
[0123] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A U-Net remote sensing image road segmentation method based on adaptive dual filters, characterized in that: The following steps are involved: Step 1: Collect remote sensing images and preprocess them to construct a training sample set; Step 2: Construct an adaptive dual-filter U-Net network, including an encoder and a decoder. Two fusion layers are set in the upsampling module of the decoder. Each fusion layer generates an adaptive low-pass filter and an adaptive high-pass filter to perform filtering fusion on the extracted feature maps. Step 3: Use the training sample set to train and optimize the adaptive dual-filter U-Net network to obtain a remote sensing image segmentation model; Step 4: Collect the remote sensing image to be segmented, input it into the remote sensing image segmentation model, and obtain the road segmentation result; The encoder extracts low-level feature maps from the remote sensing images of the training sample set; the low-level feature maps include edge information and texture information; the encoder includes four downsampling layers, the remote sensing image is used as input, each downsampling layer outputs a feature map, the output of the previous downsampling layer is used as the input of the next downsampling layer, and the first input feature map, the second input feature map, the third input feature map and the fourth input feature map are obtained respectively, and the sizes of the four feature maps decrease in sequence; The decoder includes an upsampling module, a skip connection layer, a convolution module and a segmentation module connected in sequence; the upsampling module includes four upsampling layers and two fusion layers, wherein the fourth input feature map is used as input, each upsampling layer outputs a feature map, and the output of the previous upsampling layer is used as the input of the next upsampling layer, and the fourth output feature map, the third output feature map, the second output feature map and the first output feature map are obtained respectively, and the sizes of the four feature maps are increased in sequence; the first fusion layer is arranged in parallel with the third upsampling layer, and is used to fuse the third output feature map and output the first fused feature map; the second fusion layer is arranged in parallel with the fourth upsampling layer, and is used to fuse the second output feature map and output the second fused feature map; The fourth output feature map, the first fused feature map, the third fused feature map and the first output feature map are merged with the corresponding input feature map in the skip connection layer and then input into the convolution module for fusion to obtain a third fused feature map. The third fused feature map is input into the segmentation module for road classification to obtain a road segmentation result. Each fusion layer includes an expansion module, two convolutional layers, a filter generation module, a low-pass filter module, a high-pass filter module, and a feature fusion module; The expansion module expands the size of the input output feature map by 2 times to obtain an expanded output feature map; The first convolution layer performs a convolution operation on the input output feature map to obtain the first convolution feature map, and the second convolution layer performs a convolution operation on the expanded output feature map to obtain the second convolution feature map; The filter generation module generates an adaptive low-pass filter and an adaptive high-pass filter according to the second convolution feature map; The low-pass filtering module uses an adaptive low-pass filter to filter the first convolution feature map to obtain a low-pass filtering feature map, and performs pixel random fusion with the first convolution feature map to obtain a low-frequency feature map; The high-pass filtering module uses an adaptive high-pass filter to filter the second convolution feature map to obtain a high-pass filtering feature map, and adds and fuses the high-frequency feature map with the second convolution feature map pixel by pixel to obtain a high-frequency feature map; The feature fusion module fuses the low-frequency feature map and the high-frequency feature map to obtain a fused feature map.
2. According to claim 1, a U-Net remote sensing image road segmentation method based on adaptive dual filters is characterized in that: The filter generation module includes a convolution layer, a pixel reordering layer, a softmax layer and a filter inversion layer connected in sequence; the second convolution feature map is obtained by the convolution operation of the convolution layer to obtain a third convolution feature map; the pixel reordering layer reorders the pixels of the third convolution feature map to obtain a reordered feature map; the reordered feature map is activated by the softmax layer to obtain a low-pass filter; The filter inversion layer convolves the low-pass filter with a 3×3 convolution kernel to obtain a high-pass filter.
3. The method for road segmentation of remote sensing images based on U-Net with adaptive dual filters according to claim 1, characterized in that: The specific process of realizing adaptive low-pass filter generation and low-pass filtering through the second convolution layer, filter generation module and low-pass filtering module is as follows: Step 311: the second convolutional layer uses a 1×1 convolutional layer to perform a convolution operation on the extended output feature map to obtain a second convolutional feature map; Step 312: The filter generation module performs convolution, pixel rearrangement, and activation operations on the second convolution feature map in sequence to generate an adaptive low-pass filter. It is expressed as: Among them, i and j represent the row and column coordinates in the feature map respectively; p and q represent the relative positions in the low-pass filter kernel respectively; PixelShuffle represents pixel shuffling to achieve pixel rearrangement; Softmax represents the Softmax function; the specific process is: Step 3121: Perform a convolution operation on the second convolution feature map to obtain a third convolution feature map Step 3122: Use pixel shuffle operation to perform convolution on the third convolution feature map Reshape it, reduce the height and width by half, and expand the channel by 4 times; Step 3123: Apply the Softmax function to each filter kernel separately, ensure that the weight of each filter kernel still satisfies the property of probability distribution after applying the Softmax function, and divide the channels into 4 groups, each group having a spatially transformed adaptive low-pass filter; Step 313: The low-pass filtering module filters the first convolution feature map using an adaptive low-pass filter, and performs an adaptive low-pass filtering on the filtered low-pass filter feature ψ l+1 Rearrange and upsample to obtain 2 times upsampled feature values Specifically: Step 3131: The first convolution feature map is filtered using an adaptive low-pass filter to obtain four sets of low-pass filter features. g=1,2,3,4, expressed as: Step 3132: Rearrange the 4 sets of low-pass filter features into 2 times upsampled feature values It is expressed as: Among them, PixelShuffle() represents pixel shuffling, which is an operation of converting channel information into spatial information.
4. The method for road segmentation of remote sensing images based on U-Net with adaptive dual filters according to claim 3 is characterized in that: The filtering inversion layer inverts the low-pass filter to generate an adaptive high-pass filter, and the high-pass filtering module uses the adaptive high-pass filter to filter as follows: Step 321: Invert the low-pass filter to generate an adaptive high-pass filter It is expressed as: Among them, E represents the unit nucleus; Step 322: After filtering the second convolution feature map using an adaptive high-pass filter, the residual is added to the second convolution feature map to obtain a high-frequency feature map. It is expressed as: in, Represents the feature value of the (i, j) position in the high-frequency feature map of the lth layer.
5. The method for road segmentation of remote sensing images based on U-Net with adaptive dual filters according to claim 4, characterized in that: The feature fusion module upsamples the feature values by a factor of 2 And high frequency feature map Add pixel by pixel to obtain the fused feature map; specifically: Upsample the feature values by a factor of 2 And high frequency feature map A 1×1 convolutional layer is used to compress the channels, and the compressed features are added pixel by pixel. Then, the fused feature map is obtained by upsampling, which is expressed as: in, represents the fused feature map, where r is the channel reduction rate; Represents upsampling operation; Conv 1×1 Represents a 1×1 convolution operation.
6. The method for road segmentation of remote sensing images based on U-Net with adaptive dual filters according to claim 1, characterized in that: The process of constructing the optimization target to implement model training is: Step 331: The segmentation module outputs the predicted segmentation result, and the training sample set is annotated with the real segmentation result; Step 332: Calculate the loss value based on the predicted segmentation result and the true segmentation result, take the consistency between the predicted segmentation result and the true segmentation result as the optimization goal, terminate the model training based on the optimization goal, and obtain the remote sensing image segmentation model.
7. The method for road segmentation of remote sensing images based on U-Net with adaptive dual filters according to claim 1, characterized in that: In step 3, a multi-loss function is used to optimize the network. The multi-loss function L is expressed as: L=L ce +λL b Among them, L ce represents the cross entropy loss; L b represents the boundary loss; λ represents the model parameter; Represents the predicted segmentation result; y i represents the actual segmentation result; N represents the number of samples.
Citation Information
Patent Citations
road extraction method based on remote sensing images and deep learning
CN109800736A
Road extraction method and system based on dynamic routing neural network
CN116665175A