Remote sensing image multi-scale semantic segmentation method based on coding and decoding network
By adopting a codec network-based method in remote sensing image semantic segmentation, combining ResNeXt50_32x4d backbone network, adaptive feature collaboration module (AFCM) and attention-driven advanced semantic integration module (AHSIM), as well as adaptive learning rate scheduling strategy (ALRSS), the problems of difficulty in capturing multi-scale context information in the existing technology are solved, and efficient multi-scale feature fusion and semantic segmentation are achieved, which significantly improves segmentation accuracy and model generalization capabilities.
Patent Information
- Application Number
- CN202510299594.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing remote sensing image semantic segmentation methods are difficult to effectively capture multi-scale context information, resulting in the lack of semantics for large goals and the loss of details for small goals. At the same time, the calculation cost is high, it is difficult to balance accuracy and efficiency. The learning rate scheduling strategy lacks dynamic adaptability, resulting in limited convergence speed and generalization capabilities.
A codec network-based method is adopted to extract multi-scale features through ResNeXt50_32x4d backbone network combined with packet convolution, and an adaptive feature collaboration module (AFCM) and attention-driven advanced semantic integration module (AHSIM) are introduced for feature fusion and semantic integration. At the same time, an adaptive learning rate scheduling strategy (ALRSS) is used for dynamic learning rate adjustment, and morphological post-processing is performed to remove noise and fill small holes.
It significantly improves the segmentation accuracy of complex land objects in remote sensing images, enhances the model's ability to segment large targets and recognizes details of small targets, improves the model's convergence speed and generalization ability, and reduces the dependence on post-processing.
Smart Images

Figure CN120219739A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and remote sensing image processing, and specifically relates to a multi-scale semantic segmentation method for remote sensing images based on an encoder-decoder network. Background Art
[0002] Semantic segmentation of remote sensing images is a key technology in the fields of computer vision and remote sensing image processing, aiming to achieve accurate identification of surface coverage types through pixel-level classification, and is widely used in fields such as urban planning, disaster monitoring, and land use analysis. In recent years, encoder-decoder network architectures based on deep learning (such as U-Net, DeepLab series) have significantly improved the segmentation accuracy by extracting multi-scale features through the encoder and restoring the spatial resolution in combination with the decoder. However, the scenes of remote sensing images are complex, the target scale differences are large, and the edge details are rich. The existing methods still face the following challenges: First, traditional encoder-decoder networks rely on the receptive fields of fixed convolutional kernels and are difficult to effectively capture multi-scale context information, resulting in semantic loss of large targets and detail loss of small targets; Second, the same type of ground objects in remote sensing images often present diverse appearances due to illumination and perspective changes, and the existing attention mechanisms (such as self-attention) have high computational costs and are difficult to balance accuracy and efficiency; In addition, most of the existing learning rate scheduling strategies are based on fixed decay rules and cannot dynamically adapt to the changes in the loss landscape during model training, resulting in limited convergence speed and generalization ability.
[0003] In response to the above problems, existing methods have been improved by introducing pyramid pooling, dilated convolution, and attention modules, but there are still defects such as a single feature fusion method and insufficient dynamic adaptability. For example, although the pyramid pooling module can capture global information, it ignores the fine processing of local details; dilated convolution is prone to produce grid effects while expanding the receptive field, affecting the edge segmentation accuracy; while the method based on Transformer can model long-range dependencies, but has high computational complexity and is difficult to meet the real-time processing requirements of high-resolution scenes of remote sensing images. In addition, the existing learning rate strategies lack dynamic perception of the model state during training, resulting in the model being prone to falling into local optima or overfitting. In response to this, we propose a multi-scale semantic segmentation method for remote sensing images based on an encoder-decoder network. Summary of the Invention
[0004] To solve the above technical problems, a multi-scale semantic segmentation method for remote sensing images based on an encoder-decoder network is provided, and the present technical solution solves the above problems.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A multi-scale semantic segmentation method for remote sensing images based on an encoder-decoder network, comprising the following steps:
[0007] S1. Obtain image data, preprocess the image to normalize it to a specified size, and perform data augmentation operations, where the data augmentation operations include: random rotation, flipping, and scaling;
[0008] S2. Construct a deep learning model. The model adopts an encoder-decoder structure, uses ResNeXt50_32x4d as the backbone network, and extracts multi-scale features through grouped convolution;
[0009] S3. The decoder module gradually restores high-resolution features through bilinear upsampling and makes skip connections with the features of the corresponding layer of the encoder;
[0010] S4. Fusion of local and global context information through the Adaptive Feature Collaboration Module (AFCM). The AFCM includes parallel multi-scale convolutional layers, dilated convolution, and a dense connection feature pyramid. The convolutional kernels of the parallel multi-scale convolutional layers are 1×1, 3×3, and 5×5 respectively, and the dilation rate of the dilated convolution is [3, 6, 12, 18, 30];
[0011] S5. The Attention-driven High-level Semantic Integration Module (AHSIM) adjusts the weights of the features output by the AFCM through a dynamic calibration attention mechanism. The calculation formula is:
[0012] u c =v c *X,F ex =Pooling(F tr (input))
[0013] In the formula, u c represents the c-th channel of the output feature map, v c represents the convolutional kernel for extracting the c-th channel of the input feature X, F ex represents the feature map after pooling, Pooling represents the global average pooling operation, and F tr (input) represents the result of the input feature after passing through the feature transformation operation F tr ;
[0014] S6. Adaptive Learning Rate Scheduling Strategy (ALRSS). Based on the loss change threshold Δloss = 0.001, the learning rate is dynamically adjusted according to the decay factor 0.1;
[0015] S7. Post-process the obtained semantic segmentation result. Use erosion and dilation in morphological operations to remove noise points in the segmentation result, fill small holes, and make the segmentation area smooth and complete.
[0016] Preferably, in step S1, the method of preprocessing the image to normalize it to a specified size and performing data augmentation operations is:
[0017] Among them, the formula for normalizing to a specified size is:
[0018]
[0019] In the formula, α norm represents the result of normalization, α represents the original pixel value, μ represents the pixel mean of the dataset, and σ represents the pixel standard deviation of the dataset;
[0020] Among them, the method for performing data augmentation operations is:
[0021] The angle of the random rotation is [-45°, 45°], and the pixel coordinates are transformed through the rotation matrix R. The expression of the rotation matrix is:
[0022]
[0023] In the formula, θ represents the rotation angle;
[0024] Flipping is divided into horizontal and vertical. When flipping horizontally, the abscissa transformation formula is:
[0025] x new = W - x
[0026] In the formula, x new represents the abscissa after flipping, W represents the image width, and x represents the original abscissa;
[0027] For scaling, bilinear interpolation is used. For the pixel value f(x', y') at the scaled coordinates (x', y'), it is calculated from the four adjacent pixel values f(x0, y0), f(x0, y1), f(x1, y0), f(x1, y1) of the original image according to the following formula:
[0028] f(x′, y′) = (1 - u)(1 - v)f(x0, y0) + u(1 - v)f(x1, y0) + (1 - u)vf(x0, y1) + n vf(x1, y1)
[0029] In the formula, u and v are coefficients used in the bilinear interpolation calculation.
[0030] Preferably, in step S2, the construction of the deep learning model specifically includes:
[0031] The encoder contains multiple convolutional layers and pooling layers for feature extraction. The encoder uses ResNeXt50_32x4d as the backbone network and uses grouped convolution for feature calculation. Its convolution calculation formula is specifically:
[0032] Assume that the input feature map X has a size of H×W and the number of channels is C in , then the input tensor Among them, B is the batch size. Suppose the number of groups G = 32, and the number of channels D = 4 for each group. Then the number of input channels for each group is calculated as follows:
[0033]
[0034] Perform an independent convolution operation on each group g:
[0035]
[0036] In the formula, Y g (i, j) represents the pixel value at the (i, j) position of the convolution output feature map of the g-th group. g = 1, 2,..., G represents the group index, m, n represent the coordinate indices within the convolution kernel, K is the convolution kernel size, and W g (m, n) is the convolution kernel weight of the g-th group;
[0037] The output feature maps Y of each group g Are concatenated through the channel dimension:
[0038] Y = Concat(Y1, Y2,..., Y G )
[0039] Among them, the concatenated output feature map Y has the size B × Cout × H × W, where C out = G × D;
[0040] The decoder restores the resolution through deconvolution and upsampling, and combines the encoder features for semantic segmentation prediction.
[0041] Preferably, in step S2, the skip connection of the encoder-decoder architecture is designed as follows: After upsampling each layer of the decoder, it is concatenated with the features of the corresponding layer of the encoder. The dimension of the concatenated features is:
[0042]
[0043] In the formula, C enc And C dec Are the number of channels of the corresponding layers of the encoder and decoder respectively. By fusing the shallow details and deep semantics through skip connections, the segmentation accuracy is improved.
[0044] Preferably, the multi-scale convolution output of the AFCM is fused through channel concatenation and attention weighting. The specific implementation method is as follows:
[0045] Input feature map Among them, B is the batch size, C is the number of channels, H and W are the height and width. Local to global features are extracted through 1×1, 3×3, and 5×5 convolution kernels respectively:
[0046] X1 = Conv 1×1 (X)
[0047] X3 = Conv 3×3 (X)
[0048] X5 = Conv 5×5 (X)
[0049] Wherein, Conv k×k represents a convolution operation with a kernel size of k×k;
[0050] Channel attention is applied to the 3×3 and 5×5 convolution results:
[0051]
[0052] Wherein, represents the feature after applying channel attention to the 3×3 convolution result, X3 represents the feature after the feature X is processed by the 3×3 convolution kernel, represents per-channel multiplication, σ is the Sigmoid function, which is used to normalize the weights, MLP represents a multi-layer perceptron, which is used to generate channel weights, and GAP represents global average pooling, which is used to compress the spatial dimension, represents the feature after applying channel attention to the 5×5 convolution result, X5 represents the feature after the feature X is processed by the 5×5 convolution kernel;
[0053] The weighted features are concatenated and fused with the feature along the channel dimension, and its expression is:
[0054]
[0055] The final output feature dimension is
[0056] Preferably, the specific implementation of the global average pooling GAP is:
[0057] For the input feature map The GAP output is:
[0058]
[0059] The output dimension is It is used to generate channel attention weights. The attention mechanism is guided by global context information to focus on key channels and suppress redundant information.
[0060] Preferably, in step S4, the dilation rate of the dilated convolution is set to [3, 6, 12, 18, 30], and the corresponding receptive field calculation is:
[0061] R = (d - 1)(K - 1) + K
[0062] Among them, d is the dilation rate, K is the convolution kernel size, and the final maximum receptive field is 133×133. Multi-scale context information is captured through dilated convolutions with multiple dilation rates to enhance the segmentation ability for large objects.
[0063] Preferably, in step S5, the dynamic weight generation process of the attention-driven high-level semantic integration module AHSIM includes:
[0064] Input features Compress the channels to C' through 1×1 convolution;
[0065]
[0066] The high-frequency path enhances the non-linearity through the GELU activation function:
[0067] F high = GELU(F compressed )
[0068] Where GELU(x) = xΦ(x), and Φ(x) is the cumulative distribution function of the standard Gaussian distribution;
[0069] The low-frequency path captures the global context through global average pooling:
[0070] F low = GAP(F compressed )
[0071] The outputs of the two paths are concatenated to generate a weight matrix:
[0072] W att = Softmax(Conv1×1(Concat(Fhigh,F low )))
[0073] Perform channel scaling and spatial correction on the input features:
[0074]
[0075] Calibrate the complementarity of high-frequency and low-frequency features through dynamic weights.
[0076] Preferably, in step S6, the learning rate adjustment logic of the ALRSS is specifically as follows:
[0077] Set the initial learning rate η0 = 0.01 in the initial stage, and the dynamic decay condition is to define the loss change threshold Δ loss = 0.001;
[0078] When the loss decrease amplitude in K consecutive training epochs satisfies the following formula condition, trigger the learning rate decay, where the formula is:
[0079]
[0080] In the formula, represents the average loss of the th training cycle, represents the average loss of the th training cycle;
[0081] The learning rate is updated exponentially with a decay factor γ = 0.1, and the lower limit of the learning rate is 1×10 -6 , where the learning rate update formula is:
[0082] η t+1 = η t ·γ
[0083] This strategy is based on the flatness analysis of the Loss Landscape, and its flatness ρ is defined as:
[0084]
[0085] where r is the perturbation radius, and ALRSS guides the model to converge to the low-ρ region by dynamically adjusting the learning rate.
[0086] Preferably, in step S7, the morphological operations of the post-processing include:
[0087] The erosion operation uses the structuring element S to process the segmentation result:
[0088]
[0089] The dilation operation uses the same structuring element:
[0090]
[0091] Noise points are removed by erosion, and small holes are filled by dilation to make the segmentation region smooth and complete.
[0092] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0093] The multi-scale semantic segmentation method of remote sensing images based on the encoding and decoding network proposed by the present invention effectively extracts multi-scale features by introducing ResNeXt50_32x4d as the backbone network and combining grouped convolution, improving the model's accurate recognition ability of land cover types. Through the fusion strategy of the adaptive feature collaboration module AFCM, it not only fuses local and global context information, but also enhances the model's segmentation ability for large targets through parallel multi-scale convolutional layers and dilated convolution, ensuring the accuracy and integrity of the segmentation results. The adopted adaptive learning rate scheduling strategy ALRSS can dynamically adjust the learning rate according to the loss change during training, avoiding the risk of the model falling into local optimum or overfitting, significantly improving the model's convergence speed and generalization ability. By performing morphological post-processing on the semantic segmentation results, noise points are further removed and small holes are filled, making the segmentation area smoother and more complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 is the flowchart of the method of the present invention;
[0095] Figure 2 is the model architecture diagram of the present invention;
[0096] Figure 3 is the semantic segmentation result of each category on the Potsdam test dataset;
[0097] Figure 4 is the comparison diagram of the segmentation methods on the Vaihingen test dataset;
[0098] Figure 5 is the comparison diagram of the segmentation methods on the DeepGlobe test dataset;
[0099] Figure 6 is the comparison diagram of the segmentation methods before and after integrating AFCM on the ISPRS Potsdam test set. DETAILED DESCRIPTION OF THE INVENTION
[0100] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0101] Referring to Figure 1 shown, a multi-scale semantic segmentation method of remote sensing images based on the encoding and decoding network includes:
[0102] S1: Normalize the image and randomly rotate, flip, and scale to enhance data diversity;
[0103] S2: Construct an encoding and decoding network, and use ResNeXt50_32x4d grouped convolution to extract multi-scale features;
[0104] S3: The decoder bilinearly upsamples to restore the resolution and fuses features with the encoder skip connections;
[0105] S4: The AFCM module fuses context information through multi-scale convolution, dilated convolution, and pyramid;
[0106] S5: The AHSIM dynamic calibration attention mechanism adjusts the feature weights to enhance semantic expression;
[0107] S6: The ALRSS dynamically adjusts the learning rate based on the loss change threshold to optimize the convergence process;
[0108] S7: Post-processing of morphological erosion and dilation to remove noise and smooth the segmentation region.
[0109] In the specific implementation of the present invention, first, high-resolution surface cover image data is obtained through a satellite or aerial remote sensing platform, and the original image is registered in coordinates and radiometrically corrected using professional geographic information system software. In the data preprocessing stage, the input images are uniformly adjusted to a size of 512×512 pixels, and each band is normalized using the Z-Score normalization method to eliminate the influence of illumination differences and sensor noise. To enhance the generalization ability of the model, a data augmentation strategy is implemented in real time during training: the images are randomly rotated within the angle range of [-15°, +15°], and bilinear interpolation is used to maintain geometric accuracy; horizontal or vertical flipping operations are performed with a 50% probability to simulate different shooting perspectives; scale transformation is performed using a random scaling factor of 0.8 - 1.2, and edge padding technology is combined to maintain the integrity of the target. The preprocessed image data is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 to ensure the scientific evaluation of the model.
[0110] Refer to Figure 2 As shown, the network architecture is constructed using an encoder structure improved based on ResNeXt50_32x4d. This backbone network introduces a grouped convolution mechanism with 32 groups of 4 channels on the basis of traditional residual modules. Each group of convolutions independently extracts local features and then concatenates them along the channel dimension, significantly enhancing feature diversity while keeping the number of parameters controllable. The encoder contains five downsampling stages, and the resolution is reduced by convolution with a stride of 2 in each stage. At the same time, max-pooling operations are used to retain important texture information. The decoder part is designed as a symmetric structure, and the spatial resolution is gradually restored by bilinear upsampling and is skip-connected with the features of the corresponding levels of the encoder. After aligning the number of channels using 1×1 convolution, element-wise addition is performed to achieve cross-level fusion of shallow detail features and deep semantic features. After the output of the last layer of the decoder, a 3×3 convolution is connected to refine the features and generate an initial segmentation map with the same resolution as the input.
[0111] In the feature fusion stage, an adaptive feature collaboration module is introduced. This module consists of three parallel multi-scale convolutional branches: the first branch uses a 3×3 standard convolution to capture local details, the second branch uses a 5×5 large kernel convolution to extract regional features, and the third branch configures a dilated convolution sequence with dilation rates of [3, 6, 12] to form a hierarchical receptive field expansion. The output features of each branch are dynamically weighted by a channel attention mechanism. The channel statistics are obtained through global average pooling, and a channel weight vector is generated through two fully connected layers. After normalization using the Sigmoid function, it is multiplied with the original features channel by channel to highlight the contribution of important feature channels. The weighted multi-scale features are fused across scales through a densely connected pyramid structure, and the bottom-layer features are passed to the upper layer for calculation to form a multi-level feature interaction network.
[0112] The semantic integration module uses a dynamic calibration mechanism to optimize the feature representation. The previous-stage features are input into a high-frequency path composed of the GELU activation function and a low-frequency path composed of global average pooling. After the high-frequency path compresses the channel dimension through a 1×1 convolution, the Gaussian error linear unit is used to enhance the non-linear representation ability; the low-frequency path extracts the global context vector and expands it into a spatial weight map through a fully connected layer. The outputs of the two paths are concatenated in the channel dimension and then convolved to generate a spatial-channel joint attention matrix to adaptively calibrate the original features, effectively balancing the contribution ratio of local details and global semantics. This mechanism is particularly suitable for dealing with the apparent differences of the same type of ground objects in remote sensing images caused by illumination and shadows, and improving the adaptability of the model to complex scenes.
[0113] In the training optimization stage, an adaptive learning rate scheduling strategy is implemented. The initial learning rate is set to 0.001, and the cosine annealing algorithm is used for periodic adjustment. Monitor the change in the second derivative of the validation set loss curve. When the loss reduction amplitude is less than the threshold of 0.0005 in three consecutive training cycles, the learning rate decay mechanism is automatically triggered and gradually reduced to the lower limit of 1e-6 by a factor of 0.5. This strategy effectively avoids the training oscillation caused by the traditional stepwise decay, enables the model to converge smoothly to the flat minimum value region, and improves the generalization performance. The loss function uses a joint optimization of the improved Dice loss and cross-entropy, sets a weighted ratio of 0.6:0.4 to balance the class imbalance problem, and introduces an edge-sensitive factor for small target segmentation to strengthen the gradient backpropagation intensity in the contour region.
[0114] In the post-processing link, a multi-scale morphological filtering pipeline is designed. First, the binary segmentation result is eroded twice using a 3×3 circular structuring element to eliminate isolated noise points and small artifacts; then, the same structuring element is used to perform three dilation operations to fill internal holes and smooth the edge serrations. For large continuous regions, an area threshold filter is applied to remove misdetected fragments smaller than 50 pixels. After extracting the boundary line through morphological gradient operation, sub-pixel level optimization is performed, and finally, a vector segmentation result with topological correctness is output.
[0115] As Figures 3 - 5 shown, MSFFENet achieved mIoU of 89.82%, 86.30% and 71.10% on the Potsdam, Vaihingen and DeepGlobe datasets respectively, outperforming baseline models such as DeepLab v3+ and PSPNet. As shown in Table 6, the introduction of AFCM increased the mIoU of the Potsdam dataset by 1.32%, verifying the effectiveness of multi-scale feature fusion.
[0116] In summary, the advantages of the present invention are as follows:
[0117] Through the collaborative design of the Adaptive Feature Collaboration Module (AFCM) and the Attention-driven High-level Semantic Integration Module (AHSIM), the efficient fusion of multi-scale features and the precise complementarity of high-frequency and low-frequency information are realized for the first time, significantly improving the segmentation accuracy of complex ground objects (such as buildings, roads, vehicles) in high-resolution images (the mIoU reaches 89.82%, a 5.35% improvement over the existing optimal method), and effectively solving the problems of missed detection of small targets and blurred edges (the F1-score of vehicles is increased by 9.2%, and the IoU of road edges is increased by 7.8%); combined with the dynamic optimization of the Adaptive Learning Rate Scheduling Strategy (ALRSS), the training efficiency is increased by 15% (the training time is shortened to 22.8 hours), and the model generalization is enhanced (the cross-dataset error is reduced by 18%); in addition, through the grouped convolution lightweight design of ResNeXt50_32x4d (the number of parameters is reduced by 30%) and the bilinear upsampling optimization of the decoder, the real-time inference speed of the model on edge devices is increased by 25%, while reducing the dependence on post-processing operations (the accuracy only drops by 0.8% without post-processing), providing an efficient and reliable technical solution for the real-time interpretation and engineering deployment of high-resolution satellite images.
[0118] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-scale semantic segmentation method for remote sensing images based on a codec network, characterized in that: The following steps are involved: S1, acquiring image data, preprocessing the image to normalize the image to a specified size, and performing data enhancement operations, wherein the data enhancement operations include: random rotation, flipping, and scaling; S2. Build a deep learning model. The model adopts an encoder-decoder structure, uses ResNeXt50_32x4d as the backbone network, and extracts multi-scale features through group convolution; S3, decoder module, gradually restores high-resolution features through bilinear upsampling and jumps with the features of the corresponding layers of the encoder; S4, fusing local and global context information through an adaptive feature collaboration module AFCM, wherein the AFCM includes a parallel multi-scale convolution layer, a dilated convolution, and a densely connected feature pyramid, wherein the convolution kernels of the parallel multi-scale convolution layer are 1×1, 3×3, and 5×5, respectively, and the dilated convolution expansion rate is [3, 6, 12, 18, 30]; S5. The attention-driven high-level semantic integration module AHSIM adjusts the weights of the features output by AFCM through a dynamic calibration attention mechanism. The calculation formula is: u c =v c *X,F ex =Pooling(F tr (input)) In the formula, u c represents the cth channel of the output feature map, vc represents the convolution kernel, which is used to extract the cth channel of the input feature X, and F ex Represents the feature map after pooling, Pooling represents the global average pooling operation, F tr (input) represents the result of the input feature after the feature conversion operation Ftr; S6, adaptive learning rate scheduling strategy ALRSS, based on the loss change threshold Δloss=0.001, dynamically adjusts the learning rate according to the decay factor 0.1; S7. Post-process the semantic segmentation results, use erosion and dilation in morphological operations to remove noise points in the segmentation results, fill small holes, and make the segmented area smooth and complete.
2. According to claim 1, a remote sensing image multi-scale semantic segmentation method based on a codec network is characterized in that: In step S1, the method of preprocessing the image, normalizing the image to a specified size, and performing data enhancement operations is as follows: Among them, the formula for normalization to the specified size is: In the formula, α norm represents the normalized result, α represents the original pixel value, μ represents the mean value of the dataset pixel, and σ represents the standard deviation of the dataset pixel; Among them, the method for performing data enhancement operation is: The angle of the random rotation is [-45°, 45°]. The pixel coordinates are transformed by the rotation matrix R. The expression of the rotation matrix is: Where θ represents the rotation angle; Flipping can be done horizontally or vertically. The horizontal coordinate transformation formula for horizontal flipping is: x new =W-x In the formula, x new represents the horizontal coordinate after flipping, W represents the image width, and x represents the original horizontal coordinate; Scaling uses bilinear interpolation. The pixel value f(x′,y′) at the scaled coordinate (x',y′) is calculated using the following formula through the four adjacent pixel values f(x0,y0), f(x0,y1), f(x1,y0), and f(x1,y1) of the original image: f(x′,y′)=(1-u)(1-v)f(x0,y0)+u(1-v)f(x1,y0)+(1-u)vf(x0,y1)+nvf(x1,y1) In the formula, u and v are the coefficients used in the bilinear interpolation calculation.
3. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: In step S2, the construction of the deep learning model specifically includes: The encoder contains multiple convolutional layers and pooling layers for feature extraction. The encoder uses ResNeXt50_32x4d as the backbone network and uses group convolution for feature calculation. The convolution calculation formula is as follows: Assume that the input feature map X has size H×W and the number of channels is C in , then the input tensor Where B is the batch size, set the number of groups G = 32, the number of channels per group D = 4, then the number of input channels per group is calculated as follows: Perform a separate convolution operation on each group g: Where Y g (i, j) represents the pixel value of the g-th group of convolution output feature maps at position (i, j), g = 1, 2, ..., G represents the group index, m, n represents the coordinate index within the convolution kernel, K is the convolution kernel size, W g (m,n) is the convolution kernel weight of the g-th group; Output feature map Y of each group g Concatenate across the channel dimension: And=Concat(Y1,Y2,…,Y G ) The concatenated output feature map Y has the size B×Cout×H×W, where C out =G×D; The decoder restores the resolution through deconvolution and upsampling, and combines the encoder features to make semantic segmentation predictions.
4. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: In step S2, the jump connection of the encoder-decoder architecture is designed as follows: after upsampling each layer, the decoder concatenates the features of the corresponding layer of the encoder, and the concatenated feature dimensions are: In the formula, C enc and C dec are the number of channels in the corresponding layers of the encoder and decoder respectively. The shallow details and deep semantics are fused through jump connections to improve the segmentation accuracy.
5. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: The multi-scale convolution output of the AFCM is fused by channel splicing and attention weighting, and the specific implementation method is as follows: Input feature map Where B is the batch size, C is the number of channels, H and W are the height and width, and local to global features are extracted through 1×1, 3×3, and 5×5 convolution kernels respectively: X1=Conv 1×1 (X) X3=Conv 3×3 (X) X5=Conv 5×5 (X) In the formula, Conv k×k represents a convolution operation with a kernel size of k×k; Apply channel attention to the 3×3 and 5×5 convolution results: In the formula, It represents the feature after the 3×3 convolution result is applied with channel attention, X3 represents the feature after the feature X is processed by the 3×3 convolution kernel, represents channel-by-channel multiplication, σ is the Sigmoid function, which is used to normalize the weights, MLP represents the multi-layer perceptron, which is used to generate channel weights, GAP represents the global average pooling, which is used to compress the spatial dimension, It represents the features after the 5×5 convolution result is applied with channel attention, and X5 represents the features after the feature X is processed by the 5×5 convolution kernel; The weighted features are concatenated and fused along the channel dimension, and the expression is: X fusion =Concat(X1,X 3Att ,X 5Att ) The final output feature dimension is 6. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: The specific implementation of the global average pooling GAP is: For the input feature map The GAP output is: The output dimension is Used to generate channel attention weights. Use global context information to guide the attention mechanism to focus on key channels and suppress redundant information.
7. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: In step S4, the dilation rate of the dilated convolution is set to [3, 6, 12, 18, 30], and the corresponding receptive field is calculated as: R=(d-1)(K-1)+K Among them, d is the dilation rate, K is the convolution kernel size, and the final maximum receptive field is 133 × 133. Multi-scale context information is captured through multi-dilation rate hole convolution to enhance the segmentation ability of large objects.
8. The multi-scale semantic segmentation method of remote sensing images based on a codec network according to claim 1, characterized in that: In step S5, the dynamic weight generation process of the attention-driven high-level semantic integration module AHSIM includes: Input Features Compress the channel to C′ through 1×1 convolution; The high-frequency path enhances nonlinearity through the GELU activation function: F high =GEL(F compressed ) Where, GELU(x)=xΦ(x), Φ(x) is the standard Gaussian distribution cumulative function; The low-frequency path captures global context via global average pooling: F low =GAP(F compressed ) The dual path outputs are concatenated to generate a weight matrix: W att =Softmax(Conv1×1(Concat(Fhigh,F low ))) Perform channel scaling and spatial correction on input features: The complementarity of high-frequency and low-frequency features is calibrated through dynamic weights.
9. The method for multi-scale semantic segmentation of remote sensing images based on a codec network according to claim 1, characterized in that: In step S6, the learning rate adjustment logic of the ALRSS is as follows: In the initial stage, the initial learning rate η0=0.01 is set, and the dynamic attenuation condition is to define the loss change threshold Δ loss =0.001; When the loss decreases over K consecutive training cycles meets the following formula conditions, the learning rate decay is triggered, where the formula is: In the formula, Indicates The average loss over training cycles, Indicates The average loss over training cycles; The learning rate is updated exponentially with a decay factor of γ = 0.1 and a lower limit of 1×10 -6 , where the learning rate update formula is: or t+1 =the t ·c The strategy is based on the loss landscape flatness analysis, and its flatness ρ is defined as: Where r is the perturbation radius, and ALRSS guides the model to converge to the low ρ region by dynamically adjusting the learning rate.
10. The method for multi-scale semantic segmentation of remote sensing images based on a codec network according to claim 1, characterized in that: In step S7, the post-processing morphological operation includes: The erosion operation uses the structural element S to process the segmentation result: The dilation operation uses the same structuring element: Noise points are removed by erosion, and small holes are filled by expansion to make the segmented area smooth and complete.
Citation Information
Cited By
Deep learning-based weak intercalated layer semantic segmentation method and device, and medium
CN120807923A
Thangka element detection method and system based on improved RT-DETR
CN120852936A
Image processing method, processing device and computer product
CN121353749A
Semantic segmentation method for feature image captured by unmanned aerial vehicle and neural network structure
CN121616833A
A semantic segmentation method for feature images captured by a drone and a neural network structure
CN121616833B