A small sample remote sensing image semantic segmentation method based on low-rank feature decoupling

CN118397270BActive Publication Date: 2026-09-25BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410475265.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2026-09-25
Estimated Expiration
2044-04-19

AI Technical Summary

Technical Problem

[0004]为解决上述技术问题,本发明提供了一种基于低秩特征解耦的小样本遥感图像语义分割方法,通过将基类和新类学习的特征进行解耦,能够解决模型在微调过程中对于基类目标分割效果下降的问题;提出低秩微调卷积层,将可训练的卷积核参数低秩分解,能够有效降低模型微调的参数量,进一步有效解决模型对于新类的过拟合问题,该方法对于提升小样本遥感图像分割的准确率有重要的研究意义

Benefits of technology

[0006]本发明的有益效果在于:通过采用低秩微调卷积层,提升优化效率并降低内存需求,减少微调时对于小样本新类别的过拟合风险;通过采用基类和新类两阶段解耦学习的优化方法,在保持遥感图像原有地物知识的同时,提高了模型对于新类类别的学习效果;进一步通过融合基类和新类地物分割预测结果图,实现了精确的多类别遥感图像地物分割。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397270B_ABST
    Figure CN118397270B_ABST
Patent Text Reader

Abstract

The application provides a small sample remote sensing image semantic segmentation method based on low-rank feature decoupling, comprising the following steps: training on a large-scale data set containing base-class ground object target labeling to obtain a base-class ground object segmentation model; retaining the encoder and base-class segmentation decoder structure in the base-class ground object segmentation model, introducing a new class segmentation decoder for new class ground object segmentation, and adopting a low-rank fine-tuning convolution layer for the convolution layer of the new class segmentation decoder; training the new class segmentation decoder, freezing the encoder and base-class segmentation decoder in the base-class ground object segmentation model, and only updating the low-rank fine-tuning convolution layer parameters of the new class segmentation decoder; and fusing the output base-class ground object segmentation prediction result and the new class ground object segmentation prediction result to generate a final segmentation result containing all classes. The method significantly improves the accuracy of new class ground object segmentation under small sample conditions without affecting the base-class ground object segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image interpretation and machine learning, specifically involving a few-sample remote sensing image semantic segmentation method based on low-rank feature decoupling. Background Technology

[0002] Semantic segmentation of remote sensing images is a crucial step in intelligent interpretation. It classifies remote sensing images pixel-by-pixel based on the consistency or similarity features within the objects being segmented. It has significant applications in fields such as land resource statistics and resource exploration. Current mainstream supervised deep learning methods rely on large amounts of labeled data and have achieved good performance in data-driven tasks. However, constructing high-resolution remote sensing target segmentation datasets requires precise expert judgment and consumes substantial human and material resources. Therefore, semantic segmentation of small-sample remote sensing images is of significant research value. The challenge lies in extracting segmentation cues from limited supporting images, overcoming intra-class differences and inter-class similarities of target objects, and achieving rapid generalization of new categories under small-sample conditions.

[0003] Existing few-sample image semantic segmentation algorithms are divided into meta-learning-based methods and model fine-tuning-based methods. Meta-learning-based methods rely on the support set as input during testing, which has two problems: (1) This paradigm requires the construction of a corresponding query set for target category matching and identification of sample categories in the support set for each model test; (2) This paradigm cannot perform multi-class segmentation, it can only segment a single new category target, and it cannot segment the base class targets that the model has learned. The above problems make it impossible for the meta-learning-based few-sample segmentation paradigm to be used on a large scale in real-world scenarios. Although the model fine-tuning-based method can achieve the segmentation of base class and new class targets at the same time, due to the complexity of remote sensing images, directly fine-tuning the model has the following two problems: (1) During the process of learning new categories, the accuracy of the model in recognizing base class ground objects drops significantly; (2) In the later stages of learning new categories of ground objects, the model will suffer from severe overfitting to the small-sample training set. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a few-sample remote sensing image semantic segmentation method based on low-rank feature decoupling. By decoupling the features learned from the base class and the new class, it can solve the problem of decreased segmentation performance for base class targets during model fine-tuning. Furthermore, it proposes a low-rank fine-tuning convolutional layer, which decomposes the trainable convolutional kernel parameters into low-rank components, effectively reducing the number of parameters required for model fine-tuning and further effectively solving the overfitting problem for new classes. This method has significant research value for improving the accuracy of few-sample remote sensing image segmentation.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A few-sample remote sensing image semantic segmentation method based on low-rank feature decoupling includes the following steps: Step 1: Train the base class land cover segmentation model on a large-scale dataset containing base class land cover object annotations; Step 2: Retain the encoder and base class segmentation decoder structure in the base class land cover segmentation model, and introduce a new class segmentation decoder for new class land cover segmentation. The convolutional layer of the new class segmentation decoder adopts a low-rank fine-tuned convolutional layer. Step 3: Train the new class segmentation decoder. During the training process, freeze the encoder and base class segmentation decoder in the base class land cover segmentation model, and only update the low-rank fine-tuning convolutional layer parameters of the new class segmentation decoder to decouple the base class land cover segmentation prediction results generated by the base class segmentation decoder from the new class land cover segmentation prediction results generated by the new class segmentation decoder. Step 4: Combine the segmentation prediction results of the base class land cover and the segmentation prediction results of the new class land cover output in Step 3 to generate the final segmentation result that includes all classes.

[0006] The beneficial effects of this invention are as follows: by employing low-rank fine-tuning convolutional layers, optimization efficiency is improved and memory requirements are reduced, thus reducing the risk of overfitting to new categories with small samples during fine-tuning; by adopting an optimization method of decoupled learning in two stages for base class and new class, the learning effect of the model for new categories is improved while maintaining the original ground cover knowledge of remote sensing images; furthermore, by fusing the segmentation prediction results of base class and new class ground cover, accurate multi-class remote sensing image ground cover segmentation is achieved. Attached Figure Description

[0007] Figure 1 This is a flowchart of a small-sample remote sensing image semantic segmentation method based on low-rank feature decoupling according to the present invention. Figure 2 This is a schematic diagram of the low-rank fine-tuning convolution structure of the present invention. Detailed Implementation

[0008] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0009] like Figure 1The diagram shows a flowchart of a few-sample remote sensing image semantic segmentation method based on low-rank feature decoupling according to the present invention. Step 1: Train the model on a large-scale dataset containing base class land cover annotations to obtain a pre-trained remote sensing image segmentation model, aiming to achieve accurate segmentation of base class land cover. Step 2: Retain the encoder and decoder structure in the pre-trained model, and introduce a new segmentation decoder for new class land cover segmentation, replacing its convolutional layer with a low-rank fine-tuned convolutional layer to reduce the model's overfitting to small-sample new class data. Step 3: Train the new class land cover segmentation decoder, while decoupling the prediction results of the base class from the new class land cover segmentation results, maintaining the model's segmentation effect on the base class land cover while enhancing the model's learning of the new class. Step 4: Fuse the base class land cover segmentation prediction map and the new class land cover segmentation prediction map output in Step 3 to generate a final segmentation result map containing all classes.

[0010] The specific implementation process for each step is described below: 1. Training of the base class land cover segmentation model Suppose that in a few-sample semantic segmentation task, there are sufficient labeled images of base class features and K labeled images of new class features. Let the set of base class categories be denoted as... The new class set is denoted as ,in The number of base class categories, The number of new categories (where b and n are abbreviations for base and novel, respectively), and the background category is denoted as... This refers to pixels that do not belong to any target category. To achieve base class ground object segmentation, this invention employs an encoder-decoder architecture, which uses a ResNet-50 residual network as the encoder and a Feature Pyramid Network (FPN) as the decoder. Specifically, given an input remote sensing image... Where H and W are the length and width of the image, respectively, the residual network ResNet-50 first generates four feature maps at different resolutions. Where i ranges from 1 to 4, representing the index of the feature map at different resolutions. This represents the channel dimension of the feature map output in the initial stage. Subsequently, high-level contextual information and low-level texture features are fused using FPN, and the channels are mapped to the same dimension to generate new feature maps of different resolutions. Where C=128 represents the channel dimension of the mapping output, and the specific process is as follows: , in, and The transformation function for feature dimensionality reduction corresponding to the i-th resolution is applied to feature maps of different resolutions. and new feature maps at different resolutions The above method uses 3×3 convolutions, ReLU activation function, and batch normalization layers. `upsample` represents the upsampling function, implemented using bilinear interpolation. This method fuses feature maps of different resolutions from top to bottom, ultimately yielding the fused feature map. .

[0011] Then the fused feature map Using the number of channels The 3×3 convolution can be used to obtain the base class land cover segmentation prediction results. During the training of the encoder and base class segmentation decoder, the first cross-entropy loss function is used for optimization: , In the above formula This indicates whether the pixel belongs to the base class or the background class. For the one-hot encoded label of the pixel, the base class label is pixels, ,otherwise , This represents the probability that a pixel belongs to the background category in the base class feature segmentation prediction result. This indicates that the pixels in the base class land cover segmentation prediction results belong to the base class category. By optimizing the encoder and base class feature decoder using the above loss function, a base class feature segmentation model can be trained.

[0012] 2. Low-rank parameter reset To prevent overfitting to new classes due to fine-tuning on small sample datasets, this invention replaces the convolutional layers in the original decoder with low-rank fine-tuned convolutional layers. These convolutional layers, while inheriting knowledge from the original model, are fine-tuned using only a small number of parameters. The specific implementation is as follows: Figure 2 As shown, the input and output of the low-rank fine-tuned convolutional layer are the same as those of a regular convolutional layer, but the optimized parameters are only four low-rank matrix parameters. Specifically, the low-rank fine-tuned convolutional layer decomposes the learnable convolutional kernel parameters into the product of two low-rank matrices, as shown in the following equation: , in, Given the input feature map, This represents the output feature map. Represents the low-rank convolution kernel parameters. This represents the convolution kernel parameters of the original base class segmentation decoder. Represents the dot product of matrices. This represents the convolution operation. Indicates matrix dimension transformation. , , , These are learnable matrix parameters used to reconstruct the convolution kernel parameters. , Let k be the number of channels in the input and output feature maps, and k be the receptive field of the convolution kernel. This represents the hyperparameter of the reaction matrix rank. Low-rank fine-tuned convolutions can significantly improve training efficiency and prevent the model from overfitting on small samples. To preserve the basic knowledge of the model's pre-training, this invention uses learnable matrix parameters... Initialize to 0, other learnable matrix parameters , , Random initialization is used.

[0013] 3. Decoupling learning of new target features To prevent the training of new land cover classes from affecting the accuracy of base class target recognition, this invention employs a dual-branch decoupled network design. Specifically, during the fine-tuning of the new land cover segmentation decoder, the parameters of both the encoder and the base class segmentation decoder are frozen. Therefore, the model focuses solely on training the new land cover classes during training, without affecting the performance of the base classes. Specifically, the feature map obtained after FPN encoding by the new land cover segmentation decoder... Using a number of channels The 3×3 convolution yields the segmentation prediction results for the new land cover types. During the training of the new class segmentation decoder, only the parameters of the low-rank fine-tuning convolutional layers of the new class segmentation decoder are updated, and the second cross-entropy loss function is used for optimization. , in, This indicates that the pixel belongs to a new category or the background category. For the one-hot encoded label of the pixel, the category label for the new class is pixels, ,otherwise , This represents the probability that a pixel belongs to the background category in the new land cover segmentation prediction result. This indicates that the pixels in the new land cover segmentation prediction results belong to the base class category. Furthermore, since the parameters of the base class feature segmentation decoder are frozen during the fine-tuning phase, the segmentation effect of the base class feature will not be affected by the fine-tuning of the new class feature segmentation decoder.

[0014] 4. Segmentation result fusion After the above training, for an input remote sensing image, the model inference can obtain the base class land cover segmentation prediction result. New land cover segmentation prediction results However, during the training of base class feature segmentation, new class features are classified as background, and during the fine-tuning of new class feature segmentation, base class features are classified as background. Therefore, directly stitching and fusing the above results will result in most features being classified as background. To fuse the two results, this invention proposes a dedicated fusion strategy, as shown in the following equation: , in, For the final segmentation result, if neither decoder can identify a certain land cover category, then the pixel is classified as the background category. Since the base class decoder is trained on a large amount of data, if the base class decoder determines that a pixel belongs to the base class land cover category, then the base class category label is directly assigned to that pixel. If a pixel is identified as background by the base decoder but as foreground by the new category decoder, this strategy classifies it into the new category.

[0015] Contents not described in detail in this specification are common knowledge to those skilled in the art. Although illustrative specific embodiments of the invention have been described above to facilitate understanding by those skilled in the art, it should be understood that the invention is not limited to the scope of the specific embodiments. Various modifications will be readily apparent to those skilled in the art as long as they fall within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the inventive concept are protected.

Claims

1. A semantic segmentation method for small-sample remote sensing images based on low-rank feature decoupling, characterized in that, Includes the following steps: Step 1: Train on a large-scale dataset containing base class land cover object annotations to obtain a base class land cover segmentation model; the base class land cover segmentation model adopts an encoder-decoder architecture, wherein a residual network ResNet-50 is used as the encoder and a feature pyramid network FPN is used as the decoder. Step 2: Retain the encoder and base class segmentation decoder structures in the base class land cover segmentation model, and introduce a new class segmentation decoder for new class land cover segmentation. The convolutional layer of the new class segmentation decoder adopts a low-rank fine-tuned convolutional layer, which decomposes the learnable convolutional kernel parameters into the product of two low-rank matrices. Step 3: Train the new class segmentation decoder. During the training process, freeze the encoder and base class segmentation decoder in the base class land cover segmentation model, and only update the low-rank fine-tuning convolutional layer parameters of the new class segmentation decoder to decouple the base class land cover segmentation prediction results generated by the base class segmentation decoder from the new class land cover segmentation prediction results generated by the new class segmentation decoder. Step 4: Fuse the segmentation prediction results of the base class features and the segmentation prediction results of the new class features output in Step 3 to generate the final segmentation result containing all categories. The fusion strategy is as follows: when the base class segmentation decoder identifies a feature as a base class feature, the base class label is assigned first; when the base class segmentation decoder identifies a feature as background and the new class segmentation decoder identifies it as foreground, the new class label is assigned; otherwise, the background label is assigned.

2. The method for semantic segmentation of small-sample remote sensing images based on low-rank feature decoupling according to claim 1, characterized in that, Step 1 includes: Suppose that in a few-sample semantic segmentation task, there are sufficient labeled images of base class features and K labeled images of new class features. Let the set of base class features be denoted as Ki. The set of categories of new land features is denoted as ,in, The number of categories of base land cover. The number of new land cover categories, denoted as background category. Pixels that do not belong to any target category; given an input remote sensing image Where H and W are the length and width of the input remote sensing image, respectively, the residual network ResNet-50 first generates four feature maps at different resolutions. Where i ranges from 1 to 4, representing the index of the feature map at different resolutions. To determine the channel dimension of the feature map output in the initial stage, a Feature Pyramid Network (FPN) is used to fuse high-level contextual information and low-level texture features, mapping the feature map channels to the same dimension to generate new feature maps of different resolutions. Where C represents the channel dimension of the mapped output, and the new feature maps at different resolutions. The expression is: in, and The transformation function for feature dimensionality reduction corresponding to the i-th resolution is applied to feature maps of different resolutions. and new feature maps at different resolutions The above method uses 3×3 convolutions, ReLU activation function, and batch normalization layers. `upsample` represents the upsampling function, implemented using bilinear interpolation. This method fuses feature maps of different resolutions from top to bottom, ultimately yielding the fused feature map. ; For the fused feature map Using the number of channels The 3×3 convolution yields the base class land cover segmentation prediction result. During the training of the encoder and base class segmenter, the first cross-entropy loss function is used. Optimize: In the above formula This indicates whether the pixel belongs to the base class or the background class. For the one-hot encoded label of the pixel, the base class label is pixels, ,otherwise , Indicates the prediction results of base class land cover segmentation The probability that a mid-range pixel belongs to the background category. Indicates the prediction results of base class land cover segmentation Medium pixels belong to the base class category The probability is calculated, and the base class land cover segmentation model is obtained after optimization.

3. The semantic segmentation method for small-sample remote sensing images based on low-rank feature decoupling according to claim 1, characterized in that, In step 2, the low-rank fine-tuning convolutional layer decomposes the learnable convolutional kernel parameters into the product of two low-rank matrices, as shown in the following equation: in, Given the input feature map, This represents the output feature map. Represents the low-rank convolution kernel parameters. This represents the convolution kernel parameters of the base class segmentation decoder. Represents the dot product of matrices. This represents the convolution operation. Indicates matrix dimension transformation. , , , These are learnable matrix parameters used to reconstruct the convolution kernel parameters. , Let k be the number of channels in the input and output feature maps, and k be the receptive field of the convolution kernel. The hyperparameter of the rank of the reaction matrix is ​​the learnable matrix parameter. Initialize to 0, other learnable matrix parameters , , Perform random initialization.

4. The semantic segmentation method for small-sample remote sensing images based on low-rank feature decoupling according to claim 2, characterized in that, Step 3 includes: The feature map obtained by the new class segmentation decoder after encoding by the Feature Pyramid Network (FPN) Using a number of channels The 3×3 convolution yields the new land feature segmentation prediction results. During the training of the new class segmentation decoder, only the parameters of the low-rank fine-tuning convolutional layers of the new class segmentation decoder are updated, and the second cross-entropy loss function is used. Optimize: in, This indicates that the pixel belongs to a new category or the background category. For the one-hot encoded label of the pixel, the category label for the new class is pixels, ,otherwise , This represents the probability that a pixel belongs to the background category in the new land cover segmentation prediction result. Indicates the prediction results of new land cover segmentation Medium pixels belong to the base class category The probability is determined by freezing the encoder and base class segmentation decoder during training, so that the base class segmentation decoder is not affected by the fine-tuning of the new class segmentation decoder, thus decoupling the base class land cover segmentation prediction results from the new class land cover segmentation prediction results.

5. The semantic segmentation method for small-sample remote sensing images based on low-rank feature decoupling according to claim 1, characterized in that, Step 4 includes: Using a fusion strategy to segment and predict base class land cover results New land cover segmentation prediction results The fusion process is performed using the following strategy: in, For the final segmentation result, if neither decoder can identify a pixel as a certain land cover category, then the pixel is classified as the background category. If the base class segmentation decoder identifies a pixel as belonging to the base class land cover category, then the pixel is assigned a base class category label. If a pixel is identified as background by the base class segmentation decoder and as foreground by the new class segmentation decoder, then the pixel is assigned a new class label. .

Citation Information

Patent Citations

  • FPN-based steel surface defect detection method

    CN112053357A

  • Small sample target detection and identification method based on self-supervised ion separation space

    CN117876868A