A remote sensing image semantic segmentation method and device based on a frozen visual base model
Patent Information
- Application Number
- CN202610492564.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-04-15
AI Technical Summary
[0005]为解决现有技术需要消耗海量的计算资源和存储空间,计算复杂度较高,导致模型在实际部署时推理效率低下的问题,本发明的首要目的在于提供一种避免了大规模参数微调带来的计算和显存压力,大幅提高了模型的训练效率和硬件适用性的基于冻结视觉基础模型的遥感图像语义分割方法
[0044]由上述技术方案可知,本发明的有益效果为:第一,本发明通过完全冻结预训练的DINOv3模型骨干网络的参数,避免了大规模参数微调带来的计算和显存压力,大幅提高了模型的训练效率和硬件适用性;第二,本发明构建了轻量级的特征选择与融合模块,能够自适应地筛选冻结大模型输出的多层级特征,有效应对遥感图像中复杂的地物尺度变化;第三,本发明采用非对称分阶段注入上采样策略设计了跳跃融合上采样模块,通过巧妙利用大模型浅层特征中蕴含的丰富几何与边界信息,最大程度地恢复了空间细节,显著提高了模型在多类遥感地物上的整体分割精度和平均交并比。
Smart Images

Figure CN122368476B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and remote sensing image processing technology, and in particular to a method and device for semantic segmentation of remote sensing images based on a frozen vision model. Background Technology
[0002] Semantic segmentation of remote sensing images is a core technology in fields such as Earth observation and urban planning. With the development of sensor technology, high-resolution remote sensing images contain extremely rich information about ground features. However, remote sensing images are usually characterized by complex backgrounds, drastic changes in the scale of ground features, small inter-class differences and large intra-class variances. Traditional semantic segmentation models have limited ability to extract generalized features and cope with complex urban scenes.
[0003] In recent years, large-scale visual foundational models have demonstrated powerful feature representation and generalization capabilities. However, directly applying these visual foundational models to semantic segmentation of remote sensing images and performing full parameter fine-tuning requires massive computational resources and storage space. Furthermore, existing adaptation methods typically have high computational complexity, leading to low inference efficiency during practical deployment, and excessive parameter updates can easily destroy the general knowledge representations learned by pre-trained models on massive datasets.
[0004] Therefore, there is an urgent need for a remote sensing image semantic segmentation method that can fully utilize the powerful feature extraction capabilities of large visual models, while also possessing high computational efficiency and excellent spatial detail restoration capabilities. Summary of the Invention
[0005] To address the problem that existing technologies require massive amounts of computing resources and storage space, resulting in high computational complexity and low inference efficiency during actual deployment, the primary objective of this invention is to provide a remote sensing image semantic segmentation method based on a frozen vision model that avoids the computational and memory pressure caused by large-scale parameter fine-tuning, and significantly improves the training efficiency and hardware applicability of the model.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a semantic segmentation method for remote sensing images based on a frozen visual model, the method comprising the following sequential steps:
[0007] (1) Obtain a high-resolution remote sensing image dataset and preprocess it. Divide the preprocessed dataset into a training set and a test set.
[0008] (2) Construct a semantic segmentation model for remote sensing images. The semantic segmentation model for remote sensing images includes a pre-trained visual base model, a feature selection and fusion module, and a skip fusion upsampling module. The pre-trained visual base model adopts the pre-trained DINOv3 model. The feature selection and fusion module includes a multilayer perceptron, a stitching module, a softmax function, and a weighted fusion module. The skip fusion upsampling module includes transposed convolution, batch normalization, convolution with a kernel size of 1×1, convolution with a kernel size of 3×3, a ReLU activation function, and a thinning module.
[0009] (3) The training set is input into the remote sensing image semantic segmentation model for training. First, under the condition of freezing the parameters of the pre-trained visual base model, multi-level image features are extracted through the pre-trained visual base model. Then, the extracted multi-level image features are input into the feature selection and fusion module to filter the features at different levels and aggregate contextual information, and output the fused feature map. ; Fuse feature maps The data is input to the skip fusion upsampling module, which performs step-by-step upsampling by combining shallow spatial detail information, and then performs feature recovery to generate a high-resolution segmentation feature map, thus obtaining the trained model.
[0010] (4) Input the remote sensing image to be tested into the trained model and output the final semantic segmentation prediction result of the remote sensing image.
[0011] In step (1), the high-resolution remote sensing image dataset is any one of the ISPRS Vaihingen dataset, ISPRS Potsdam dataset, and LoveDA dataset; the preprocessing includes cropping, rotation, horizontal flipping, and vertical flipping.
[0012] In step (3), the specific processing steps of the feature selection and fusion module are as follows:
[0013] (4a) Receive multi-level feature maps from the frozen pre-trained visual base model Channel-level attention weights are obtained through a multilayer perceptron. The weight calculation process is as follows:
[0014] ;
[0015] in, It indicates the first The output features of the layers; MLP stands for Multilayer Perceptron. For the first The original weights of the layer, For batch size, The spatial dimension is determined by flattening the feature map; after concatenating the original weights of all layers, the softmax function is used for normalization to ensure that the sum of the weights for each sample and each spatial location is 1.
[0016] ;
[0017] in, Output the total number of layers for the backbone network, and satisfy the following condition for any sample: , represents the weights corresponding to the output features of the i-th layer after normalization; k is an intermediate variable, taking values from 1 to I;
[0018] (4b) Utilize The input feature map is filtered by channel features to suppress redundant background information and enhance feature channels sensitive to specific remotely sensed features. The filtered multi-level features are aligned in spatial dimensions, and a weighted summation operation is used to aggregate contextual information to generate a fused feature map containing global context. :
[0019]
[0020] In the formula, The feature embedding dimension.
[0021] In step (3), the specific processing steps of the skip fusion upsampling module are as follows:
[0022] (5a) Received fusion feature map Dimensional transformation and deformation are performed, and spatial information interaction is carried out on these newly fused, pixel-by-pixel customized features to obtain multi-dimensional features containing multi-layer information as the starting point for upsampling;
[0023] (5b) An asymmetric, phased injection upsampling strategy is adopted, which injects supplementary features in the first two injections and does not inject supplementary features in the last two injections. The upsampling starting point and supplementary features are fused through selective skip connections.
[0024] The first upsampling is represented as:
[0025] ;
[0026] The second upsampling is represented as:
[0027] ;
[0028] The third upsampling is represented as:
[0029] ;
[0030] The fourth upsampling is represented as:
[0031] ;
[0032] in, and All are supplementary features, among which, The highest-level output feature with the strongest semantic information in the pre-trained DINOv3 model output; The mid-level features with the strongest structural information in the output of the pre-trained DINOv3 model; For splicing operations; , , , These are the inputs for the first, second, third, and fourth upsampling operations, respectively. , , , These are the outputs for the first, second, third, and fourth upsampling operations, respectively. This is an interpolation operation used to standardize dimension size; For a transposed convolution with a kernel size of 2×2 and a stride of 2, This indicates a batch normalization operation. For convolutions with a kernel size of 1×1, It is the ReLU activation function; For the residual convolution module, the calculation process is as follows:
[0033] ;
[0034] in, , All are convolutions with a kernel size of 3×3. , Both indicate batch normalization operations. , These are the input and output of the residual convolution module, respectively.
[0035] (5c) Input the upsampled features into the thinning module for feature smoothing and channel dimensionality reduction to gradually restore the spatial boundary details of the remote sensing image:
[0036]
[0037] in, For a convolution with a kernel size of 3×3, For GELU activation function, The output after four upsampling steps is used as the input to the refinement module; To refine the module's output;
[0038] Finally, the data is classified using a 1×1 kernel convolution to generate a high-resolution segmentation feature map.
[0039] Step (4) specifically refers to: inputting the high-resolution remote sensing image to be tested into the trained model to obtain the segmentation feature map, using the Softmax function to calculate the probability distribution of each pixel belonging to different land cover categories, and selecting the category label corresponding to the maximum probability as the final predicted category of the pixel, and outputting the complete remote sensing image semantic segmentation map.
[0040] Another object of the present invention is to provide an electronic device comprising:
[0041] Processor; and
[0042] A memory storing computer program instructions that, when executed by the processor, cause the processor to perform the remote sensing image semantic segmentation method based on a frozen vision model as described above.
[0043] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the remote sensing image semantic segmentation method based on a frozen visual model as described above.
[0044] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, by completely freezing the parameters of the pre-trained DINOv3 model backbone network, the present invention avoids the computational and memory pressure caused by large-scale parameter fine-tuning, and greatly improves the training efficiency and hardware applicability of the model; Second, the present invention constructs a lightweight feature selection and fusion module, which can adaptively filter and freeze multi-level features output by large models, effectively dealing with complex land cover scale changes in remote sensing images; Third, the present invention adopts an asymmetric staged injection upsampling strategy to design a skip fusion upsampling module, which cleverly utilizes the rich geometric and boundary information contained in the shallow features of large models to restore spatial details to the greatest extent, and significantly improves the overall segmentation accuracy and average intersection-union ratio of the model on multiple types of remote sensing land cover. Attached Figure Description
[0045] Figure 1 This is a flowchart of the method of the present invention;
[0046] Figure 2 This is a schematic diagram of the skip fusion upsampling module in this invention;
[0047] Figure 3 This is a comparative diagram of the semantic segmentation results of remote sensing images obtained by the present invention. Detailed Implementation
[0048] like Figure 1As shown, a semantic segmentation method for remote sensing images based on a frozen visual model is proposed. This method includes the following sequential steps:
[0049] (1) Obtain a high-resolution remote sensing image dataset and preprocess it. Divide the preprocessed dataset into a training set and a test set.
[0050] (2) Construct a semantic segmentation model for remote sensing images. The semantic segmentation model for remote sensing images includes a pre-trained visual base model, a feature selection and fusion module, and a skip fusion upsampling module. The pre-trained visual base model adopts the pre-trained DINOv3 model. The feature selection and fusion module includes a multilayer perceptron, a stitching module, a softmax function, and a weighted fusion module. The skip fusion upsampling module includes transposed convolution, batch normalization, convolution with a kernel size of 1×1, convolution with a kernel size of 3×3, a ReLU activation function, and a thinning module.
[0051] (3) The training set is input into the remote sensing image semantic segmentation model for training. First, under the condition of freezing the parameters of the pre-trained visual base model, multi-level image features are extracted through the pre-trained visual base model. Then, the extracted multi-level image features are input into the feature selection and fusion module to filter the features at different levels and aggregate contextual information, and output the fused feature map. ; Fuse feature maps The data is input to the skip fusion upsampling module, which performs step-by-step upsampling by combining shallow spatial detail information, and then performs feature recovery to generate a high-resolution segmentation feature map, thus obtaining the trained model.
[0052] (4) Input the remote sensing image to be tested into the trained model and output the final semantic segmentation prediction result of the remote sensing image.
[0053] In step (1), the high-resolution remote sensing image dataset is any one of the ISPRS Vaihingen dataset, ISPRS Potsdam dataset, and LoveDA dataset; the preprocessing includes cropping, rotation, horizontal flipping, and vertical flipping.
[0054] In step (3), the specific processing steps of the feature selection and fusion module are as follows:
[0055] (4a) Receive multi-level feature maps from the frozen pre-trained visual base model Channel-level attention weights are obtained through a multilayer perceptron. The weight calculation process is as follows:
[0056] ;
[0057] in, It indicates the first The output features of the layers; MLP stands for Multilayer Perceptron. For the first The original weights of the layer, For batch size, The spatial dimension is determined by flattening the feature map; after concatenating the original weights of all layers, the softmax function is used for normalization to ensure that the sum of the weights for each sample and each spatial location is 1.
[0058] ;
[0059] in, Output the total number of layers for the backbone network, and satisfy the following condition for any sample: , represents the weights corresponding to the output features of the i-th layer after normalization; k is an intermediate variable, taking values from 1 to I;
[0060] (4b) Utilize The input feature map is filtered by channel features to suppress redundant background information and enhance feature channels sensitive to specific remotely sensed features. The filtered multi-level features are aligned in spatial dimensions, and a weighted summation operation is used to aggregate contextual information to generate a fused feature map containing global context. :
[0061]
[0062] In the formula, The feature embedding dimension.
[0063] like Figure 2 As shown, in step (3), the specific processing steps of the skip fusion upsampling module are as follows:
[0064] (5a) Received fusion feature map Dimensional transformation and deformation are performed, and spatial information interaction is carried out on these newly fused, pixel-by-pixel customized features to obtain multi-dimensional features containing multi-layer information as the starting point for upsampling;
[0065] (5b) An asymmetric, phased injection upsampling strategy is adopted, which injects supplementary features in the first two injections and does not inject supplementary features in the last two injections. The upsampling starting point and supplementary features are fused through selective skip connections.
[0066] The first upsampling is represented as:
[0067] ;
[0068] The second upsampling is represented as:
[0069] ;
[0070] The third upsampling is represented as:
[0071] ;
[0072] The fourth upsampling is represented as:
[0073] ;
[0074] in, and All are supplementary features, among which, The highest-level output feature with the strongest semantic information in the pre-trained DINOv3 model output; The mid-level features with the strongest structural information in the output of the pre-trained DINOv3 model; For splicing operations; , , , These are the inputs for the first, second, third, and fourth upsampling operations, respectively. , , , These are the outputs for the first, second, third, and fourth upsampling operations, respectively. This is an interpolation operation used to standardize dimension size; For a transposed convolution with a kernel size of 2×2 and a stride of 2, This indicates a batch normalization operation. For convolutions with a kernel size of 1×1, It is the ReLU activation function; For the residual convolution module, the calculation process is as follows:
[0075] ;
[0076] in, , All are convolutions with a kernel size of 3×3. , Both indicate batch normalization operations. , These are the input and output of the residual convolution module, respectively.
[0077] In the first two upsampling processes, the missing semantic and structural information of the starting feature is supplemented. In the last two upsampling processes, no feature supplementation is performed to prevent information redundancy and unnecessary waste of computing resources. At the same time, it also avoids the pollution of the final segmentation boundary by the high-frequency noise introduced by shallow features.
[0078] (5c) Input the upsampled features into the thinning module for feature smoothing and channel dimensionality reduction to gradually restore the spatial boundary details of the remote sensing image:
[0079]
[0080] in, For a convolution with a kernel size of 3×3, For GELU activation function, The output after four upsampling steps is used as the input to the refinement module; To refine the module's output;
[0081] Finally, the data is classified using a 1×1 kernel convolution to generate a high-resolution segmentation feature map.
[0082] Step (4) specifically refers to: inputting the high-resolution remote sensing image to be tested into the trained model to obtain the segmentation feature map, using the Softmax function to calculate the probability distribution of each pixel belonging to different land cover categories, and selecting the category label corresponding to the maximum probability as the final predicted category of the pixel, and outputting the complete remote sensing image semantic segmentation map.
[0083] like Figure 3 As shown, when processing remote sensing images of complex urban scenes (including impermeable surfaces, buildings, low vegetation, trees, vehicles, etc.) using this invention, a high-precision semantic segmentation map with clear boundaries and accurate category recognition can be generated through the collaborative work of freezing the DINOv3 large model and feature selection fusion, asymmetric skip upsampling, and refinement modules. Figure 3 In the diagram, (a) shows the high-resolution remote sensing image to be tested, containing categories such as impervious surfaces, buildings, low vegetation, and trees; (b) shows the actual category segmentation image; and (c) shows the segmentation result of this invention, with white markings corresponding to the impervious surface category, blue markings corresponding to the building category, cyan markings corresponding to the low vegetation category, and green markings corresponding to the tree category. Comparing images (b) and (c), it can be seen that the segmentation result of this invention is basically consistent with the actual segmentation image in terms of pixel category classification, and the segmentation result boundary is clear. Even on very narrow roads, this invention can accurately segment the data.
[0084] In summary, this invention avoids the computational and memory pressure caused by large-scale parameter fine-tuning by completely freezing the parameters of the pre-trained DINOv3 model backbone network, thus significantly improving the model's training efficiency and hardware applicability. This invention constructs a lightweight feature selection and fusion module that can adaptively filter and freeze multi-level features from the output of large models, effectively addressing complex scale variations of ground features in remote sensing images. This invention employs an asymmetric, staged injection upsampling strategy to design a skip fusion upsampling module, which cleverly utilizes the rich geometric and boundary information contained in the shallow features of large models to maximize the recovery of spatial details, significantly improving the overall segmentation accuracy and average intersection-over-union ratio of the model across multiple types of remote sensing ground features.
[0085] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A semantic segmentation method for remote sensing images based on a frozen vision model, characterized in that: The method includes the following steps in sequence: (1) Obtain a high-resolution remote sensing image dataset and preprocess it. Divide the preprocessed dataset into a training set and a test set. (2) Construct a semantic segmentation model for remote sensing images. The semantic segmentation model for remote sensing images includes a pre-trained visual base model, a feature selection and fusion module, and a skip fusion upsampling module. The pre-trained visual base model adopts the pre-trained DINOv3 model. The feature selection and fusion module includes a multilayer perceptron, a stitching module, a softmax function, and a weighted fusion module. The skip fusion upsampling module includes transposed convolution, batch normalization, convolution with a kernel size of 1×1, convolution with a kernel size of 3×3, a ReLU activation function, and a thinning module. (3) The training set is input into the remote sensing image semantic segmentation model for training. First, under the condition of freezing the parameters of the pre-trained visual base model, multi-level image features are extracted through the pre-trained visual base model. Then, the extracted multi-level image features are input into the feature selection and fusion module to filter the features at different levels and aggregate contextual information, and output the fused feature map. ; Fuse feature maps The data is input to the skip fusion upsampling module, which performs step-by-step upsampling by combining shallow spatial detail information, and then performs feature recovery to generate a high-resolution segmentation feature map, thus obtaining the trained model. (4) Input the remote sensing image to be tested into the trained model and output the final semantic segmentation prediction result of the remote sensing image; In step (3), the specific processing steps of the feature selection and fusion module are as follows: (4a) Receive multi-level feature maps from the frozen pre-trained visual base model Channel-level attention weights are obtained through a multilayer perceptron. The weight calculation process is as follows: ; in, It indicates the first The output features of the layers; MLP stands for Multilayer Perceptron. For the first The original weights of the layer, For batch size, The spatial dimension is determined by flattening the feature map; after concatenating the original weights of all layers, the softmax function is used for normalization to ensure that the sum of the weights for each sample and each spatial location is 1. ; in, Output the total number of layers for the backbone network, and satisfy the following condition for any sample: , represents the weights corresponding to the output features of the i-th layer after normalization; k is an intermediate variable, taking values from 1 to I; (4b) Utilize The input feature map is filtered by channel features to suppress redundant background information and enhance feature channels sensitive to specific remotely sensed features. The filtered multi-level features are aligned in spatial dimensions, and a weighted summation operation is used to aggregate contextual information to generate a fused feature map containing global context. : ; In the formula, The feature embedding dimension.
2. The remote sensing image semantic segmentation method based on a frozen vision model according to claim 1, characterized in that: In step (1), the high-resolution remote sensing image dataset is any one of the ISPRS Vaihingen dataset, ISPRS Potsdam dataset, and LoveDA dataset; the preprocessing includes cropping, rotation, horizontal flipping, and vertical flipping.
3. The remote sensing image semantic segmentation method based on a frozen vision model according to claim 1, characterized in that: In step (3), the specific processing steps of the skip fusion upsampling module are as follows: (5a) Received fusion feature map Dimensional transformation and deformation are performed, and spatial information interaction is carried out on these newly fused, pixel-by-pixel customized features to obtain multi-dimensional features containing multi-layer information as the starting point for upsampling; (5b) An asymmetric, phased injection upsampling strategy is adopted, which injects supplementary features in the first two injections and does not inject supplementary features in the last two injections. The upsampling starting point and supplementary features are fused through selective skip connections. The first upsampling is represented as: ; The second upsampling is represented as: ; The third upsampling is represented as: ; The fourth upsampling is represented as: ; in, and All are supplementary features, among which, The highest-level output feature with the strongest semantic information in the pre-trained DINOv3 model output; The mid-level features with the strongest structural information in the output of the pre-trained DINOv3 model; For splicing operations; , , , These are the inputs for the first, second, third, and fourth upsampling operations, respectively. , , , These are the outputs for the first, second, third, and fourth upsampling operations, respectively. This is an interpolation operation used to standardize dimension size; For a transposed convolution with a kernel size of 2×2 and a stride of 2, This indicates a batch normalization operation. For convolutions with a kernel size of 1×1, It is the ReLU activation function; For the residual convolution module, the calculation process is as follows: ; in, , All are convolutions with a kernel size of 3×3. , Both indicate batch normalization operations. , These are the input and output of the residual convolution module, respectively. (5c) Input the upsampled features into the thinning module for feature smoothing and channel dimensionality reduction to gradually restore the spatial boundary details of the remote sensing image: ; in, For a convolution with a kernel size of 3×3, For GELU activation function, The output after four upsampling steps is used as the input to the refinement module; To refine the module's output; Finally, the data is classified using a 1×1 kernel convolution to generate a high-resolution segmentation feature map.
4. The remote sensing image semantic segmentation method based on a frozen vision model according to claim 1, characterized in that: Step (4) specifically refers to: inputting the high-resolution remote sensing image to be tested into the trained model to obtain the segmentation feature map, using the Softmax function to calculate the probability distribution of each pixel belonging to different land cover categories, and selecting the category label corresponding to the maximum probability as the final predicted category of the pixel, and outputting the complete remote sensing image semantic segmentation map.
5. An electronic device, comprising: processor; as well as A memory storing computer program instructions that, when executed by the processor, cause the processor to perform the remote sensing image semantic segmentation method based on a frozen visual fundamental model as described in any one of claims 1-4.
6. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the remote sensing image semantic segmentation method based on a frozen visual fundamental model as described in any one of claims 1-4.
Citation Information
Patent Citations
Coding-decoding open vocabulary semantic segmentation method based on CLIP model
CN120563834A
Remote sensing image change detection method and device and storage medium
CN121616958A