An Efficient Spatial Target Segmentation Method Based on Vision Foundation Model

Through the efficient fine-tuning of parameters, prototype perception learning and consistency regularization strategies of visual basic models and object query guidance, the morphological differences and scarcity of labeled data in satellite image segmentation are solved, and efficient and robust satellite image segmentation is achieved, suitable for real-time applications of edge devices.

CN120236210BActive Publication Date: 2025-08-01NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510728912.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing satellite image segmentation technology faces the problems of significant differences in satellite morphology, scarce labeling data, computing resource constraints and inefficient model training. It is difficult to achieve high-quality segmentation under very small amounts of labeling data, and it is difficult to deploy efficiently on edge devices.

Method used

The visual basic model is used to combine the efficient fine-tuning of parameter guidance, prototype-aware learning and consistency regularization strategies, and efficient and robust satellite image segmentation is achieved by inserting learnable object query matrix and designing multi-grained prototype vectors.

Benefits of technology

Implementing high-quality satellite image segmentation with a very small amount of labeled data improves the generalization ability and adaptability of the model, and can run in real time on edge devices to meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236210B_ABST
    Figure CN120236210B_ABST
Patent Text Reader

Abstract

The present invention discloses an efficient spatial target segmentation method based on a vision foundation model. By leveraging the powerful prior knowledge of the vision foundation model, combining parameter-efficient fine-tuning guided by object queries and prototype-aware learning, and a consistency regularization strategy, efficient and high-quality satellite image segmentation is achieved under the condition of extremely small amounts of labeled data. The present invention can effectively solve problems such as large differences in satellite morphology, scarce labeled data, and low model training efficiency, significantly improving the generalization ability and practicality of satellite image segmentation, and providing reliable basic support for downstream tasks such as attitude estimation and visual measurement of satellite systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of aerospace technology, and particularly relates to an efficient space target segmentation method based on a vision foundation model. Background Art

[0002] With the development of aerospace technology, satellite imaging has been increasingly widely used in fields such as earth observation, navigation and positioning, and intelligent remote sensing. As a basic link in satellite image analysis, satellite image segmentation is crucial for downstream tasks such as attitude estimation and visual measurement, directly affecting the reliability and effectiveness of precise control and mission planning, and is one of the key technologies to promote the progress of intelligent aerospace technology.

[0003] However, the existing satellite image segmentation technologies face the following technical problems:

[0004] 1) Significant differences in satellite morphology: There are significant differences in the morphological structures of different types of satellites, resulting in the difficulty for a model trained on one dataset to effectively adapt to other types of satellites, greatly limiting the generalization ability of the model in practical applications. Traditional methods are usually designed for specific types of satellites and are difficult to handle unseen satellite morphologies.

[0005] 2) Scarce labeled data: Obtaining high-quality labeled data for satellite image segmentation is costly and time-consuming, especially the labeled data for new types of satellites is often extremely limited and cannot support the complete training process of traditional deep learning models.

[0006] 3) Computational resource constraints: In actual deployment environments, especially on edge devices, computational resources are often strictly limited and it is difficult to run complex satellite image segmentation models.

[0007] 4) Low model training efficiency: Existing methods need to retrain the model from scratch, which is time-consuming and relies on a large amount of labeled data, and it is difficult to meet the requirements of rapid iterative updates.

[0008] The existing satellite image segmentation technologies mainly include the following categories:

[0009] 1) Methods based on convolutional neural networks (CNNs): Such as DASPOCNET and OCRNet, etc. These methods mainly rely on complex convolutional architectures and attention mechanisms. Although they have achieved good performance on the seen satellite types, their generalization ability is limited and the computational overhead is large.

[0010] 2) Methods based on Transformers: Such as models based on MiT-B3 and MiT-B5. These methods use self-attention mechanisms to capture global dependencies, but they are still difficult to effectively handle the large differences in satellite morphologies and perform poorly in cross-domain adaptability.

[0011] 3) Traditional domain adaptation methods: Such as unsupervised domain adaptation (UDA) technology, which reduces the feature distribution difference between the source domain and the target domain through mechanisms such as adversarial learning. However, in the case of extremely few or no annotations, the performance still drops significantly.

[0012] These existing methods are difficult to solve the following problems simultaneously: (1) How to achieve high-quality segmentation under the condition of extremely few or even single annotated images; (2) How to make full use of the prior knowledge of pre-trained vision foundation models; (3) How to effectively handle the huge differences between satellite morphologies; (4) How to implement a computationally efficient and easily deployable segmentation system. Summary of the Invention

[0013] To overcome the deficiencies of the prior art, the present invention provides an efficient spatial target segmentation method based on a vision foundation model. By utilizing the powerful prior knowledge of the vision foundation model, combining object query-guided parameter-efficient fine-tuning and prototype-aware learning, as well as a consistency regularization strategy, it realizes efficient and high-quality satellite image segmentation under the condition of extremely few annotated data. The present invention can effectively solve problems such as large differences in satellite morphologies, scarce annotated data, and low model training efficiency, significantly improving the generalization ability and practicality of satellite image segmentation, and providing reliable basic support for downstream tasks such as attitude estimation and visual measurement of satellite systems.

[0014] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0015] Step 1: Feature extraction and evaluation of the vision foundation model;

[0016] Step 2: Parameter-efficient fine-tuning based on object queries;

[0017] Step 3: Prototype-aware learning and distribution consistency discrimination;

[0018] Step 4: Consistency learning method and robustness enhancement.

[0019] Preferably, the specific content of Step 1 is as follows:

[0020] Step 1-1: Evaluation of multiple vision foundation models;

[0021] Evaluate the performance of multiple vision foundation models in the satellite image segmentation task. All vision foundation models use the same decoder DAFormer to generate semantic predictions;

[0022] The mathematical representation of the evaluation process is as follows:

[0023]

[0024] Among them, represents the The feature extraction part of a vision foundation model represents the input satellite image, represents the generated pixel-level prediction result; represents the decoder head;

[0025] Step 1-2: Benchmark performance establishment;

[0026] Determine the benchmark performance of each vision foundation model through cross-validation, using the mean intersection over union (mIoU) as the evaluation metric:

[0027]

[0028] wherein, represents the number of classes, represents the predicted region of the model for class and represents the true annotation region of class ;

[0029] Step 1-3: Feature visualization and analysis;

[0030] Use the t-SNE dimensionality reduction method to perform visual analysis on the features and evaluate the representation ability of each model for satellite components:

[0031]

[0032] wherein, represents the feature representation after dimensionality reduction for visual analysis; represents the t-SNE dimensionality reduction method.

[0033] Preferably, the vision foundation model includes CLIP, EVA02, SAM, and DINOv2.

[0034] Preferably, the specific content of step 2 is as follows:

[0035] Step 2-1: Object query initialization and insertion;

[0036] Initialize the learnable object query matrix and insert it between different layers of the vision foundation model:

[0037]

[0038] wherein, represents the object query matrix of the th layer, represents the length of the object query, represents the th layer feature dimension;

[0039] Step 2-2: Query-guided feature refinement;

[0040] Feature refinement is achieved through the attention interaction between queries and features:

[0041]

[0042] Among them, represents the original features extracted by the th layer of the visual backbone model, represents the refined features, The function ensures the probability weighting of feature components; represents matrix transpose;

[0043] Step 2-3: Feature fusion and transmission;

[0044] Concatenate and fuse the original features and the refined features:

[0045]

[0046] Among them, is the concatenated feature, and respectively represent the weights and biases of the MLP, which are shared among different layers, represents concatenating two feature vectors;

[0047] Step 2-4: Forward propagation and gradient update;

[0048] During forward propagation, the parameters of the visual backbone model are frozen, and only the object query matrix and the parameters , participate in gradient update:

[0049]

[0050] Among them, represents the loss function of the segmentation task; respectively represent the gradient of the loss function with respect to the query matrix, the gradient of the loss function with respect to the weight parameter, and the gradient of the loss function with respect to the bias parameter.

[0051] Preferably, the specific steps of step 3 are as follows:

[0052] Step 3-1: Prototype generation and probability-driven optimization;

[0053] Generate initial prototype vectors for each satellite component based on the labeled data and feature representations:

[0054]

[0055] Among them, Represents the category index, Represents the indicator function, Represents the image annotation, Represents the initial prototype vector of the Represents the number of pixel rows of the input image, Represents the number of pixel columns of the input image, Represents the feature vector at the image position (i,j);

[0056] Calculate the semantic similarity between the global object query and the initial prototype vector, and establish the correlation reflecting the alignment degree between the query and different satellite components:

[0057]

[0058] Among them, Represents the global object query, Represents the similarity between the k-th query and the m-th category prototype;

[0059] Through the optimization process driven by the probability distribution, obtain the final set of prototype vectors:

[0060]

[0061] Among them, is the weight coefficient;

[0062] Step 3-2: Distribution consistency discrimination;

[0063] For each pixel, first extract its feature representation, and then calculate the distribution similarity vector between this feature and the global object query:

[0064]

[0065] Among them, represents the distribution similarity vector between the i-th pixel in the image and the global object query; ; represents the k-th correction coefficient;

[0066] Use the cosine similarity to calculate the similarity between the distribution similarity vector and the final prototype vector :

[0067]

[0068] Among them, represents the cosine similarity function;

[0069] Finally, the pixel is mapped to the category with the highest similarity:

[0070]

[0071] wherein, represents the category index that makes the expression in the parentheses reach the maximum value , represents the th predicted category label of the pixel;

[0072] Step 3-3: Multi-granularity segmentation architecture;

[0073] By designing prototype vectors of different granularities, the segmentation control from coarse-grained to fine-grained is realized;

[0074] Step 3-4: Training objective;

[0075] During the training process, the model weights are updated through the following loss function:

[0076]

[0077] wherein, represents the total number of pixels in the image; represents the th true category label of the pixel.

[0078] Preferably, the specific step 4 is as follows:

[0079] Step 4-1: Scale consistency regularization;

[0080] Perform multi-scale transformation on the input image and force the model to maintain prediction consistency at different scales:

[0081]

[0082] wherein, represents the scaled image, and respectively represent the feature representations of the original image and the scaled image;

[0083] Apply the smooth L1 loss function to enhance the feature consistency regularization:

[0084]

[0085] wherein, is defined as:

[0086]

[0087] Step 4-2: Overall training objective;

[0088] The final training objective is as follows:

[0089]

[0090] Among them, is the balance parameter, which is used to control the intensity of consistency regularization.

[0091] A computer program that causes a computer to execute the above-mentioned efficient spatial object segmentation method.

[0092] An electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned efficient spatial object segmentation method.

[0093] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned efficient spatial object segmentation method is implemented.

[0094] A chip, including: a processor, which is used to call and run a computer program from a memory, so that a device installed with the chip executes the above-mentioned efficient spatial object segmentation method.

[0095] The beneficial effects of the present invention are as follows:

[0096] 1. Because the parameter-efficient fine-tuning method based on object queries is adopted, the present invention can solve the problems of traditional methods relying on a large amount of labeled data and computing resources, and achieve the effect of efficient training under a single labeled image. By inserting learnable object queries between different layers of the vision base model and keeping the main parameters frozen, the present invention greatly reduces the number of trainable parameters (as low as 5% of the full amount of parameters), and at the same time effectively retains the prior knowledge of the vision base model, avoiding the overfitting problem under limited data conditions.

[0097] 2. Because the prototype-aware learning and distribution consistency discrimination method is adopted, the present invention can solve the problem that the traditional pixel classification method has weak generalization ability due to the significant morphological differences of satellites, and achieve a more efficient cross-domain adaptation effect. By transforming the traditional pixel-to-pixel segmentation paradigm into a pixel-to-prototype mapping and introducing a probability-driven prototype generation strategy, the present invention effectively handles the morphological changes of satellite components among different types, and at the same time realizes flexible segmentation control from coarse-grained to fine-grained.

[0098] 3. Due to the adoption of the multi-scale and context consistency regularization strategy, the present invention can solve the problem of insufficient model generalization ability under limited supervision and achieve a more robust segmentation performance effect. By forcing the model to maintain prediction consistency under different scales and context conditions, the present invention significantly improves the model's adaptability to scale changes, perspective changes, and imaging condition changes.

[0099] 4. Due to the adoption of the lightweight design and deployment optimization strategy, the present invention can solve the problem that complex models are difficult to deploy in resource-constrained environments and achieve efficient operation on edge devices. Through parameter-efficient design and inference optimization, the present invention realizes real-time operation on edge computing platforms such as NVIDIA Jetson AGX Xavier, meeting the actual application requirements.

[0100] In summary, the present invention breaks through the bottleneck of traditional satellite image segmentation technology. Through intelligent analysis and segmentation driven by the vision foundation model, it significantly improves the efficiency, accuracy, and adaptability of satellite image segmentation, providing reliable technical support for downstream tasks such as satellite attitude estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 is the flowchart of the method of the present invention;

[0102] Figure 2 is the effect diagram of the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0103] The present invention will be further described below in conjunction with the drawings and embodiments.

[0104] The present invention provides an efficient satellite image segmentation technology based on a vision foundation model. By utilizing the powerful prior knowledge of the vision foundation model, combining parameter-efficient fine-tuning guided by object queries and prototype-aware learning, and a consistency regularization strategy, it realizes efficient and high-quality satellite image segmentation under extremely small amounts of labeled data.

[0105] The present invention can effectively solve problems such as large differences in satellite morphology, scarce labeled data, and low model training efficiency, significantly improving the generalization ability and practicality of satellite image segmentation, and providing reliable basic support for downstream tasks such as satellite system attitude estimation and visual measurement.

[0106] Step 1: Feature extraction and evaluation of the vision foundation model;

[0107] The purpose of this step is to evaluate and select a suitable vision foundation model as the backbone network for feature extraction, laying a foundation for subsequent parameter-efficient fine-tuning.

[0108] Step 1-1: Evaluation of multiple vision foundation models;

[0109] Evaluate the performance of multiple vision backbone models on satellite image segmentation tasks, including CLIP, EVA02, SAM, and DINOv2, etc. To ensure a fair comparison, all models use the same decoder head (DAFormer) to generate semantic predictions;

[0110] The mathematical representation of the evaluation process is as follows:

[0111]

[0112] Step 1-2: Establish baseline performance;

[0113] Determine the baseline performance of each vision backbone model through cross-validation, using the mean Intersection over Union (mIoU) as the evaluation metric:

[0114]

[0115] Step 1-3: Feature visualization and analysis;

[0116] Use the t-SNE dimensionality reduction method to visualize and analyze the features, and evaluate the representation ability of each model for satellite components:

[0117]

[0118] The input of this step is multiple vision backbone models and satellite image datasets, and the output is the performance evaluation results and feature analysis reports of each vision backbone model, providing a basis for subsequent model selection.

[0119] Step 2: Parameter-efficient fine-tuning based on object queries;

[0120] This step mainly proposes a parameter-efficient fine-tuning method. By introducing object queries, efficient learning of satellite-specific semantic knowledge is achieved while freezing the main parameters of the vision backbone model.

[0121] Step 2-1: Object query initialization and insertion;

[0122] Initialize the learnable object query matrix and insert it between different layers of the vision backbone model:

[0123]

[0124] Step 2-2: Query-guided feature refinement;

[0125] Achieve feature refinement through the attention interaction between queries and features:

[0126]

[0127] Step 2-3: Feature fusion and transmission;

[0128] Concatenate and fuse the original features and the refined features:

[0129]

[0130] Step 2-4: Forward propagation and gradient update;

[0131] During the forward propagation process, the parameters of the visual base model are kept frozen, and only the object query matrix and the parameters 、 participate in the gradient update:

[0132]

[0133] The input of this step is the visual base model selected in the previous step and a small number of labeled target satellite images, and the output is the refined feature representation and the trained object query matrix, effectively capturing the satellite-specific semantic knowledge. By only updating a small number of parameters (object query and fusion layer), this step greatly improves the training efficiency and avoids overfitting, while retaining the powerful prior knowledge of the visual base model.

[0134] Step 3: Prototype-aware learning and distribution consistency discrimination;

[0135] In this step, by introducing a prototype-aware learning framework, the traditional pixel-to-pixel segmentation paradigm is transformed into a pixel-to-prototype mapping, effectively handling the satellite morphology difference problem and achieving flexible control of the segmentation granularity.

[0136] Step 3-1: Prototype generation and probability-driven optimization;

[0137] Generate the initial prototype vectors of each satellite component based on the labeled data and the feature representation:

[0138]

[0139] Calculate the semantic similarity between the global object query and the initial prototype vectors, and establish the correlation reflecting the alignment degree between the query and different satellite components:

[0140]

[0141] Obtain the final set of prototype vectors through a probability distribution-driven optimization process:

[0142]

[0143] Step 3-2: Distribution consistency discrimination;

[0144] For each pixel, first extract its feature representation, and then calculate the distribution similarity vector between this feature and the global object query:

[0145]

[0146] Calculate the similarity between the distribution similarity vector and the final prototype vector using cosine similarity:

[0147]

[0148] Finally, the pixel is mapped to the category with the highest similarity:

[0149]

[0150] Step 3-3: Multi-granularity segmentation architecture;

[0151] By designing prototype vectors of different granularities, flexible segmentation control from coarse-grained to fine-grained is achieved. For example, for a satellite antenna, a coarse-grained "antenna" category can be designed, as well as multiple fine-grained sub-categories such as "antenna 1", "antenna 2", "antenna 3", etc.;

[0152] Step 3-4: Training objective;

[0153] During the training process, the model weights are updated through the following loss function:

[0154]

[0155] The input of this step is the refined features and object query matrix output in Step 2, as well as a small amount of labeled data; the output is a set of probability-driven prototype vectors and a trained distribution consistency discriminant model. By transforming the traditional pixel classification paradigm into a prototype matching paradigm, this step effectively addresses the problem of large morphological variations of satellite components and achieves flexible segmentation control from coarse-grained to fine-grained;

[0156] Step 4: Consistency learning method and robustness enhancement;

[0157] In this step, by introducing multi-scale and context consistency regularization strategies, the generalization ability and robustness of the model under limited supervision are further improved.

[0158] Step 4-1: Scale consistency regularization;

[0159] Perform multi-scale transformation on the input image and force the model to maintain prediction consistency at different scales:

[0160]

[0161] Apply the smooth L1 loss function to enhance feature consistency:

[0162]

[0163] Among them, is defined as:

[0164]

[0165] Step 4-2: Overall training objective;

[0166] The final training objective is:

[0167]

[0168] The input of this step is the model and training data output from the previous step, and the output is a robust satellite image segmentation model enhanced by consistency regularization. By forcing the model to maintain prediction consistency under different scales and context conditions, this step effectively improves the generalization ability and robustness of the model, enabling it to better adapt to complex and changing satellite observation conditions.

[0169] Step 5: Model deployment and optimization;

[0170] This step mainly focuses on the actual deployment and performance optimization of the model to ensure the efficient operation of the system in resource-constrained environments.

[0171] Step 5-1: Analysis of the number of parameters and computational complexity:

[0172] Evaluate the number of parameters and computational complexity of the model and compare with traditional methods:

[0173]

[0174] Among them, represents the number of trainable parameters of the method of the present invention, represents the number of parameters of the full-scale fine-tuning baseline.

[0175] Step 5-2: Inference time optimization;

[0176] Further optimize the inference time through techniques such as model pruning and knowledge distillation:

[0177]

[0178] Among them, represents the inference time, represents the model parameters, represents the hardware platform, represents the input data scale.

[0179] Step 5-3: Edge device deployment:

[0180] Deploy and optimize the model on edge computing platforms such as NVIDIA Jetson AGX Xavier to ensure real-time performance:

[0181]

[0182] Among them, represents the number of frames processed per second, which measures the real-time performance of the model;

[0183] The input of this step is the model trained in the previous step, and the output is the satellite image segmentation system optimized for the target deployment environment. Through parameter quantity control, inference optimization, and platform adaptation, this step ensures the efficient and real-time operation of the system, meeting the actual application requirements.

[0184] Test results:

[0185] Table 1: IOU comparison of different methods under the SatelliteDataset → Speed+ generalization setting

[0186] Method Number of trainable parameters (M) Training time (h) mIoU (%) Main body Solar panel Antenna Average DASPOCNET 42.3 8.5 80.0 84.2 58.8 74.3 OCRNet 45.6 9.2 80.3 83.9 58.5 74.2 Domain generalization (DG) 35.6 4.6 67.9 69.6 0.0 45.8 Unsupervised domain adaptation (UDA) 35.6 5.3 85.0 72.3 1.9 53.1 The present invention (single sample) 1.8 0.3 88.5 80.1 72.6 80.4 The present invention (five samples) 1.8 0.4 93.5 89.9 80.7 88.0 The present invention (ten samples) 1.8 0.5 95.2 92.8 85.6 91.2

[0187] In the table, the main body represents the IoU of the main structure segmentation of the satellite, the solar panel represents the IoU of the solar panel area segmentation, and the antenna represents the IoU of the component segmentation; the average represents the average value of the IoUs of the three categories.

[0188] It can be seen from the test results that the method of the present invention is significantly superior to the existing methods when only using a single labeled image. Especially in the most challenging task of antenna segmentation, the mIoU of the method of the present invention reaches 72.6%, while the traditional domain generalization method and the unsupervised domain adaptation method are only 0.0% and 1.9% respectively. More importantly, the number of trainable parameters of the present invention is only 1.8M, and the training time is only 0.3 hours, reducing the number of parameters by more than 95% and the training time by more than 90% compared with the traditional methods, fully demonstrating the efficiency and effectiveness of the method.

[0189] As the number of training samples increases, the performance of the method of the present invention is further improved. When using 10 samples, the average mIoU reaches 91.2%, exceeding the best supervised baseline. This result fully verifies the scalability and potential of the method of the present invention.

Claims

1. An efficient spatial target segmentation method based on a vision foundation model, characterized in that, It includes the following steps: Step 1: Visual base model feature extraction and evaluation; Evaluate the performance of multiple visual base models on the satellite image segmentation task. All visual base models use the same decoder DAFormer to generate semantic predictions; Determine the baseline performance of each visual base model through cross-validation, using the mean intersection over union (mIoU) as the evaluation metric; Use the t-SNE dimensionality reduction method to visually analyze the features and evaluate the representation ability of each model for satellite components; Step 2: Parameter-efficient fine-tuning based on object queries; Step 2-1: Object query initialization and insertion; Initialize the learnable object query matrix and insert it between different layers of the visual base model: ; Among them, represents the object query matrix of the -th layer, represents the length of the object query, represents the -dimensionality of the layer features; Step 2-2: Query-guided feature refinement; Achieve feature refinement through the attention interaction between the query and the features: ; Among them, represents the original features extracted from the layer of the visual base model, represents the refined features, The function ensures the probability weighting of the feature components; [[ID=⑨]] [[ID=⑩]] represents matrix transpose; Step 2-3: Feature fusion and transfer; Concatenate and fuse the original features and the refined features: ; Among them, is the concatenated feature, and represent the weights and biases of the MLP respectively, which are shared between different layers, represents concatenating two feature vectors; Step 2-4: Forward propagation and gradient update; During the forward propagation process, the parameters of the vision backbone model remain frozen, and only the object query matrix and the parameters , participate in the gradient update: ; Among them, represents the loss function for the segmentation task; respectively represent the gradient of the loss function with respect to the query matrix, the gradient of the loss function with respect to the weight parameter, and the gradient of the loss function with respect to the bias parameter; Step 3: Prototype-aware learning and distribution consistency discrimination; Generate initial prototype vectors for each satellite component based on the annotated data and feature representations; Calculate the semantic similarity between the global object query and the initial prototype vectors, and establish the correlation reflecting the alignment degree between the query and different satellite components; Obtain the final set of prototype vectors through an optimization process driven by probability distribution; For each pixel, first extract its feature representation, and then calculate the distribution similarity vector between this feature and the global object query; Use cosine similarity to calculate the similarity between the distribution similarity vector and the final prototype vectors; The pixel is mapped to the category with the highest similarity; Achieve segmentation control from coarse-grained to fine-grained by designing prototype vectors with different granularities; During the training process, the model weights are updated through the loss function; Step 4: Consistency learning method and robustness enhancement; Perform multi-scale transformation on the input image and force the model to maintain prediction consistency at different scales; Apply the smooth L1 loss function to enhance feature consistency regularization; Determine the overall training objective.

2. The efficient spatial target segmentation method based on a vision foundation model according to claim 1, wherein The specific content of Step 1 is as follows: Step 1-1: Evaluation of multiple visual base models; Evaluate the performance of multiple visual base models on the satellite image segmentation task. All visual base models use the same decoder DAFormer to generate semantic predictions; The mathematical representation of the evaluation process is as follows: ; Among them, represents the feature extraction part of the th visual basic model, represents the input satellite image, represents the generated pixel-level prediction result; represents the decoding head; Step 1-2: Establishment of baseline performance; Determine the baseline performance of each visual base model through cross-validation, using the mean intersection over union (mIoU) as the evaluation metric: ; Among them, represents the number of categories, represents the prediction area of the model for the category ; represents the true annotation area of the category; Step 1-3: Feature visualization and analysis; Use the t-SNE dimensionality reduction method to visually analyze the features and evaluate the representation ability of each model for satellite components: ; Among them, represents the feature representation after dimensionality reduction and is used for visual analysis; represents the t-SNE dimensionality reduction method.

3. The efficient spatial target segmentation method based on a vision foundation model according to claim 2, wherein The visual base models include CLIP, EVA02, SAM, and DINOv2.

4. An efficient spatial target segmentation method based on a vision foundation model according to claim 3, characterized in that, The specific content of Step 3 is as follows: Step 3-1: Prototype generation and probability-driven optimization; Generate initial prototype vectors for each satellite component based on the annotated data and feature representations: ; Among them, represents the class index, represents the indicator function, represents the image annotation, represents the initial prototype vector of the represents the number of pixel rows of the input image, represents the number of pixel columns of the input image, represents the feature vector at the image position (i,j); Calculate the semantic similarity between the global object query and the initial prototype vectors, and establish the correlation reflecting the alignment degree between the query and different satellite components: ; Among them, represents a global object query, represents the th query and the similarity with the th category prototype; represents the number of layers of the network; Through an optimization process driven by probability distribution, a final set of prototype vectors is obtained: ; Among them, is the weight coefficient; Step 3-2: Distribution consistency discrimination; For each pixel, first extract its feature representation, and then calculate the distribution similarity vector between this feature and the global object query: ; Among them, represents the distribution similarity vector of the -th pixel in the image and the global object query; ; represents the k-th correction coefficient; Calculate the similarity between the distribution similarity vector and the final prototype vector using cosine similarity : ; Among them, represents the cosine similarity function; Finally, the pixel is mapped to the category with the highest similarity: ; Among them, represents the category index that makes the expression in the parentheses reach the maximum value , represents the -th predicted category label of the pixel; Step 3-3: Multi-granularity segmentation architecture; By designing prototype vectors with different granularities, segmentation control from coarse-grained to fine-grained is achieved; Step 3-4: Training objective; During the training process, the model weights are updated through the following loss function: ; Among them, represents the total number of pixels in the image; represents the true class label of the 5. An efficient spatial target segmentation method based on a vision-based foundation model according to claim 4, characterized in that The specific content of step 4 is as follows: Step 4-1: Scale consistency regularization; Perform multi-scale transformation on the input image and force the model to maintain prediction consistency at different scales: ; Among them, represents the scaled image, and respectively represent the feature representations of the original image and the scaled image; Apply the smooth L1 loss function to enhance feature consistency regularization: ; Among them, is defined as: ; Step 4-2: Determine the overall training objective; The final training objective is: ; Among them, is a balance parameter used to control the strength of consistency regularization.

6. An electronic device, characterized in that, Including: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 5 is implemented.

8. A chip, characterized in that, Including: A processor, which is used to call and run a computer program from the memory, so that the device installed with the chip executes the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Fine-grained multi-mode prompt learning method based on visual language pre-training model

    CN119538179A

  • Few-sample defect detection method based on prototype prompt fine-tuning visual basic model SAM

    CN119624927A