A specific object remote sensing extraction method and system based on visual transformer

CN121170616BActive Publication Date: 2026-09-29WENCHANG AEROSPACE SUPERCOMPUTING SMART TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261530.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-09-29
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

[0004]分块与拼接策略僵硬:通常采用固定重叠比例和概率均值拼接,无视模型置信度与局部不确定性,容易出现拼接边缘断裂或伪影问题;

Benefits of technology

[0027]本发明中,通过对云量、气溶胶光学厚度(AOD)、降雨等环境参数进行评估和筛选,系统在数据输入端即实现质量控制,显著降低因云雾、降雨影响导致的误判,提升模型输入样本的可靠性,从而提高整个目标提取流程的准确率。其次,本发明采用自适应分块机制,根据质量评分为高、中、低质量区分别设置不同的重叠比例与多时相输入方式,高质量区实现快速精准处理,中质量区通过密集重叠与多时相信息增强识别稳健性,低质量区则交由人工复核,有效避免无效计算与噪声干扰,系统运行更高效、结果更可信。此外,多源时相Transformer模型通过光学与时序双分支融合,结合交叉注意力实现空间纹理与时序变化的综合理解,不仅增强了对撂荒地边缘、形态及动态状态变化的捕捉,对轻度云遮或影像干扰环境也表现出较强的鲁棒性。解码器输出的不确定度图及分类结果能够细致量化模型信心,根据熵值判定像素可信度,并在拼接阶段依据不确定度进行后续重推理或人工复核,从理论上确保边界连贯性和整体精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170616B_ABST
    Figure CN121170616B_ABST
Patent Text Reader

Abstract

The application discloses a specific object remote sensing extraction method and system based on a visual Transformer, and belongs to the technical field of remote sensing image processing and deep learning. The method first extracts quality parameters such as cloud blocking rate, AOD and wet state from high-resolution images, calculates a comprehensive quality score for each sub-block, and classifies the sub-blocks into high-quality, medium-quality and low-quality areas; different processing strategies are adopted for different quality areas; a CNN+Transformer double-branch model is trained, and finally a high-precision specific object classification map is generated, and an uncertainty heat map and a to-be-reviewed label are output. The system has the technical advantages of high precision, strong robustness, explainable results and closed-loop optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of remote sensing image processing and deep learning, and in particular to a method and system for remote sensing extraction of specific objects based on visual Transformer. Background Technology

[0002] In recent years, semantic segmentation technology for remote sensing images has developed rapidly, especially with significant progress made using visual Transformer-based methods. Traditional convolutional neural networks (CNNs) excel at capturing local texture and structural information, but they often suffer from imprecise edge detection and misclassification when dealing with features that are complex and dispersed (such as abandoned land, roads, and building debris). Furthermore, existing object extraction techniques fail to adequately consider the robustness of models to environmental disturbances such as clouds, shadows, and haze, specifically in the following ways:

[0003] Quality parameters are not fully incorporated: Environmental factors such as aerosol concentration (AOD), terrain slope, and precipitation conditions are not processed, which may lead to identification bias due to inputting low-quality sub-maps.

[0004] rigid segmentation and stitching strategies: They usually use fixed overlap ratios and probability mean stitching, ignoring model confidence and local uncertainty, which can easily lead to stitching edge breakage or artifacts.

[0005] No compensation mechanism for low-to-medium quality areas: For areas with slight occlusion or partial shadow, existing technologies often adopt a method of skipping all or making rough predictions, resulting in the omission of potential target objects;

[0006] Lack of a closed-loop system design: The lack of manual review, feedback mechanisms, and dynamic update strategies makes it impossible to continuously optimize model performance in response to errors. Summary of the Invention

[0007] The purpose of this invention is to overcome the aforementioned deficiencies in the existing technology and propose a method and system for remote sensing extraction of specific objects based on visual Transformer. Starting from the data quality level, it combines multi-source temporal information, adaptive block strategy and uncertainty-driven stitching mechanism to achieve a high-precision, robust and interpretable method for remote sensing extraction of specific objects.

[0008] In a first aspect, embodiments of this application provide a method for remote sensing extraction of specific objects based on visual Transformer, including the following steps:

[0009] Step S1: Acquire the remote sensing images to be processed, including the main source remote sensing image and the auxiliary source time series image;

[0010] Step S2: Segment the main source remote sensing image to obtain sub-blocks, extract multi-source parameter information for each sub-block, and calculate the sub-block comprehensive quality score s based on the multi-source parameter information;

[0011] Step S3: Classify the sub-blocks according to the overall quality score s, dividing them into high-quality areas, medium-quality areas, and low-quality areas;

[0012] Step S4: Apply corresponding processing strategies to the sub-blocks based on the quality classification results;

[0013] Step S5: Feed the sub-blocks of the high-quality region and the medium-quality region into the trained Transformer model in sequence, and output the classification probability map and uncertainty map of the specific object.

[0014] Step S6: In the order of sliding windows, stitch together the classification results of all sub-blocks to restore the complete map, and at the same time generate the full map uncertainty layer and the area to be verified label layer.

[0015] In step S2, the multi-source parameter information includes cloud cover ratio, aerosol optical thickness, and humidity state. In step S4, if the quality classification result is a high-quality region, the latest main-source remote sensing image is used as the input to the Transformer model.

[0016] In step S4, if the quality classification result is a medium quality region, a multi-temporal input mode is adopted, and the latest temporal phase map, the lowest cloud / AOD / no precipitation image, and the latest historical auxiliary source temporal phase map are selected and stitched together as the input of the Transformer model.

[0017] The latest time-phase image is the main source remote sensing image obtained in step S1. The lowest cloud / AOD / no precipitation image is the auxiliary source time-series image with the lowest total cloud cover and aerosol concentration and no precipitation in the set historical time. The latest historical auxiliary source time-phase image is the one with the most recent shooting time among the historical auxiliary source time-series images of the corresponding area of ​​the sub-block.

[0018] The sum of cloud cover and aerosol concentration is calculated by recording the cloud occlusion ratio + AOD value of each image, and selecting the auxiliary source time series image with the lowest cloud occlusion ratio + AOD value and no rainfall.

[0019] If the quality classification result is a low-quality area, obtain the auxiliary source time-series image corresponding to the location of the low-quality area in the historical auxiliary source time-phase image set, treat it as the main source remote sensing image, and perform the operations described in steps S2-S3 to obtain its quality classification result; if it meets the standard of a high-quality area or a medium-quality area, process it according to the corresponding processing strategy. If the quality classification results of all auxiliary source time-series images in the historical auxiliary source time-phase image set fail to meet the standard, skip this area and mark it as "awaiting manual review".

[0020] Secondly, this application provides a specific object remote sensing extraction system based on visual Transformer, including: a main source image acquisition module: used to acquire main source remote sensing images;

[0021] Secondary source multi-temporal image acquisition module: used to acquire secondary source time-series images and store them in the system database;

[0022] Data processing module: used to execute steps S2-S6;

[0023] Database: Used for storing images, quality classification results during processing, and uncertainty maps;

[0024] Display module: Provides a WebGIS platform and mobile APP for displaying result maps and uncertain areas, and supports manual verification of input;

[0025] Closed-loop feedback and model update module: After receiving the results of manual review, it automatically adjusts the sub-block quality score weights and uncertainty thresholds.

[0026] Beneficial effects

[0027] In this invention, by evaluating and screening environmental parameters such as cloud cover, aerosol optical thickness (AOD), and rainfall, the system achieves quality control at the data input end, significantly reducing misjudgments caused by clouds, fog, and rainfall, improving the reliability of the model's input samples, and thus increasing the accuracy of the entire target extraction process. Secondly, this invention employs an adaptive block-based mechanism, setting different overlap ratios and multi-temporal input methods for high, medium, and low-quality regions based on their quality scores. High-quality regions achieve rapid and accurate processing, medium-quality regions enhance robustness through dense overlap and multi-temporal information, and low-quality regions are subject to manual review, effectively avoiding invalid calculations and noise interference, resulting in more efficient system operation and more reliable results. Furthermore, the multi-source temporal Transformer model, through the fusion of optical and temporal branches and combined with cross-attention, achieves a comprehensive understanding of spatial texture and temporal changes, not only enhancing the capture of abandoned land edges, morphology, and dynamic state changes, but also demonstrating strong robustness to environments with slight cloud cover or image interference. The uncertainty map and classification results output by the decoder can quantify the model confidence in detail, determine the confidence of pixels based on the entropy value, and perform subsequent re-inference or manual verification based on the uncertainty during the stitching stage, theoretically ensuring boundary coherence and overall accuracy. Attached Figure Description

[0028] Figure 1 This is a flowchart of a method for remote sensing extraction of specific objects based on the visual Transformer.

[0029] Figure 2Flowchart for implementing corresponding processing strategies for sub-blocks based on quality classification results;

[0030] Figure 3 This is a schematic diagram of the system in Example 2. Detailed Implementation

[0031] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be noted that these descriptions are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0032] Example 1

[0033] like Figure 1 As shown, in one embodiment, taking abandoned land as an example, the steps of the technical solution of the present invention are as follows:

[0034] Step S1: Obtain the remote sensing images to be processed, including the main source remote sensing image and the auxiliary source time series image.

[0035] In remote sensing target extraction, it is necessary to ensure high-quality input data. The data is collected and evaluated from two aspects: optical data quality and environmental impact measurement. The remote sensing images to be processed include main source remote sensing images and auxiliary source time-series images.

[0036] Main source remote sensing imagery: High-resolution satellite imagery (approximately 1 meter / pixel) is used to acquire clear textures and structural details of specific targets, preferably GF-2 imagery. This level of accuracy can be used to precisely identify the edges of abandoned land, the texture of weeds, and small-area land features, but it suffers from high acquisition costs and large storage requirements.

[0037] Secondary source temporal images: Preferably, Sentinel-2 imaging can be used to provide multi-temporal perspectives and enhance the ability to capture long-term change features. Its image size is smaller, it can be acquired frequently, and it is suitable for dynamic monitoring. By introducing secondary source temporal images and utilizing their high frequency of periodic imaging, occasional occlusion can be buffered, making the model more adaptable to sudden environmental changes (such as clouds and shadows) and improving output stability.

[0038] Step S2: Segment the main source remote sensing image to obtain sub-blocks, extract multi-source parameter information for each sub-block, and calculate the sub-block comprehensive quality score s based on the multi-source parameter information.

[0039] In remote sensing processing, large-area images are divided into several sub-blocks, typically 512×512 pixels, though no specific limit is specified here. Each sub-block represents a small square in the image, and the model analyzes, scores, and segments these sub-blocks separately. Multi-source parameter information is extracted from the main source remote sensing image, including cloud cover ratio, aerosol optical thickness, and humidity status.

[0040] Cloud occlusion ratio: refers to the proportion of pixels covered by clouds or cloud shadows in a remote sensing sub-block (such as a 512×512 pixel block) out of the total number of pixels.

[0041] Using the main source remote sensing image from step S1, which contains multiple spectral bands (such as red, green, blue, near-infrared, etc.), these bands are input into a lightweight neural network model specifically designed for cloud detection, such as CD-CTFM. It combines CNN (extracting local textures) and Transformer (learning full-image features). The model outputs the probability that each pixel is a "cloud" (the cloud probability value is between 0 and 1).

[0042] The higher the probability of a pixel being associated with a cloud, the greater the likelihood that it is occluded by a cloud or cloud shadow. Clouds can completely obscure ground features, making specific objects underneath, such as abandoned land textures, invisible. By identifying these occluded areas, these areas can be excluded, and missing information can be avoided as a basis for judgment, thus preventing incorrect identification and improving overall accuracy.

[0043] Aerosol Optical Depth (AOD): Aerosol optical depth (AOD) is a core atmospheric parameter used to measure the degree to which tiny particles in the atmosphere (such as dust, smoke, and haze) block and scatter sunlight. The higher the AOD value, the less light can pass through. NASA states that the meaning of AOD values ​​is as follows:

[0044] AOD < 0.1: The sky is very clear;

[0045] AOD≈1.0: The air is turbid and visibility is low;

[0046] AOD>3.0: Very severe smog, significantly blocking the sun.

[0047] In this embodiment, the AOD image is interpolated to the same grid (consistent with the size of the remote sensing sub-block), and the average AOD within each sub-block (e.g., 512×512 pixels) is calculated. Haze reduces image contrast and blurs land textures. Excluding haze areas helps the model more accurately identify land features and avoids misidentifying blurred texture areas as abandoned land.

[0048] Moisture Status Marker: Analyze meteorological data to obtain rainfall records for the past two days. For each sub-block, view the historical rainfall information for that area. If the cumulative rainfall in that area (sub-block) is ≥5 mm in the past two days, it is considered a "moisture-affected area".

[0049] If a sub-block is marked as a wet-affected area, the system will label that sub-block as "wet". Ground moisture alters light reflection; excluding or processing wet sub-blocks separately can prevent "non-abandoned" moist soil from being mistakenly identified as "abandoned land," thereby improving overall extraction accuracy and reducing interference.

[0050] Cloud obscuration rate, AOD, and humidity status are mapped proportionally to the 0-1 range, and the overall quality score s is calculated using the following formula:

[0051] s=1-(w1×cloud+w2×AOD+w3×rain)

[0052] The specific parameters are explained below:

[0053] cloud: cloud cover percentage (e.g., 0.3 means 30% of the area is covered by clouds);

[0054] AOD: Average haze value of the sub-block (e.g., 0.25);

[0055] rain: Wet status indicator (0 / 1).

[0056] w1, w2, and w3 are weights, and preferably they can be set to w1>w2>w3 according to the influence of each parameter on the result, and their sum is set to 1.

[0057] The closer s is to 1, the less environmental interference the sub-block has and the higher the observation quality, making it suitable for direct processing. The closer s is to 0, the more environmental factors there are and the lower the block quality is, requiring optimization or manual review.

[0058] Step S3: Classify the sub-blocks according to the overall quality score s, dividing them into high-quality areas, medium-quality areas, and low-quality areas.

[0059] The initial overlap between sub-blocks is set to 20% (approximately 102 pixels); this overlap ensures that the model can see the context of adjacent areas at the boundaries, improving the continuity and accuracy of stitching.

[0060] Based on the overall quality score s, the sub-blocks are divided into three categories:

[0061] High-quality zone: s≥0.8—clear information, high quality;

[0062] Medium quality region: 0.6≤s<0.8 — There is some occlusion or interference, but optimization is still possible;

[0063] Low-quality region: s<0.6—Severe interference, not recommended for direct use, triggers compensation process.

[0064] Step S4: Apply corresponding processing strategies to the sub-blocks based on the quality classification results.

[0065] Different processing strategies apply to different quality classification results for sub-blocks. (See attached document) Figure 2 The processing flow based on the quality classification results is shown.

[0066] For high-quality areas, which indicate high imaging quality and low interference, the main source remote sensing image is directly used for subsequent processing.

[0067] Although the medium-quality region contains minor disturbances such as clouds, haze, or rainfall, it still contains usable information. Skipping it directly might overlook important data, therefore additional processing is needed to compensate for the deficiencies. The specific operational procedure is as follows:

[0068] The overlap ratio has been increased from 20% to 50%, meaning that adjacent sub-blocks will share approximately 256 pixels.

[0069] Multi-temporal input: Select the latest temporal phase image + the lowest cloud / AOD / no precipitation image + the latest historical auxiliary source temporal phase image, and stitch them together as input.

[0070] Latest temporal image: refers to the main source remote sensing image obtained in step S1, which provides the latest land cover status and texture information, and is the most critical ground feature image.

[0071] Lowest Cloud / AOD / No Rainfall Image: This refers to the auxiliary source time series image with the lowest cloud cover, lowest aerosol concentration, and no rainfall within a historical period. It is the clearest image among historical auxiliary source time series images for that area, ensuring the model's recognition accuracy. The historical period can be set according to the database storage capacity, such as 7 days or 1 month. The specific selection strategy is to use a cloud detection model to evaluate the cloud cover ratio and calculate the AOD value. Record the cloud cover ratio + AOD value for each image, and select the image with the lowest cloud cover ratio + AOD score that has not experienced rainfall. If the image with the lowest cloud cover ratio + AOD score has experienced rainfall, then select the image with the second lowest cloud cover ratio + AOD score that has not experienced rainfall.

[0072] Latest historical auxiliary source time-series image: The most recently captured image among the historical auxiliary source time-series images of this area. Under normal circumstances, it is updated every 3-7 days and can serve as a backup to compensate for occasional interference.

[0073] The above steps can buffer situations where the latest image is suddenly obscured by clouds or has poor image quality, while also providing an alternative solution to improve the stability of the model's judgment.

[0074] Input merging: Multi-temporal images are stitched together into a 512×512×9 (three-temporal × 3-channel RGB) structure. The multi-temporal information improves the tolerance to cloud cover and haze, and the overlapping areas add context to assist in the localization of stitching edges.

[0075] For low-quality areas, a compensation process is triggered. Based on the time series, the nearest-time-series image of the low-quality area from the historical auxiliary source time-phase image set is acquired and treated as the main source remote sensing image. Steps S2-S3 are then performed on this image to obtain its quality classification result. If it meets the criteria for a high-quality or medium-quality area, the corresponding subsequent operations are performed on this image. If it does not meet the criteria, earlier auxiliary source time-series images are acquired until the classification result meets the criteria for a high-quality or medium-quality area. If the classification results of all auxiliary source time-series images fail to meet the criteria, this area is skipped and marked as "awaiting manual review".

[0076] Simultaneously, the system marks this sub-block as "awaiting manual review" and displays it as a warning area on WebGIS or mobile devices. The interface allows users to view multi-temporal maps and SAR auxiliary maps, and manually determine its abandoned land attributes to ensure the accuracy of classification to the greatest extent possible. Manual input results are encoded into the system and stored as trusted tags. Table 1 below shows the processing methods for different classification results.

[0077] Table 1. Processing methods for different classification results

[0078] Step S5: Feed the sub-blocks of the high-quality region and the medium-quality region into the trained Transformer model in sequence, and output the classification probability map and uncertainty map of the specific object.

[0079] This embodiment adopts a CNN+Transformer hybrid structure: First, CNN is used to extract local texture features and process the main source remote sensing image and the auxiliary source time series image respectively. Then, Transformer (including ViT or SwinTransformer variants) is used to further fuse global information and use a self-attention mechanism to obtain global context information, thereby forming a comprehensive understanding of the boundaries, structure and texture of ground features.

[0080] During the processing of medium-quality sub-blocks, three images are concatenated to form a multi-temporal input; this ensures that the model simultaneously observes the current state, recent changes, and clear references, improving robustness to slight occlusion or illumination changes. In the decoding stage, the fused high-dimensional features are decoded and restored to a 512×512 resolution spatial feature map, achieving dual-channel output of the classification probability map and the uncertainty map.

[0081] Classification probability map: each pixel corresponds to the probability of "abandoned land" or "non-abandoned land", preferably, if the probability P of a pixel is ≥ 0.5, it is determined as abandoned land.

[0082] Uncertainty map: MC-Dropout or MC-Frequency Dropout is used to calculate pixel-level entropy, and estimate the confidence level of the model's judgment. The higher the entropy, the less confident the model is in the classification of this pixel.

[0083] Step S6: splicing and restoring all sub-blocks' classification results into a complete map according to the sliding window sequence, and generating a full-map uncertainty layer and a region-to-be-rechecked marking layer at the same time.

[0084] For each overlapping sub-block region, the system adopts different strategies according to the uncertainty u of pixels in the region, sets a low uncertainty threshold t1 and a high uncertainty threshold t2, preferably, the value of threshold t1 is 0.2, and the value of the high uncertainty threshold is 0.5.

[0085] If u<t1 (low uncertainty): it means the model output is credible, and the overlapping region of sub-blocks is directly fused by averaging;

[0086] If t1≤u≤t2 (medium uncertainty): trigger secondary inference, and select the result with higher confidence for splicing;

[0087] If u>t2 (high uncertainty): the model is uncertain, mark this region as to be rechecked, and it will not participate in automatic splicing.

[0088] splicing and restoring all sub-blocks' classification results into a complete map according to the sliding window sequence, and generating a full-map uncertainty layer and a region-to-be-rechecked marking layer at the same time. The uncertainty layer displays the confidence level of the model for each pixel classification in the form of a heat map, and the to-be-rechecked marking layer clearly marks the high uncertainty region for manual rechecking or subsequent supplementary sampling, so as to assist subsequent processing.

[0089] For the "to be rechecked" region, the system can automatically perform resampling or prompt manual rechecking. After the rechecking result is fed back to the system, it is used to adjust the uncertainty threshold and update the model. All inference results (classification, position, uncertainty, rechecking mark) are stored in a spatial database (such as PostGIS), which facilitates spatiotemporal query and management. A visualization platform is established to display classification maps, difficult regions and previous rechecking history, and supports on-site positioning, manual verification and data annotation.

[0090] Example 2

[0091] Figure 3 is a schematic diagram of the system provided by the second embodiment of the present application, specifically including:

[0092] Main source image acquisition module: Used to acquire high-resolution optical images to obtain the latest texture and structural information of the target area. Preferably, GF-2 imagery is used.

[0093] Auxiliary source multi-temporal image acquisition module: used to acquire multi-temporal remote sensing images that are easy to update and dynamically monitor, and store them in the system database. Preferably, Sentinel-2 platform images are used.

[0094] Data processing module: Used to execute the processing steps S2-S6 above.

[0095] Database module: Stores the acquired remote sensing images, all classification results, uncertainty values, and verification labels into the spatial database.

[0096] Display module: Provides interfaces between the WebGIS platform and mobile APP to display result maps and uncertain areas, and supports manual verification of input.

[0097] Closed-loop feedback and model update module: After receiving the results of manual review, it automatically adjusts the sub-block quality score weights and uncertainty thresholds; it periodically adds the reviewed and labeled data to the training set, triggers model fine-tuning, and improves future recognition accuracy.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0099] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0100] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.

Claims

1. A method for remote sensing extraction of specific objects based on visual Transformer, characterized in that, Includes the following steps: Step S1: Acquire the remote sensing images to be processed, including the main source remote sensing image and the auxiliary source time series image; Step S2: Segment the main source remote sensing image to obtain sub-blocks, extract multi-source parameter information for each sub-block, and calculate the sub-block comprehensive quality score s based on the multi-source parameter information; Step S3: Classify the sub-blocks according to the overall quality score s, dividing them into high-quality areas, medium-quality areas, and low-quality areas; Step S4: Take appropriate processing strategies for the sub-blocks based on the quality classification results; If the quality classification result is a high quality region, the main source remote sensing image in step S1 is used as the input of the Transformer model. If the quality classification result is a medium quality region, a multi-temporal input mode is adopted, and the latest temporal phase map, the lowest cloud / AOD / no precipitation image, and the latest historical auxiliary source temporal phase map are selected and stitched together as the input of the Transformer model. The latest time-phase image is the main source remote sensing image obtained in step S1. The lowest cloud / AOD / no precipitation image is the auxiliary source time-series image with the lowest total cloud cover and aerosol concentration and no precipitation in the set historical time. The latest historical auxiliary source time-phase image is the one with the most recent shooting time among the historical auxiliary source time-series images of the corresponding area of ​​the sub-block. If the quality classification result is a low-quality area, obtain the auxiliary source time series image corresponding to the location of the low-quality area with the most recent acquisition time in the historical auxiliary source time phase image set, regard it as the main source remote sensing image, and perform the operations described in steps S2-S3 to obtain its quality classification result; if it meets the standard of high-quality area or medium-quality area, then process it according to the corresponding processing strategy. If the quality classification results of all auxiliary source time series images in the historical auxiliary source time phase image set fail to meet the standards, skip this area and mark it as "awaiting manual review". Step S5: Feed the sub-blocks of the high-quality region and the medium-quality region into the trained Transformer model in sequence, and output the classification probability map and uncertainty map of the specific object. Step S6: In the order of sliding windows, stitch together the classification results of all sub-blocks to restore the complete map, and at the same time generate the full map uncertainty layer and the area to be verified label layer.

2. The method according to claim 1, characterized in that, In step S2, the multi-source parameter information includes cloud obstruction ratio, aerosol optical thickness, and humidification status.

3. The method according to claim 1, characterized in that, The sum of cloud cover and aerosol concentration is calculated by recording the cloud occlusion ratio + AOD value of each image, and selecting the auxiliary source time series image with the lowest cloud occlusion ratio + AOD value and no rainfall.

4. A remote sensing extraction system for specific objects based on visual Transformer, characterized in that, include: Main source image acquisition module: used to acquire main source remote sensing images; Secondary source multi-temporal image acquisition module: used to acquire secondary source time-series images and store them in the system database; Data processing module: used to execute steps S2 to S6 as described in any one of claims 1-3; Database module: Used for storing images, quality classification results during processing, and uncertainty maps; Display module: Provides a WebGIS platform and mobile APP for displaying result maps and uncertain areas, and supports manual verification of input; Closed-loop feedback and model update module: After receiving the results of manual review, it automatically adjusts the sub-block quality score weights and uncertainty thresholds.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic equipment

    CN119068345A

  • High-precision spectral remote sensing atmospheric correction and cloud removal processing method fusing multiple time phases, atmospheric parameters and ground reflectivity

    CN119540110A