Remote sensing image general basic model construction and analysis method based on grid coding

Through the general basic model of remote sensing images based on grid encoding, the problems of high manual labeling cost, difficult data integration and low computing efficiency in remote sensing image analysis are solved, efficient integration of multi-source data and cross-task generalization are achieved, classification accuracy and segmentation performance are improved, and it is suitable for diversified downstream tasks such as smart cities and agricultural monitoring.

CN120495899APending Publication Date: 2025-08-15NAT SUPERCOMPUTING SHENZHEN CENT (SHENZHEN CLOUD COMPUTING CENT) +1
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510917991.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional remote sensing image analysis methods have problems such as high manual annotation cost, difficulty in integrating multi-source data, low processing efficiency of massive data, and inconsistent data identification and coding, making it difficult to achieve the ability to generalize across tasks, fields, and modalities.

Method used

A general basic model of remote sensing images based on grid encoding is adopted. Through multimodal data preprocessing, integrated space-time grid coding, multi-scale feature extraction and cross-task adaptation, a general basic model of remote sensing images based on grid encoding is built, and multi-source remote sensing data is unified to achieve efficient integration and cross-scale generalization.

Benefits of technology

It improves the efficiency of multi-source data integration, enhances the generalization ability of the model, has classification accuracy of more than 90%, has more than 80% segmented IoU, has performed excellently in high-score series data tests, and has improved the accuracy rate in global mapping tasks by 7%, and has significantly reduced the dependence of manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495899A_ABST
    Figure CN120495899A_ABST
Patent Text Reader

Abstract

The invention discloses a grid coding-based remote sensing image general basic model construction and analysis method, which comprises the following steps of: pre-processing multi-modal data, eliminating cloud layer influence, aligning space coordinates and unifying data scales; generating a unified identification code with geographical meaning by using efficient space-time grid integrated coding; performing multi-scale feature extraction by means of a visual Transform and a convolutional neural network, and optimizing multi-source data fusion in combination with an anchoring feature extraction strategy; and finally, cross-task downstream application adaptation is realized on the basis of a multi-task Transform architecture. According to the method, the problems of high manual annotation cost, difficulty in multi-source data integration, low mass data processing efficiency and the like in traditional remote sensing image analysis are effectively solved. Experiments show that the multi-source data integration efficiency is improved, the model generalization ability is enhanced, the classification precision in high-resolution series data tests reaches 90% or above, the segmentation IoU exceeds 80%, and an efficient solution is provided for remote sensing image analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to remote sensing image analysis technology, and in particular to a remote sensing image universal basic model construction and analysis method based on grid coding. Background Art

[0002] Remote sensing technology, as a crucial means of acquiring information about the Earth's surface, plays an irreplaceable role in numerous fields, including resource exploration, environmental monitoring, disaster warning, and urban planning. With the rapid growth of the number of remote sensing satellites in orbit worldwide, now exceeding 1,200, the amount of remote sensing data generated annually reaches tens of petabytes, characterized by massive volume, multi-source heterogeneity, and multidimensionality.

[0003] Traditional remote sensing image analysis methods face significant challenges, mainly in the following aspects: Manual labeling is costly: Remote sensing data labeling relies on background knowledge and is more expensive. For example, in high-resolution remote sensing images, multiple objects are distributed in adjacent spaces with complex and varied scales and shapes, and have large scale differences. At the same time, due to the diversity of land objects, remote sensing images have large intra-class differences and high inter-class similarities, which poses a challenge to multi-label remote sensing image classification. Identifying and labeling target labels requires greater manpower, and fine-grained distinction between different land objects requires more precise expertise and double the workload. This makes manual labeling time-consuming and labor-intensive, and it is difficult to ensure data consistency on a global scale.

[0004] Difficulty integrating multi-source data: Traditional methods rely on single-modality, task-specific models, making it difficult to effectively integrate multimodal data (such as optical, radar, and hyperspectral data), time series, and geographic prior knowledge. This limits the model's generalization across multiple tasks and large datasets. The heterogeneity of multi-source data further exacerbates the fusion challenge, making it difficult to uniformly process data of varying resolutions, formats, and quality. For example, geometric differences between optical and SAR data, overlapping spectral ranges of different sensors, and insufficient temporal synchronization all increase the complexity of data processing.

[0005] Inefficient processing of massive data volumes: Processing terabytes or even petabytes of data requires enormous computing resources, making traditional single-machine processing difficult to meet real-time requirements, such as the minute-level response required for disaster monitoring. Remote sensing geoscience computing is both computationally and data-intensive, and data-intensive computing leads to I / O and memory-intensive computing. Therefore, personal computers are no longer sufficient for current large-scale remote sensing applications. The mismatch between hard drive read / write, network transmission, memory access, and computer processing speed creates a computing resource bottleneck that limits the overall system's computational efficiency.

[0006] Natural factors such as cloud cover reduce data validity: the global average cloud cover exceeds 66%, reaching over 80% in tropical regions, severely impacting the availability of optical remote sensing data. Cloud obstruction can be divided into thin cloud interference and thick cloud cover. Thin clouds can be considered a mixture of ground objects and clouds, while thick clouds completely block surface information, resulting in a complete loss of the imaging area. Accurate detection and removal of clouds and fog are key to improving the utilization of remote sensing image data. However, the physical and imaging properties of clouds are complex and diverse, with rich textures and shapes. Methods based on single or multiple spatial features are often only effective for specific types of remote sensing data, while their performance for other sensor data varies greatly.

[0007] In addition, there is still no unified identification standard for remote sensing data, which creates a bottleneck for the comprehensive application of massive remote sensing data from multiple sources.

[0008] Inconsistent identification codes across management systems: Different remote sensing data management systems use independent identification systems for the same remote sensing data. For example, the "Geospatial Data Cloud," the "Gaofen Science and Education Platform," and the "Land Observation Satellite Data Service Platform" have different naming conventions for Gaofen satellite data. This lack of direct correlation in identification makes cross-database management of remote sensing data difficult. When users try to find the same image across different systems, they can only rely on information such as latitude and longitude ranges to search and filter, greatly increasing query complexity.

[0009] The identification codes of different remote sensing data within the same system are not unified: Even within the same remote sensing data management system, there are significant differences in the identification codes of different remote sensing data. For example, the "Geospatial Data Cloud" system identifies different satellite data such as Landsat5, EO-1, Sentinel-1, ZY-3, and GF-1. These data identifiers usually only contain information such as the satellite name, sensor type, shooting time, and product serial number, and lack actual geographical meaning. This means that the spatial coverage of remote sensing data cannot be seen from the identification code itself, and additional metadata queries are required, affecting query efficiency. In addition, this inconsistency also leads to a lack of direct association between multi-phase, multi-scale, and multi-source remote sensing data in the same area, making data integration and query inconvenient.

[0010] At the same time, remote sensing data is typically owned and stored by various space agencies worldwide. This data exhibits significant heterogeneity in terms of format, management methods, access methods, and data usage specifications. Given the enormous volume of remote sensing data and the significant time overhead associated with transmitting it over wide area networks, organizing and serving this heterogeneous data from diverse space agencies is an extremely complex task. Ideally, remote sensing data should cover all time, location, and observation characteristics globally. However, this is unrealistic and unnecessary in practice. From a data spatial perspective, data from various sensors and remote sensing data analysis models are essentially fragments within a high-dimensional remote sensing data space. An ideal remote sensing data infrastructure should be able to reconstruct all data from these fragments using data processing techniques such as interpolation or inversion to compensate for missing data. However, this reconstruction is limited by algorithmic deficiencies, uncertainty in accuracy, and the enormous computational complexity. The large volume of remote sensing data and the diverse and complex processing methods create an urgent need for high performance, high throughput, and low cost. Applications such as natural disaster monitoring and early warning require rapid response to disasters, demanding real-time performance on the order of minutes or even seconds. Global change research requires processing massive amounts of remote sensing data. The effectiveness of these applications depends not only on peak computing speed but also on the scale of the problems that can be solved. Furthermore, the emergence of the next-generation Digital Earth is forcing remote sensing to retain traditional professional applications while providing intuitive Earth information services to the public. This necessitates high concurrency in remote sensing computing. The increasing volume of data can easily lead to algorithmic performance bottlenecks. Traditional centralized data storage and management strategies and serial spatiotemporal analysis algorithms are increasingly unable to meet the demands for efficient storage and real-time processing and analysis of spatiotemporal big data.

[0011] In recent years, intelligent remote sensing interpretation technology has rapidly developed. However, most of these models are specialized and difficult to generalize to different tasks, resulting in a waste of resources. Foundational models, as a universal and generalizable solution, have attracted significant attention in the remote sensing field. These foundational models, integrated from multiple expert models, acquire extensive knowledge by learning from a large amount of data and tasks, capturing more details, enabling them to solve a variety of downstream tasks and generalizing better to new datasets. Remote sensing data exhibits complex and heterogeneous characteristics such as multimodality, multitemporality, and multiscale. Existing masking strategies primarily focus on modeling spatial features while neglecting spectral feature modeling, resulting in an inability to fully exploit the spectral dimension of spectral data. Data acquired by different sensors exhibit geometric, spectral, and temporal differences, such as geometric differences between optical and SAR data, overlapping spectral ranges between different sensors, and insufficient temporal synchronization, all of which increase the complexity of data processing. Although a large body of work has achieved remarkable results in perceptual recognition and cognitive prediction tasks using single- or multi-temporal remote sensing data, a comprehensive review is still lacking to provide systematic guidance for foundational remote sensing models. Models must generalize across tasks, domains, and modalities to meet the broad needs of remote sensing applications. Basic remote sensing models typically have massive parameter counts. For example, the SkySense model has 2.06 billion parameters, requiring 80 A100 GPUs for training. This results in high model training costs and enormous demands on computing resources. Traditional deep learning models have a large number of parameters, severely limiting their applicability and interpretability. Furthermore, convolutional neural network (CNN) training suffers from gradient dispersion, low training efficiency, and optimization difficulties. High-quality data remains the primary challenge for large models. Remote sensing data annotation is costly, and achieving the full diversity required by the models is difficult.

[0012] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0013] The main purpose of the present invention is to overcome the defects in the above-mentioned background technology and provide a method for constructing and analyzing a universal basic model of remote sensing images based on grid coding.

[0014] To achieve the above object, the present invention adopts the following technical solutions: A method for constructing and analyzing a universal basic model of remote sensing images based on grid coding includes the following steps: S1. Multimodal Data Preprocessing: Temporal correction is performed on heterogeneous remote sensing data to eliminate cloud effects. Geometric correction is performed using multi-scale feature fusion and template matching to align spatial coordinates and project them into a unified coordinate system. Dynamic slicing technology is used to unify the data scale. After noise processing, standardized image data is generated. S2. Efficient spatiotemporal grid integration encoding: Dynamically determines the quadtree grid level based on the resolution threshold, combines the geographic embedding module to encode the latitude and longitude of the four corners of the image and the ground sampling distance, and integrates the time code with metadata code to generate a unified identification code and grid code feature vector with geographical meaning; S3. Multi-scale feature extraction: We use the multi-head self-attention mechanism of the visual Transformer to capture global dependencies, combine it with a convolutional neural network to extract local texture features, and use an anchored feature extraction strategy to dynamically generate masks to optimize multi-source data fusion. S4. Cross-task downstream application adaptation: Based on the multi-task Transformer architecture, feature transfer learning technology is used to adapt features to image classification, semantic segmentation, and change detection tasks.

[0015] Furthermore, in step S1: The geometric correction is aimed at complex terrain or low-quality images, and uses multi-scale feature fusion and global geometric constraints to enhance matching robustness; Perform multi-look processing, speckle filtering, and DEM terrain correction on SAR images; The unified coordinate system projection adopts the UTM partition mechanism, and cross-partition images are based on the coverage area center or the user-specified coordinate system.

[0016] Furthermore, in step S2, the unified identification code is composed of metadata code, spatial grid code and unit code, wherein: The spatial grid code includes the positioning grid code and the span code that records the grid span in the longitude and latitude directions; The positioning grid code is mapped to a quaternary one-dimensional code at a preset level according to the image resolution.

[0017] Furthermore, in step S2: The grid coding uses fine-grained grids for urban areas and coarse-grained grids for sparse areas, thereby optimizing storage and retrieval efficiency.

[0018] Furthermore, in step S2: The resolution threshold is divided into: high resolution is mapped to a fine grid level, medium resolution is mapped to a medium grid level, and low resolution is mapped to a coarse grid level; The boundary resolution uses fuzzy logic to calculate the membership weight to determine the grid level, or is adaptively adjusted through multi-scale feature fusion.

[0019] Furthermore, in step S2: According to the mapping rules between remote sensing image resolution and grid level, a dynamic hierarchical division model is established through logarithmic calculation. For local high-contrast areas in the image, such as areas with complex textures or significant boundary features of land objects, feature evaluation is performed based on the gradient amplitude or information entropy index in the area, and the degree of grid division is dynamically adjusted to enhance the ability to retain spatial detail information.

[0020] Furthermore, the anchor feature extraction strategy in step S3 includes: The data were classified into the following categories according to the collection time and source: same time but different sources, same source but different time and other situations; Generate the same mask matrix for data from different sources at the same time and implement consistent masking; Generate complementary mask matrices for data from the same source at different times to implement mutually exclusive masking; Generate random masks independently for the rest of the data; Three spectral bands are randomly selected to train the band-independent feature encoder.

[0021] Furthermore, in step S3: In the multi-scale feature extraction step, the multi-head self-attention mechanism of the visual Transformer is used to construct an associative calculation model of the query, key, and value matrices to capture the global semantic dependencies of the remote sensing imagery. At the same time, combined with a convolutional neural network, multi-scale convolution kernels or dilated convolution operations are used to extract local texture features of different scales in the image, forming a complementary fusion of global semantics and local texture features.

[0022] Furthermore, in step S3: The mask ratio is dynamically controlled at 75%-90%, and the mask generation adopts a random block mask or a grid mask algorithm.

[0023] Furthermore, in step S4: The cross-task adaptation adopts feature distillation technology to fuse ViT and CNN output features, and shares encoder weights to deploy task-specific decoders.

[0024] The present invention has the following beneficial effects: This paper proposes a method for constructing and analyzing a universal basic model for remote sensing imagery based on grid coding. This method constructs a universal basic model for remote sensing imagery based on grid coding, and through the collaboration of grid coding and a multi-task architecture, achieves efficient integration and cross-scale generalization of multi-source data. This method effectively overcomes the limitations of traditional methods in terms of labeling cost, data heterogeneity, and computational efficiency, and enhances the automated processing and cross-task generalization capabilities of multi-source heterogeneous remote sensing data. On the one hand, by leveraging grid coding technology to uniformly organize multi-source remote sensing data and incorporate images of different resolutions, time series, and spectral bands into a unified framework, this method not only addresses the issues of inconsistent remote sensing data identification coding, lack of practical geographic meaning, and difficulty in integrating data from the same region, but also improves the efficiency of multi-source data integration through a quadtree structure and geographic embedding module. Experimental results show that the present invention can encode and index 1TB of data in less than one hour, with an encoding consistency error of less than 5%. On the other hand, by combining modules such as multimodal data preprocessing, multi-scale feature extraction and cross-task adaptation, we build a spatiotemporal and spatial-spectral integrated representation capability, capture global and local features through the fusion of visual Transformer and convolutional neural network, and optimize multi-source data fusion using anchor feature extraction strategy. The model has achieved classification accuracy exceeding 90% and segmentation IoU exceeding 80% in high-resolution series data tests, and cross-scale feature similarity exceeding 85%. Its performance is better than existing models such as SatMAE, and the accuracy rate in global land cover mapping tasks has increased by 7%. It only takes 4 minutes to complete global mapping on a supercomputer, which significantly improves efficiency.

[0025] In addition, the model significantly reduces reliance on manual labeling and reduces data preparation costs through weak supervision and self-supervised learning. At the same time, with its multi-task Transformer architecture and transfer learning technology, it enhances the generalization ability for diverse downstream tasks such as smart cities, agricultural monitoring, and surface cover mapping. It can better cope with challenges such as the heterogeneity, temporal synchronization, and spectral differences of multimodal data, providing an efficient solution for large-scale remote sensing applications.

[0026] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematic diagram of a spatiotemporal spectrum image according to an embodiment of the present invention.

[0028] Figure 2 This is a flow chart of a data preprocessing method according to an embodiment of the present invention.

[0029] Figure 3 2 is a diagram showing the effect of data pruning according to an embodiment of the present invention, where (a) is the discarded image and (b) is the retained image.

[0030] Figure 4This is a diagram of the basic model algorithm architecture of an embodiment of the present invention.

[0031] Figure 5 This is a flow chart of a geographic information encoding method according to an embodiment of the present invention.

[0032] Figure 6 This is a schematic diagram of a downstream scenario application of an embodiment of the present invention.

[0033] Figure 7 This is an overall flow chart of the method for constructing and analyzing a universal basic model for remote sensing images based on grid coding in the present invention. DETAILED DESCRIPTION

[0034] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.

[0035] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0036] The present invention aims to overcome the limitations of traditional remote sensing image analysis methods in terms of labeling costs, data heterogeneity, and computational efficiency, and to solve the problem of the lack of a unified identification standard for remote sensing data. It provides a method for constructing and analyzing a universal basic model for remote sensing images based on grid coding. The present invention designs a universal basic model device and key algorithms for remote sensing images based on grid coding, combining multimodal data processing, grid coding technology, and deep learning architecture to achieve efficient remote sensing image analysis and overcome the limitations of traditional methods in terms of labeling costs, data heterogeneity, and computational efficiency. By constructing a representation capability that integrates time, space, and spectrum, remote sensing images of different resolutions, time series, and spectral bands are incorporated into a unified framework to improve the generalization and adaptability of the model. Different from traditional single-task models, the present invention uses grid coding to uniformly organize multi-source remote sensing data and construct a universal model that supports multiple downstream tasks.

[0037] See Figure 7 The embodiment of the present invention provides a method for constructing and analyzing a universal basic model of remote sensing images based on grid coding, comprising the following steps: Step S1. Multimodal data preprocessing: Temporal correction is performed on heterogeneous remote sensing data to eliminate cloud effects. Geometric correction assisted by multi-scale feature fusion and template matching is used to align spatial coordinates and project them into a unified coordinate system. Dynamic slicing technology is used to unify the data scale. After noise processing, standardized image data is generated.

[0038] In some embodiments, in step S1: for complex terrain or low-quality images, the geometric correction adopts multi-scale feature fusion and global geometric constraints to enhance the robustness of matching; performs multi-look processing, speckle filtering, and DEM terrain correction on SAR images; the unified coordinate system projection adopts the UTM zoning mechanism, and the cross-zone images are based on the center of the coverage area or the user-specified coordinate system. When dynamically slicing, the block size can be dynamically adjusted according to the resolution.

[0039] Step S2. Efficient spatio-temporal grid integrated coding: Dynamically determine the quadtree grid level based on the resolution threshold, combine the geographic embedding module to encode the longitude and latitude of the four corners of the image and the ground sampling distance, fuse the time coding and metadata coding, and generate a unified identification code with geographical meaning and a grid coding feature vector.

[0040] In some embodiments, the unified identification code consists of metadata coding, spatial grid coding, and unit code, where: the spatial grid coding includes positioning grid coding and span code for recording the longitude and latitude grid spans; the positioning grid coding is mapped to a quaternary one-dimensional coding of a preset level according to the image resolution.

[0041] In some embodiments, for the grid coding, fine-grained grids are used for urban areas and coarse-grained grids are used for sparse areas to optimize the storage and retrieval efficiency.

[0042] In some embodiments, the resolution thresholds are divided as follows: high resolution (GSD ≤ 5 meters) is mapped to the fine grid level, medium resolution (5 meters < GSD ≤ 30 meters) is mapped to the medium grid level, and low resolution (GSD > 30 meters) is mapped to the coarse grid level; the boundary resolution uses fuzzy logic to calculate the membership degree to weightedly determine the grid level, or adaptively adjusts through multi-scale feature fusion.

[0043] In some embodiments, in step S2, according to the mapping rule between the remote sensing image resolution and the grid level, a dynamic level division model is established through logarithmic calculation; for local high-contrast regions in the image, such as regions with complex textures or significant feature boundaries of ground objects, feature evaluation is performed based on the gradient magnitude or information entropy index within the region, and the fineness of the grid division is dynamically adjusted to enhance the ability to retain spatial detail information.

[0044] Step S3. Multi-scale feature extraction: Capture global dependencies through the multi-head self-attention mechanism of the Vision Transformer, combine the convolutional neural network to extract local texture features, and adopt an anchor feature extraction strategy to dynamically generate masks to optimize the multi-source data fusion; In some embodiments, the anchor feature extraction strategy includes: classifying data according to acquisition time and source into: same time but different sources, same source but different time and other situations; generating the same mask matrix for data from different sources at the same time to implement consistency masking; generating complementary mask matrices for data from the same source but different time to implement mutual exclusion masking; independently generating random masks for the remaining data; and randomly selecting three spectral bands to train band-independent feature encoders.

[0045] In some embodiments, the mask ratio is dynamically controlled at 75%-90%, and the mask generation adopts a random block mask or a grid mask algorithm.

[0046] In some embodiments, an association calculation model of the query Q, key K, and value V matrices is constructed through the multi-head self-attention mechanism of the visual Transformer to capture the global semantic dependencies of the remote sensing image; at the same time, combined with a convolutional neural network, multi-scale convolution kernels or dilated convolution operations are used to extract local texture features of different scales in the image, forming a complementary fusion of global semantics and local texture features.

[0047] The multi-head self-attention mechanism is as follows: Calculate feature dependencies, is the dimension of the key vector, Q, K, and V are the query, key, and value matrices respectively.

[0048] Step S4. Cross-task downstream application adaptation: Based on the multi-task Transformer architecture, features are adapted to image classification, semantic segmentation, and change detection tasks through transfer learning technology.

[0049] In some embodiments, the cross-task adaptation uses feature distillation technology to fuse ViT and CNN output features, and shares encoder weights to deploy task-specific decoders.

[0050] The following further describes specific embodiments of the present invention and experimental verification.

[0051] In response to the problems existing in the existing remote sensing data identification system, namely, the identification coding is not unified, the coding has no actual geological meaning, and the coding is difficult to integrate data from the same area, the present invention proposes a method for constructing and analyzing a universal basic model of remote sensing images based on grid coding, designs a universal basic model device and key algorithms for remote sensing images based on grid coding, and combines multimodal data processing, grid coding technology and deep learning architecture to achieve efficient remote sensing image analysis. The present invention uses grid coding to uniformly organize multi-source remote sensing data and construct a universal model that supports multiple downstream tasks, which is different from the traditional single-task model. The technical points of the present invention include multimodal data preprocessing, efficient spatiotemporal grid integrated coding, multi-scale feature extraction and integration of cross-task adaptation modules.

[0052] The design principles of this invention focus on resolving three core issues with existing remote sensing data identification systems: First, assigning a unified identification code to remote sensing data of varying sources, types, and resolutions, thereby addressing the data sharing and management challenges faced by various departments due to inconsistent data identification. Second, the code is embedded with rich spatial information, directly reflecting the geographic location and spatial extent of the data, avoiding reliance on additional metadata to query geographic information and thus improving query efficiency. Third, the code design facilitates the rapid identification and integration of remote sensing data from different sensors and at different temporal phases within the same geographic region, supporting the comprehensive analysis and application of multi-temporal, multi-scale, and multi-source remote sensing data. Overall, this invention designs a coding model suitable for the unified identification of remote sensing data without overturning or reinventing the existing remote sensing data organization system. By establishing a mapping relationship between remote sensing data and its corresponding Earth subdivision facets, using a unified identification code as a link, and while preserving the original data structure, this method achieves the unified organization, efficient management, and application of remote sensing data, fundamentally resolving the challenge of unified and efficient organization and management of multi-source remote sensing imagery.

[0053] Unified identification and coding model for multi-source remote sensing data Unified identification code for multi-source remote sensing data = metadata code + spatial grid code + unit code.

[0054] The model consists of three parts: metadata code (MetadataCode), spatial grid code (SpatialGridCode) and unit code (OrganizationCode).

[0055] Metadata encoding: This refers to the identification code used in remote sensing image metadata files. It bridges the gap between traditional encoding methods and the unified identification code used for multi-source remote sensing data. It reflects the continuity of the existing organizational structure for remote sensing data and adheres to the overall design principle of "not overthrowing or reinventing the existing organizational structure." Metadata encoding follows a unified naming convention and can include information such as satellite name, sensor type, longitude, latitude, acquisition time, product level, product number, orbit number, scene number, resolution, receiving station, and supplementary information. The presence and order of these components can be adjusted based on specific circumstances.

[0056] Spatial grid coding: This uses a spatial grid to represent the geographic location and spatial extent of remote sensing data. It is the only component of the unified identification and coding model for multi-source remote sensing data that has actual geographic meaning. It consists of two parts: the location grid code (Location Code) and the span code (Span Code), which are connected by a hyphen (-).

[0057] Positioning grid code: This is a quaternary one-dimensional code for the specific level of the positioning grid. The selection of the positioning grid level comprehensively considers the spatial resolution of the remote sensing image, classifying remote sensing images into three categories: high spatial resolution (<1m), medium spatial resolution (1m-500m), and low spatial resolution (>500m). These correspond to GeoSOT's level 20 (64m), level 15 (2km), and level 9 (128km) grids, respectively. The theoretical and actual grid sizes of these levels are the same, enabling seamless, non-overlapping global coverage with similarly sized subdivision patches, making them easy to use as standard grids for spatial computation and spatial data organization. The selection of the positioning grid location is based on the hemisphere in which the remote sensing image is located, with the specific corner points of the image's bounding rectangle being used as the positioning grid.

[0058] Span code: It consists of six decimal digits, the first three digits are the longitudinal span code, and the last three digits are the latitudinal span code, which records the number of grids spanned by the remote sensing image in the longitudinal and latitudinal directions.

[0059] Organization Code: This is a four-digit decimal number used to identify the source organization of remote sensing data. In practical applications, if the source organization is not required, the organization code can be omitted.

[0060] Multimodal remote sensing data preprocessing Multimodal remote sensing data preprocessing is the first step of this model, which aims to convert the input heterogeneous remote sensing data into a standardized format to provide high-quality data for subsequent grid encoding and feature extraction.

[0061] Processing parameters: Input data includes high-resolution optical satellite imagery (such as the Gaofen series, with a resolution of 2 meters), multispectral imagery (such as Sentinel-2, with a resolution of 10 meters), drone aerial photography (with a resolution of 0.5 meters), and SAR imagery (such as Sentinel-1, with a resolution of 5 meters). Metadata is also required, including acquisition time, sensor type (optical, SAR, etc.), geographic coordinates (latitude and longitude, ground feature sampling distance (GSD), and spectral band information.

[0062] The processing process specifically includes: Atmospheric and albedo time series correction: The 6S model is used to perform atmospheric and albedo time series correction to eliminate cloud and atmospheric influences, keeping the correction error within 5%. This is crucial to ensure data comparability across different time phases and observation conditions.

[0063] Geometric and morphological correction: feature matching is performed using SIFT (Scale-Invariant Feature Transform) and RANSAC (Random Sample Consensus) algorithms, spatial coordinates are aligned, and then projected into a unified coordinate system.

[0064] Complex terrain or low-quality image data: When feature points are sparse or image quality is low (e.g., cloud cover, shadows, sensor noise), traditional matching methods based on local features may fail. To this end, the present invention adopts the following strategies: Multi-scale feature fusion: In addition to SIFT features, more robust feature descriptors (such as SURF or ORB) are introduced, and a multi-scale fusion strategy is adopted to capture image features at different scales and enhance the matching success rate.

[0065] Auxiliary correction based on template matching: For areas with extremely sparse feature points, auxiliary correction can be performed in combination with active template matching algorithms based on correlation or mutual information. This is particularly suitable for areas with small local texture changes.

[0066] Integrating global geometric constraints: In the RANSAC iteration process, in addition to minimizing the reprojection error, global geometric constraints (such as the relative positions of macro objects such as image edges, roads, and water bodies) are also introduced to further eliminate mismatched points and improve correction accuracy.

[0067] Manual interaction and correction: When automated processing cannot meet the accuracy requirements, manual interaction is allowed, and professionals can manually correct a small number of key control points to ensure high-precision geometric correction.

[0068] Image geometric correction: To address the inherent side-view imaging characteristics and speckle noise, multi-look processing and speckle filtering (such as Lee or Frost filtering) are performed before SIFT / RANSAC matching to reduce the impact of noise and enhance the stability of feature point extraction. At the same time, terrain correction is performed in conjunction with high-precision DEM (digital elevation model) data to eliminate geometric distortion caused by terrain fluctuations and ensure accurate registration of SAR and optical images.

[0069] Unified Coordinate System Projection Method: Preprocessed image data is uniformly projected into the UTM (Universal Transverse Mercator) coordinate system. The UTM coordinate system uses the Transverse Mercator projection, dividing the Earth into 60 regions. Each region uses an independent coordinate system. This minimizes projection distortion within a local area and provides highly accurate planar coordinates, facilitating subsequent quantitative analysis and multi-source data fusion. For images spanning multiple UTM zones, the zone containing the center of the covered area is used as the primary projection zone. Alternatively, other commonly used geographic coordinate systems (such as the WGS84 latitude and longitude coordinate system) can be selected for projection based on user needs.

[0070] Data scale unification: Batch Normalization and dynamic slicing techniques are used to unify the data scale, normalizing the resolution to 10 meters or the user-specified scale. This enables the processing of remote sensing images with different resolutions within a unified framework, overcoming the challenges posed by the heterogeneity of multi-source data.

[0071] Noise processing: The block-matching 3D filtering algorithm is used for noise processing to further improve the signal-to-noise ratio of the image and ensure data quality. High-quality input data is the basis for subsequent model performance.

[0072] Obtained parameters: The output after preprocessing is standardized image data (in TIFF format), which contains corrected pixel values and unified metadata (such as timestamps, coordinates, and band information).

[0073] Efficient spatio-temporal grid integration coding In the efficient spatio-temporal grid integration coding step, multi-source remote sensing data is uniformly organized through grid coding technology, and rich spatio-temporal geographical information is embedded. The specific implementation is as follows: Processing parameters: The input is the image data and metadata after preprocessing, including longitude and latitude, GSD, acquisition time, sensor attributes, etc.

[0074] Multi-level spatial grid generation: A quadtree structure is used to generate multi-level spatial grids, supporting dynamic adaptive partitioning. This means that the grid granularity can be adjusted according to terrain density. For example, fine grids are used in urban areas and coarse grids are used in sparse areas, thereby optimizing data storage and retrieval efficiency while ensuring accuracy.

[0075] Definition threshold for high and low resolutions and determination of grid levels: Resolution threshold: Preferably, the image resolution is divided into three levels: High resolution: GSD ≤ 5 meters Medium resolution: 5 meters < GSD ≤ 30 meters Low resolution: GSD > 30 meters Grid level mapping: High-resolution images (GSD ≤ 5 meters) correspond to the fine levels of the quadtree (e.g., level 12 or above), with a grid side length of 10 meters.

[0076] Medium-resolution images (5 meters < GSD ≤ 30 meters) correspond to the medium levels of the quadtree (e.g., levels 9 - 11), with a grid side length of 30 meters.

[0077] Low-resolution images (GSD > 30 meters) correspond to the coarse levels of the quadtree (e.g., level 8 or below), with a grid side length of 100 meters.

[0078] Processing near the boundary: When the image resolution is near the boundary between high and low resolution, the following strategy is adopted: Fuzzy logic judgment: Fuzzy logic is introduced to calculate the membership of the image to different resolution levels according to the GSD value, and the final grid level is determined according to the weighted membership.

[0079] Multi-scale fusion: Extract features from two adjacent grid levels simultaneously, and adaptively fuse features from different levels based on local image features (such as texture complexity and information entropy).

[0080] Meshing considering local characteristics of image data: Local high-contrast areas: In local high-contrast areas, a finer grid is used to preserve more detail information. For example, the local grid level is dynamically adjusted by calculating the gradient magnitude or variance of the local area and comparing it with the global threshold.

[0081] Complex object distribution: For areas with complex object distribution (such as urban areas interlaced with farmland), adaptive meshing based on object type or semantic segmentation results is used. For example, the mesh size can be adjusted based on the complexity of the object type, or an irregular mesh can be used to better fit the object boundaries.

[0082] Calculation method: Input preprocessed image data (raster data) and metadata (GSD).

[0083] According to the GSD value, the initial grid level L is calculated using the following formula: L = log2(MaxTileSize / GSD) + BaseLevel Among them, MaxTileSize is the maximum grid side length (for example, 100 meters), and BaseLevel is the base level (for example, level 8).

[0084] Fuzzy logic judgment (optional): Calculate the membership of GSD to different resolution levels μ_high, μ_mid, μ_low.

[0085] Final grid level L_final = μ_high * L_high + μ_mid * L_mid + μ_low * L_low.

[0086] Local adaptive adjustment (optional): Calculate the gradient magnitude G or information entropy E of the local area.

[0087] If G >G_threshold or E >E_threshold, then the local grid level L_local = L + ΔL (ΔL is a positive number).

[0088] Outputs a quadtree mesh with adaptive levels.

[0089] Geographic Information Embedding: Combined with the geographic embedding module, geographic information such as longitude and latitude, GSD, etc. is embedded and encoded to generate a spatial feature vector. Geographic information embedding is a key component of the basic model. It regards the remote sensing image as a set of square grids and encodes them to obtain more representative geocoding features. Geographic information embedding queries the nearest grid level based on the GSD of the remote sensing image and embeds the longitude and latitude information of the four corner points. In addition, it accurately encodes the GSD through the ground-scale position encoding vector, thereby helping the model understand the spatial range and frequency characteristics of the image. For example, low GSD images contain more high-frequency details. This explicit integration of geographic information enhances the model's understanding of spatial patterns, thereby improving its generalization ability in global-scale applications.

[0090] Time coding: Record the image acquisition time and use timestamps or periodic coding (such as sine functions) to represent temporal features. This enables the model to capture changes in ground features over time and supports multi-temporal analysis.

[0091] Metadata encoding: describes sensor attributes (such as band and resolution) to form a unified feature vector. This ensures the compatibility of data from different sensors.

[0092] Efficient storage and retrieval: The encoding process supports parallel processing and fault tolerance, ensuring efficient storage and fast retrieval. The encoding consistency error is less than 5%, and encoding and indexing 1TB of data takes less than 1 hour.

[0093] Parameters obtained: The output is grid-encoded data, where each grid cell contains a feature vector (such as [lat, lon, time, band1, band2, ..., bandN]).

[0094] Multi-scale feature extraction Multi-scale feature extraction aims to extract rich, multi-level features from grid-encoded data to support various subsequent remote sensing analysis tasks.

[0095] Processing parameters: The input is grid encoded data, that is, the feature vector of each grid cell.

[0096] The processing process specifically includes: Global Dependency Capture: The Visual Transformer (ViT) uses a multi-head self-attention mechanism to capture global dependencies and extract the overall semantic features of the image. ViT effectively models long-range dependencies in images, which is crucial for understanding the overall semantics of complex scenes. Its core is the self-attention mechanism. For the input feature sequence X = [x1, x2, ..., xN], a linear transformation is performed to generate the query (Query) Q, key (Key) K, and value (Value) matrices. The attention calculation formula is as follows: Among them, d k is the dimension of the key vector. The multi-head self-attention mechanism projects the input into multiple different subspaces, calculates the attention separately, and then concatenates the results and linearly transforms them again.

[0097] Local texture feature extraction: This approach combines convolutional neural networks (CNNs) to extract local texture features through multi-scale convolution and pooling. CNNs excel at capturing local details and texture information, complementing ViT. Multi-scale convolutions capture local information at different scales by using kernels of varying sizes or dilated convolutions. Pooling operations (such as max pooling or average pooling) are used to reduce the dimensionality of feature maps and extract key features.

[0098] Feature fusion and optimization: Optimize feature expression through cross-level feature distillation technology, fuse global and local features, improve the model's ability to represent complex scenes, and fuse features extracted by different modules (ViT and CNN) to optimize feature representation by minimizing the difference between the student model output and the teacher model output.

[0099] Anchored Feature Extraction Strategy: An anchored feature extraction strategy is designed to dynamically adjust the mask ratio (75%-90%) to capture spatiotemporal feature relationships and enhance pre-training effectiveness. This anchored feature extraction strategy is a key contribution of the proposed model, aiming to address the challenges faced by existing remote sensing self-supervised learning methods when processing heterogeneous multi-source data. Traditional random masking strategies can leak high-resolution information when reconstructing image patches at the same location but different resolutions, leading to shortcut learning. The anchored feature extraction strategy dynamically adjusts the masking strategy to adapt to different input image sets based on the metadata (spatial resolution, temporal, and spectral information) of preselected anchor images. This strategy optimizes feature learning for the spatial, temporal, and spectral diversity unique to remote sensing data, preventing the model from performing simple reconstruction tasks.

[0100] Data classification standards and methods: This strategy classifies input data by parsing the metadata information of each remote sensing image to clarify its acquisition time (accurate to the day), sensor type (such as Sentinel-2, Landsat-8, Gaofen-1, etc.) and specific batch or source identifier.

[0101] Data from different sources at the same time: When two or more images are collected at the same time (for example, on the same date, with a small error of 1-2 days allowed), but have different sensor types or source identifiers, they are considered "data from different sources at the same time." For example, images of the same area collected on the same day by Sentinel-2 and Landsat-8 satellites.

[0102] Same-Source, Different-Time Data: When two or more images have the same sensor type or source identifier but were acquired at different times, they are considered "same-source, different-time data." For example, images of the same area acquired by the Sentinel-2 satellite on different dates (e.g., January 1, 2023, and February 1, 2023).

[0103] Other cases: Image data that does not meet the above two conditions are classified as "other cases", for example, images from different times and different sources.

[0104] Specific mask implementation and parameter determination: The mask generation algorithm is primarily based on the concepts of random block masking or grid masking, and is implemented by dynamically adjusting the proportion of the masked area. The mask parameters are determined based on balancing the difficulty of the reconstruction task with the effectiveness of feature learning, avoiding information leakage or excessive masking that leads to low learning efficiency.

[0105] Consistent mask policy: Applicable scenario: data from different sources at the same time.

[0106] Implementation method: For multiple heterogeneous images with the same spatial location and the same acquisition time (for example, Sentinel-2 and Landsat-8 image the same area on the same day), generate exactly the same mask M. This means that in these images, pixels at the same spatial location are either masked or retained.

[0107] Mask generation: First, a random binary mask matrix M∈{0,1}H×W is generated for one of the images, where H and W are the number of image patches. The mask ratio controls the percentage of the masked area, for example, sum(M=0) / (H×W) is between 75% and 90%. Then, the mask is applied to all images from different sources at the same time. ,in represents the jth original remote sensing image, represents the j-th image after applying the mask, j=1,…,k, where ⊙ represents element-wise multiplication.

[0108] This strategy forces the model to learn sensor-independent, cross-modal geographic semantic information from images from different sources, preventing the model from reconstructing low-resolution information solely from high-resolution information.

[0109] Mutex mask strategy: Applicable scenario: data from the same source at different times.

[0110] Implementation: For multiple time series images of the same spatial location and source but acquired at different times (for example, Sentinel-2 imaging the same area on different days), mutually exclusive masks are generated. This means that if a pixel is masked in an image from one time phase, the pixel at the same spatial location is retained in images from other time phases, and vice versa.

[0111] Mask Generation: For N phases of images It1, It2, …, ItN, a random mask M1 is first generated for the first phase. Then, masks complementary to the previous phase mask are generated for subsequent phases. For two phases, M2 = 1 − M1. For more than two phases, allocation can be performed using a strategy that ensures that each pixel position is masked at least once and retained at least once in all phases, for example, by using a hash function or cyclic shift to generate mutually exclusive masks.

[0112] This strategy forces the model to learn the dynamic characteristics of objects changing over time and the contextual information of temporal changes from images of different phases, rather than simple pixel reconstruction.

[0113] Random mask strategy: Applicable scenarios: other situations (data from different time periods and sources).

[0114] Implementation: For each image, a random mask is generated independently.

[0115] Mask generation: Randomly select image blocks from the image to mask, and the mask ratio is dynamically adjusted (75%-90%). The size and number of mask blocks can be dynamically adjusted based on the image resolution and scene complexity to ensure the challenge of the task.

[0116] Therefore, as a basic self-supervised learning strategy, spatial context and local texture features are learned from a single image.

[0117] Dynamic band selection strategy: In order to train a band-independent feature encoder, the present invention adopts a dynamic band selection strategy in the pre-training stage.

[0118] Band selection strategy: In each training iteration, for the input anchor image, three bands are randomly selected from all available spectral bands as input channels. For example, for a remote sensing image containing B bands, let the band set be {b1, b2, …, bB}, and each iteration randomly selects a subset of three bands {bi, bj, bk} as input.

[0119] The advantage is that this random selection prevents the model from relying on any specific band combination to complete the reconstruction task, forcing it to learn more generalized and robust feature representations that are independent of the specific sensor band configuration. This means that the model can adapt to any multispectral sensor without requiring sensor-specific parameters or complex band normalization steps, greatly enhancing the model's versatility and applicability.

[0120] The output is a multi-scale feature map containing global semantics, local texture and enhanced details, with cross-scale feature similarity exceeding 85%.

[0121] Cross-task downstream application adaptation The cross-task downstream application adaptation module aims to apply the extracted multi-scale features to various specific remote sensing analysis tasks and achieve efficient task switching and model fine-tuning.

[0122] Processed parameters: The input is a multi-scale feature map, as well as task-specific data (such as classification labels, detection boxes, segmentation masks).

[0123] Processing process: Multi-task Transformer Architecture: Based on the multi-task Transformer architecture, the model supports multiple downstream tasks such as image classification, semantic segmentation, object detection, and change detection through adaptive decision heads and cross-modal fusion technology. This unified architectural design enables the model to flexibly address the needs of different remote sensing applications.

[0124] Efficient fine-tuning: Using feature alignment, transfer learning, and model distillation techniques, we significantly reduce fine-tuning data requirements (to less than 10%) and shorten training time by 30%. This significantly reduces the human and material costs of model deployment and application, and addresses the high cost of manual labeling.

[0125] Pre-trained weights for the model: The pre-trained weights for the model enhance its adaptability to multi-source data. As a foundational model, the rich spatiotemporal features learned from multi-scale datasets provide powerful initialization capabilities for downstream tasks.

[0126] Parameters obtained: The output is task-specific results, such as classification labels, detection boxes, segmentation masks, etc.

[0127] Examples and Verification The system consists of three modules: data integration, query and retrieval, and image distribution. The data integration module supports the integration and import of various remote sensing data, the generation of identification codes, and the creation of logical index files. The query and retrieval module supports longitude and latitude range queries, specific image queries, and rectangular selection queries. The image distribution module supports image stitching and cropping, small patch distribution, and image transmission and downloading.

[0128] System operating environment: Software: Linux operating system (CentOS7 or Ubuntu20.04), relying on Python3.7+, PyTorch1.10+, TensorFlow2.6+, CUDA11.3+, PostgreSQL / PostGIS (for geospatial data management), and Docker / Kubernetes (supporting distributed deployment).

[0129] Hardware: 64-core CPU (≥2.5GHz), 4 GPUs (40GB+ video memory each), 512GB+ RAM, hybrid storage (2TB+NVMeSSD, 10TB+HDD), 10Gbps network (supports RDMA).

[0130] The experimental data used over 20 types of remote sensing data, including both real and simulated data, generated by commonly used satellites and drones at home and abroad. High-resolution imagery includes remote sensing data from drones (Spark, Mavic, and Phantom); medium-resolution imagery includes data from the Gaofen series of satellites (GF1-GF4), the Ziyuan-3 satellites (ZY3-01 and ZY3-02), and the LandSat series of satellites (LandSat1-LandSat5); and low-resolution imagery includes data from the Fengyun series of satellites (FY1-FY4).

[0131] The model uses the ViT-Large architecture as its backbone network, employs a progressive training strategy, and employs a batch size of 1024, training for 130 epochs on eight NVIDIA A800 GPUs. The representativeness of the experimental environment and the rigor of the validation are crucial components of this research. The construction of the prototype system not only demonstrates the practical operational basis of the theoretical approach but also hints at the rigor of its validation. Furthermore, training the model on high-performance computing resources demonstrates the substantial computational resource requirements of the underlying model and its ability to process massive amounts of complex data. Testing with diverse remote sensing data enhances the generalizability and persuasiveness of the experimental results. These findings provide reliable context for subsequent performance analysis and demonstrate the practical value and technical depth of the research.

[0132] Unified identification code generation: After loading 50,000 remote sensing data items, the system immediately constructed a large, segmented index table and loaded the remote sensing data from various databases into it. Unified identification codes were then generated. The generation time for 50,000 codes was 2736.443 milliseconds, significantly less than 3 seconds. The location grid code for high-resolution remote sensing data is 20 bits, for medium-resolution data it is 15 bits, and for low-resolution data it is 9 bits. This is consistent with the grid level selection rules, demonstrating the correctness of the unified identification code generation for remote sensing data.

[0133] When the query data volume is 50,000, the experimental results of changing the query range size show that: In terms of query efficiency, the grid method is more efficient than the comparative method (Oracle Spatial + R-tree index), and the efficiency improvement increases with decreasing query scope. When the query scope is small (2°×2°), the grid query takes 11.017ms, while the comparative query takes 2459.932ms, resulting in a 223.29-fold increase in grid query efficiency. When the query scope is reduced to 1°×1°, the efficiency improvement is nearly three orders of magnitude. This is attributed to the grid method's use of the multi-scale encoding feature to perform a preliminary screening of all remote sensing images before the actual query, significantly reducing the query workload. As the query scope gradually decreases, the number of images excluded from the preliminary screening increases, and the advantages of the grid query gradually increase.

[0134] In terms of query accuracy, grid queries have a very high accuracy rate, generally approaching 100%. For example, when the query range is moderate or large, the grid method retrieves slightly more images than the comparison method. This is because the grid method queries in grid units, and each grid occupies a certain latitude and longitude range, similar to a buffer zone for the query range. The accuracy of grid queries increases as the grid size decreases.

[0135] In addition, a comprehensive performance evaluation of the common base model was conducted, covering various downstream tasks such as image classification, semantic segmentation, and change detection, and experiments were conducted on 7 downstream datasets with diverse spatial, temporal, and spectral coverage.

[0136] Image classification: On the AID and BigEarthNet datasets, our model outperforms all competing methods, including those pre-trained for Sentinel-2 (such as SeCo and CACo), highlighting the effectiveness of our model in processing diverse remote sensing information with a unified model.

[0137] Semantic Segmentation: On the Sen1Floods11 and CropSeg datasets, our model achieved significant improvements over methods such as SeCo and CACo. In particular, on datasets covering diverse land cover types, our model demonstrated excellent generalization capabilities. For example, on Sen1Floods11, our method achieved improvements in mIoU (mean Intersection Over Union) by 9.77% / 12.03% and 4.25% / 11.98% over SeCo and CACo, respectively.

[0138] Change Detection: Our model outperforms competing methods on the LEVIR-CD, OSCD, and DynamicEarthNet datasets. On LEVIR-CD, OSCD, and DynamicEarthNet, our model achieves 1.25% improvement in mean Intersection Over Union (MIoU), 2.23% improvement in F1, and 1.3% improvement in mIoU over the next best results, respectively. These improvements highlight the effectiveness of our model in leveraging multi-temporal remote sensing imagery within a unified framework.

[0139] Compared to vision-based models such as DINOv2, the proposed model demonstrates superior performance in remote sensing tasks, achieving the highest accuracy or mIoU / F1 scores in classification, segmentation, and change detection tasks. For example, on the BigEarthNet dataset, DINOv2 achieves a mAP of 81.2% with 3-band input, lagging behind SeCo's ResNet-50 (82.6%). This is primarily due to the domain differences and insufficient spectral sensitivity of the pre-trained dataset. The proposed model, pre-trained on this dataset and combining GEM with full-band flexibility, achieves a mAP of 74.3% with 12-band input, surpassing SeCo's 72.6% (3-band), highlighting its advantages in remote sensing design.

[0140] In summary, the present invention provides a method for constructing and analyzing a universal basic model for remote sensing images based on grid coding. Compared with the existing technologies, the present invention has the following significant advantages in remote sensing image analysis: Reduced dependency on annotations: Through weakly supervised learning and self-supervised learning, the dependency on manual annotations is significantly reduced, thereby reducing data preparation costs.

[0141] Improved Data Integration Efficiency: Grid coding technology improves the efficiency of multi-source data integration through a quadtree structure and geographic embedding modules, supporting rapid retrieval and fusion. Experimental results show that coding consistency error is less than 5%, and the encoding and indexing time for 1TB of data is reduced to less than 1 hour.

[0142] Enhanced model generalization capability: Multi-scale feature extraction and cross-task adaptation enhance the model's generalization capability, making it suitable for various scenarios such as smart cities, agricultural monitoring, and land cover mapping.

[0143] Excellent performance: In tests using high-resolution data, the proposed method achieved classification accuracy exceeding 90%, segmentation Intersection over Union exceeding 80%, and cross-scale feature similarity exceeding 85%, outperforming existing models such as SatMAE. In the global land cover mapping task, the proposed method achieved a 7% improvement in accuracy, completing the global map in just 4 minutes on a supercomputer, significantly improving efficiency.

[0144] Addressing challenges of heterogeneous data: Compared with traditional methods, this invention can better address issues such as data heterogeneity, temporal synchronization, and spectral differences when processing multimodal data, providing an efficient solution for large-scale remote sensing applications.

[0145] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0146] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0147] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.

[0148] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc or a read-only optical disc (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0149] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0150] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0152] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0153] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0154] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0155] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0156] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0157] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for constructing and analyzing a universal basic model of remote sensing images based on grid coding, characterized in that: The following steps are involved: S1. Multimodal Data Preprocessing: Temporal correction is performed on heterogeneous remote sensing data to eliminate cloud effects. Geometric correction using multi-scale feature fusion and template matching is used to align spatial coordinates and project them into a unified coordinate system. Dynamic slicing technology is used to unify data scales and generate standardized image data. S2. Efficient spatiotemporal grid integration encoding: Dynamically determines the quadtree grid level based on the resolution threshold, combines the geographic embedding module to encode the latitude and longitude of the four corners of the image and the ground sampling distance, and integrates the time code with metadata code to generate a unified identification code and grid code feature vector with geographical meaning; S3. Multi-scale feature extraction: We use the multi-head self-attention mechanism of the visual Transformer to capture global dependencies, combine it with a convolutional neural network to extract local texture features, and use an anchored feature extraction strategy to dynamically generate masks to optimize multi-source data fusion. S4. Cross-task downstream application adaptation: Based on the multi-task Transformer architecture, feature transfer learning technology is used to adapt features to image classification, semantic segmentation, and change detection tasks.

2. The method according to claim 1, wherein In step S1: The geometric correction is aimed at complex terrain or low-quality images, and uses multi-scale feature fusion and global geometric constraints to enhance matching robustness; Perform multi-look processing, speckle filtering, and DEM terrain correction on SAR images; The unified coordinate system projection adopts the UTM partition mechanism, and cross-partition images are based on the coverage area center or the user-specified coordinate system.

3. The method for constructing and analyzing a universal basic model of remote sensing images based on grid coding according to claim 1, wherein: In step S2, the unified identification code is composed of metadata code, spatial grid code and unit code, where: The spatial grid code includes the positioning grid code and the span code that records the grid span in the longitude and latitude directions; The positioning grid code is mapped to a quaternary one-dimensional code at a preset level according to the image resolution.

4. The method according to any one of claims 1 to 3, wherein In step S2: The grid coding uses fine-grained grids for urban areas and coarse-grained grids for sparse areas, thereby optimizing storage and retrieval efficiency.

5. The method according to any one of claims 1 to 3, wherein In step S2: The resolution threshold is divided into: high resolution is mapped to a fine grid level, medium resolution is mapped to a medium grid level, and low resolution is mapped to a coarse grid level; The boundary resolution uses fuzzy logic to calculate the membership weight to determine the grid level, or is adaptively adjusted through multi-scale feature fusion.

6. The method according to any one of claims 1 to 3, wherein: In step S2: According to the mapping rules between remote sensing image resolution and grid level, a dynamic hierarchical division model is established through logarithmic calculation. For local high-contrast areas in the image, such as areas with complex textures or significant boundary features of land objects, feature evaluation is performed based on the gradient amplitude or information entropy index in the area, and the degree of grid division is dynamically adjusted to enhance the ability to retain spatial detail information.

7. The method according to any one of claims 1 to 3, wherein: The anchor feature extraction strategy in step S3 includes: The data were classified into the following categories according to the collection time and source: same time but different sources, same source but different time and other situations; Generate the same mask matrix for data from different sources at the same time and implement consistent masking; Generate complementary mask matrices for data from the same source at different times to implement mutually exclusive masking; Generate random masks independently for the rest of the data; Three spectral bands are randomly selected to train the band-independent feature encoder.

8. The method according to any one of claims 1 to 3, wherein: In step S3: In the multi-scale feature extraction step, the multi-head self-attention mechanism of the visual Transformer is used to construct an associative calculation model of the query, key, and value matrices to capture the global semantic dependencies of the remote sensing imagery. At the same time, combined with a convolutional neural network, multi-scale convolution kernels or dilated convolution operations are used to extract local texture features of different scales in the image, forming a complementary fusion of global semantics and local texture features.

9. The method according to any one of claims 1 to 3, wherein: In step S3: The mask ratio is dynamically controlled at 75%-90%, and the mask generation adopts random block mask or grid mask algorithm.

10. The method according to any one of claims 1 to 3, wherein In step S4: The cross-task downstream application adaptation uses feature distillation technology to fuse ViT and CNN output features, and shares encoder weights to deploy task-specific decoders.

Citation Information

Cited By

  • Low-altitude remote sensing image AI large model identification training method and system

    CN121191001A

  • A method and system for training and recognizing large-scale AI models of low-altitude remote sensing images

    CN121191001B

  • Distributed self-adaptive framing output method based on large-range remote sensing image

    CN121214183A

  • Method for encoding and decoding spatio-temporal data of three-dimensional space

    CN121327055A

  • Natural field type extraction method based on remote sensing large model pre-training and multi-granularity boundary supervision

    CN121505458A