General remote sensing data interpretation system, method, equipment and medium

By providing a general remote sensing data interpretation system, the problems that are difficult to deal with in the existing technology of multi-task and multi-modal requirements are solved, efficient and flexible model training and verification are achieved, and the accuracy and reliability of remote sensing data interpretation are improved.

CN120014476APending Publication Date: 2025-05-16XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056238.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology is difficult to cope with the needs of multi-task and multi-modal remote sensing data interpretation. The existing models need to be adjusted and optimized in a large number of ways to face new tasks or data sets, and the data set quality is uneven, which affects the stability and generalization capabilities of the model.

Method used

It provides a general remote sensing data interpretation system, including remote sensing data module, basic model module, interpretation module, unified training module and model verification module. Through the unified training module, the data set, model and interpretation algorithm are automatically loaded according to the training instructions specified by the user, and intelligent model training is carried out, and strict verification and evaluation is carried out through the model verification module.

Benefits of technology

It realizes efficient processing of multi-task and multi-modal remote sensing data interpretation, improves the flexibility and customization of the model, ensures the accuracy and reliability of the model, and enhances the practical application effect in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014476A_ABST
    Figure CN120014476A_ABST
Patent Text Reader

Abstract

The invention discloses a general remote sensing data interpretation system, method and device and a medium, and belongs to the technical field of image processing, and the system comprises a remote sensing data module which is used for storing a remote sensing data set; the basic model module is used for integrating basic models; the interpretation module is used for integrating interpretation algorithms; the unified training module is used for training the basic model according to the training instruction; the training instruction at least comprises one or more combinations for loading a remote sensing data set, one or more combinations for loading a basic model and one or more combinations for loading an interpretation algorithm; the model verification module is used for verifying and evaluating the basic model trained each time; the technical problems that the remote sensing data interpretation tasks are numerous, the algorithm difference needed by the tasks is large, and the multi-task and multi-mode requirements are difficult to meet are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing and relates to a universal remote sensing data interpretation system, method, equipment and medium. Background Art

[0002] In the field of remote sensing science, researchers have long faced the challenge of complex combinations of multiple models and tasks. With the rapid development of remote sensing technology, the interpretation requirements for remote sensing data in different application scenarios are becoming increasingly diverse, which requires researchers to be able to select and train appropriate models for specific tasks. However, this process is often time-consuming and laborious, and requires a lot of resource investment, including but not limited to computing resources, data resources, and human resources.

[0003] Traditionally, remote sensing data interpretation models are often designed specifically for specific tasks or data sets. Although this design approach ensures the pertinence and performance of the model to a certain extent, it also greatly limits the versatility and scalability of the model. When faced with new tasks or data sets, existing models often require a lot of adjustments and optimizations, and may even need to be redesigned, which undoubtedly increases the complexity and cost of research.

[0004] In addition, remote sensing data interpretation tasks are diverse in practical applications, and the algorithms required for each task vary greatly. Traditional methods are unable to cope with these multi-task and multi-modal requirements, and it is difficult to achieve efficient processing while ensuring performance. This limitation not only limits the application scope of remote sensing data interpretation, but also affects its actual application effect in different fields.

[0005] On the other hand, existing remote sensing datasets are usually scattered and unsystematic, which brings great difficulties to the training and evaluation of models. The quality of datasets varies, making it difficult to ensure the stability and reliability of models on different datasets. At the same time, existing datasets often fail to cover a wide range of practical application scenarios, resulting in limited generalization capabilities of models in practical applications. Summary of the invention

[0006] The purpose of the present invention is to solve the technical problem that there are many remote sensing data interpretation tasks in the prior art, and the algorithms required for each task are quite different, making it difficult to cope with multi-task and multi-modal requirements, and to provide a universal remote sensing data interpretation system, method, device and medium.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides a universal remote sensing data interpretation system, comprising: Remote sensing data module, used to store remote sensing data sets; Basic model module, used to integrate basic models; An interpretation module, used to integrate interpretation algorithms; A unified training module, used to train the basic model according to the training instructions; the training instructions at least include loading one or more combinations of remote sensing data sets, loading one or more combinations of basic models, and loading one or more combinations of interpretation algorithms; The model verification module is used to verify and evaluate the basic model after each training.

[0008] Furthermore, the training instructions also include optimizer configuration, learning rate configuration and loss function configuration.

[0009] Furthermore, the basic model includes a basic model based on a CNN architecture, a basic model based on a Transformer architecture, and a basic model based on a Mamba architecture.

[0010] Furthermore, the training instructions are configured through a YAML file.

[0011] Furthermore, the unified training module also performs video memory management.

[0012] Furthermore, the interpretation algorithm includes scene classification, semantic segmentation, change detection, reference image segmentation, visual localization, reference expression understanding and segmentation, visual question answering, object detection, instance segmentation and rotated box object detection.

[0013] A second aspect of the present invention provides a general remote sensing data interpretation method, comprising the following steps: Configure training instructions; Based on the training instructions, the basic model is trained; After each training session, each interpretation algorithm was evaluated.

[0014] Furthermore, each interpretation algorithm is evaluated, and its evaluation indicators include mIOU, mF1, OA, mAcc1, mAcc5, oIoU, PR@.5, PR@.6, PR@.7, PR@.8 and PR@.9.

[0015] A third aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned general remote sensing data interpretation method when executing the computer program.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the general remote sensing data interpretation method described above is implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The invention discloses a universal remote sensing data interpretation system. The remote sensing data module realizes efficient storage and management of remote sensing data sets, supports rapid loading and access to large-scale, multi-type remote sensing data, and provides a solid foundation for subsequent data processing and interpretation. The basic model module integrates a variety of basic models, providing users with a rich selection space. Users can freely select and combine different basic models according to actual needs to build an interpretation model that best suits their application scenarios, achieving high flexibility and customization. The interpretation module integrates a variety of advanced interpretation algorithms, covering multiple fields of remote sensing data processing and interpretation, and can perform efficient and accurate interpretation for different types of remote sensing data to meet the needs of different application scenarios, further improving the accuracy and efficiency of interpretation. The unified training module can automatically load the combination of remote sensing data sets, basic models and interpretation algorithms according to the training instructions specified by the user, and perform intelligent model training. The model verification module can strictly verify and evaluate the basic model after each training to ensure the accuracy and reliability of the model.

[0018] Furthermore, the present invention discloses a universal remote sensing data interpretation method, which first allows users to configure training instructions according to specific needs. This step ensures the flexibility and pertinence of the training process. Users can freely choose a combination of remote sensing data sets, basic models, and interpretation algorithms to meet the needs of different application scenarios. Based on the configured training instructions, the system can automatically train the basic model. During the training process, the system can monitor the training status in real time and automatically adjust the training parameters to ensure that the model can learn the best feature representation. After each training is completed, the system will evaluate each interpretation algorithm. The evaluation process covers multiple aspects such as the accuracy, efficiency, and stability of the algorithm to ensure the reliability and effectiveness of the interpretation algorithm in practical applications. Through comprehensive evaluation, users can intuitively understand the performance of each interpretation algorithm, so as to make more informed choices and optimize the interpretation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 This is a general remote sensing data interpretation system architecture diagram of the present invention; Figure 2 This is a block diagram of the universal remote sensing data interpretation method of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention described and marked in the drawings here can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0024] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0025] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0026] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] The present invention is further described in detail below in conjunction with the accompanying drawings: See also Figure 1The present invention discloses a universal remote sensing data interpretation system, comprising: Remote sensing data module, used to store remote sensing data sets; Basic model module, used to integrate basic models; An interpretation module, used to integrate interpretation algorithms; A unified training module, used to train the basic model according to the training instructions; the training instructions at least include loading one or more combinations of remote sensing data sets, loading one or more combinations of basic models, and loading one or more combinations of interpretation algorithms; The model verification module is used to verify and evaluate the basic model after each training.

[0028] All basic model algorithms and interpretation algorithms of downstream tasks of the universal remote sensing data interpretation system of the present invention can be arbitrarily combined with one click, and a variety of (currently 1800+) different interpretation algorithms are provided for practical application tasks, which can support arbitrary combination training on a variety of (currently 30+) data sets, and can have a variety of (currently 50000+) practical application models, and 5000+ models have been trained and stored; in this embodiment, the one-click arbitrary combination of all basic models and the interpretation algorithms of downstream tasks is performed by setting the training modules such as the data set to be used for training, the basic model, the downstream task interpretation algorithm, the optimizer, the learning rate and the loss function in the key-value pair format of the YAML format file.

[0029] In the training instructions of this embodiment, the configuration of the remote sensing data set can be completed by setting parameters such as dataset (the name of the data set used), data_root (data set storage path), crop_size (the cropping size of the input image), img_suffix (image name prefix), mask_suffix (image noun suffix), need_bands (the number of input image channels), and reduce_zero_label (whether to ignore the zero background class).

[0030] In this embodiment, the basic models include basic models based on CNN, Transformer, Mamba and other architectures, among which the basic models based on CNN architecture include: FocalNet, InternImage, UniRepLKNet, LSKNet, ConvNeXt, SMLFR, the basic models based on Transformer architecture include: MiT, Swin Transformer, Swin Transformerv2, ScaleMAE, SatlasPretrain, Cross ScaleMAE, SatMAE++, RSP, SpectralGPT, HyperSigma, GFM, Vision Transformer, and the basic models based on Mamba architecture include: VMamba. At the same time, each basic model has models of different network scales, such as Tiny, Small, Base, Large, XLarge, Huge, etc.

[0031] The interpretation algorithms in this embodiment include scene classification, semantic segmentation, change detection, reference image segmentation, visual positioning, visual question answering, object detection, instance segmentation and rotating box object detection.

[0032] The training instructions of this embodiment complete the following actions: remote sensing data set loading, basic model loading, interpretation algorithm loading, optimizer configuration, learning rate configuration, loss function configuration and other initialization work. The training process covers the evaluation of various tasks, including scene classification, semantic segmentation, change detection, reference image segmentation, visual positioning, reference expression understanding and segmentation, visual question and answer, target detection, instance segmentation, and rotating box target detection. The universal remote sensing data interpretation system will perform different training tasks according to different training instructions. The unified training module will also perform video memory management to ensure the efficiency and stability of the training process.

[0033] The model verification module of this embodiment optimizes the parameters of the basic model by setting the verification steps of each epoch, saves the best model checkpoint based on the evaluation indicators of the verification set (such as mIOU, mF1, OA, mAcc1, mAcc5, oIoU, PR@.5, PR@.6, PR@.7, PR@.8 and PR@.9, etc.), and performs regular model evaluation and result recording.

[0034] The universal remote sensing data interpretation system of the present invention can arbitrarily combine the basic model and the interpretation algorithm with one click. The remote sensing data module fully considers the modalities and scenes involved in practical applications, covering multiple modal data such as optical data, SAR data, infrared data, multispectral data, digital elevation data, text data, etc. (some multimodal data with paired data), including images collected from multiple platforms such as satellites, UAVs, and aviation, involving diverse scenes such as towns, villages, hydrology, vegetation, farmland, transportation, military, and ports, providing a broad and fair comparison benchmark for different interpretation algorithms. The design of the basic model module allows the system to integrate multiple basic models, providing users with a rich choice space. With the continuous advancement of remote sensing technology and the emergence of new interpretation algorithms, the system can quickly adapt to and integrate these new technologies, thereby maintaining its cutting-edge technology and practicality. The interpretation module can integrate multiple interpretation algorithms, which can perform efficient and accurate interpretation for different types of remote sensing data. The coordinated work of the unified training module and the model verification module realizes an integrated process from model training to verification and evaluation. The user only needs to enter the training instructions, and the system can automatically complete the loading, training, verification and evaluation of the model, which greatly simplifies the operation process and reduces the technical threshold. At the same time, the model verification module can also conduct a comprehensive and objective evaluation of the basic model after each training, providing users with a reliable model performance reference.

[0035] See also Figure 2 The present invention provides a general remote sensing data interpretation method, comprising the following steps: S1, configure training instructions; according to actual needs, clarify the specific goals of remote sensing data interpretation, such as land cover classification, building detection, etc. According to the training goals, select a suitable remote sensing dataset. The dataset should contain a variety of remote sensing images and corresponding labels (such as category labels, location information, etc.) so that the model can learn different features and patterns. According to the characteristics of the dataset and the training goals, configure the training parameters, including learning rate, batch size, number of training rounds, etc. These parameters will directly affect the training effect and performance of the model. Select a combination of basic model and interpretation algorithm; set the training modules such as dataset, basic model, downstream task interpretation algorithm, optimizer, learning rate and loss function to be used in training through the key-value pair format of YAML format files.

[0036] S2, based on the training instructions, train the basic model; initialize the basic model according to the selected interpretation algorithm. Train the basic model according to the configured training instructions and parameters. During the training process, evaluate the training progress and effect of the model by monitoring indicators such as loss function and accuracy. If necessary, adjust the training parameters or optimize the model structure.

[0037] S3, after each training is completed, each interpretation algorithm is evaluated. According to the training objectives, appropriate evaluation indicators are selected, such as mIOU, mF1, OA, mAc1, mAc5 and oIoU, to quantify the performance of each interpretation algorithm; each interpretation algorithm is evaluated based on the evaluation indicators and verification results. The evaluation results will be used to determine which interpretation algorithm performs best on a specific task.

[0038] An embodiment of the present invention provides a general remote sensing data interpretation method, comprising the following steps: Step S1: Construct an efficient and convenient remote sensing constructible unified model framework GRSFM (i.e., General Remote Sensing Data Interpretation System).

[0039] All basic model algorithms and downstream task interpretation algorithms in the remote sensing buildable unified model framework GRSFM can be combined arbitrarily with one click. It provides 1,800+ different interpretation algorithms for practical application tasks, supports any combination training on 30+ data sets, and can have 50,000+ practical application models. 5,000+ model training and storage have been completed; All basic model algorithms and downstream task interpretation algorithms can be combined arbitrarily with one click by setting configuration files in YAML format. The training modules such as the data set, basic model, downstream task interpretation algorithm, optimizer, learning rate, loss function, etc. to be used for training are set in the format of key-value pairs.

[0040] Step S11: Data set configuration.

[0041] By setting parameters such as dataset (name of the dataset used), data_root (dataset storage path), crop_size (cropping size of the input image), img_suffix (image name prefix), mask_suffix (image noun suffix), need_bands (number of input image channels), and reduce_zero_label (whether to ignore the zero background class), you can complete the configuration of the dataset used for training.

[0042] Step S12: Basic model configuration.

[0043] By setting parameters such as model (basic model name), you can complete the configuration of the basic model used for training.

[0044] Step S13: downstream task interpretation algorithm configuration.

[0045] By setting parameters such as decoder (name of downstream task interpretation algorithm), you can complete the configuration of the downstream task interpretation algorithm used for training. Sometimes, according to the needs of the model, the neck parameter is also set to further process the feature map output by the basic model network so as to better pass it to the downstream interpretation algorithm for final prediction.

[0046] Step S14: optimizer, learning rate, and loss function configuration.

[0047] By setting parameters such as optimizer, lr_scheduler (learning rate), criterion (loss function), etc., you can complete the configuration of the optimizer, learning rate, and loss function used in training.

[0048] Step S2: Construct a high-quality remote sensing application database under a unified framework.

[0049] We have selected more than 30 public data sets from top journals, top conferences, and international scientific research competitions in the past 10 years, with a total annotation volume of more than 800K images or image-text pairs. We have fully considered the modalities and scenarios involved in practical applications, covering multiple modal data such as optical data, SAR data, infrared data, multispectral data, digital elevation data, text data, etc. (some multimodal data with paired data), including images collected from multiple platforms such as satellites, UAVs, and aviation, involving diverse scenarios such as towns, villages, hydrology, vegetation, farmland, transportation, military, and ports, providing a broad and fair comparison benchmark for different interpretation algorithms; The specific dataset information supported by the remote sensing application database is shown in Table 1. The table contains multiple datasets for different tasks. The tasks mainly include semantic segmentation, change detection, reference image segmentation, visual positioning, object detection, instance segmentation, and rotation box object detection. The specific tasks are as follows:

[0050] Table 1 The Potsdam dataset is used for semantic segmentation tasks of optical modalities and digital surface modalities. It contains 6 categories, 3456 training samples and 2016 validation samples. It supports multimodal processing and is published in the ISPRS journal.

[0051] The Potsdam_woclutter dataset is used for semantic segmentation tasks of optical modalities and digital surface modalities. It contains 5 categories, 3456 training samples and 2016 validation samples, supports multimodal processing, and is published in the ISPRS journal.

[0052] The DeepGlobe dataset is also used for semantic segmentation in the optical modality. It contains 7 categories, 5760 training samples and 1467 validation samples. It was used in the CVPR 2018 challenge.

[0053] The LoveDA dataset is also used for semantic segmentation in the optical modality. It contains 7 categories, 2522 training samples and 1669 validation samples, and was published in the NIPS 2021 journal.

[0054] The WHU_OPT_SAR dataset combines optical and SAR modalities for semantic segmentation, contains 7040 training samples and 1760 validation samples, supports multimodal processing, and is published in the JAG 2022 journal.

[0055] The DFC24_T1 dataset is used for DEM tasks. It has 2 categories, 1,304 training samples, and 327 validation samples. It supports multimodal processing and was used in the IEEE GRSS 2024 Challenge.

[0056] The Agriculture_Vision dataset is used for semantic segmentation and change detection tasks combining infrared and optical images. It contains 9 categories, 56,944 training samples and 18,334 verification samples. It was used in the CVPR 2024 challenge.

[0057] The SPARCS dataset is a multispectral dataset containing 5 categories, 1024 training samples and 256 validation samples, published in the RS 2014 journal.

[0058] The SegMunich dataset is also a multispectral dataset, containing 13 categories, 7872 training samples and 1974 validation samples, and was published in the TPAMI 2024 journal.

[0059] The LEVIR-CD+ dataset is used for change detection tasks in optical modality. It contains 2 categories, 10,192 training samples and 5,568 validation samples. It was published in the RS 2020 journal.

[0060] The SIGFloods dataset is applied to change detection in the SAR mode. It contains 2 categories, 4,300 training samples and 1,060 validation samples. It was published in the ISPRS 2024 journal.

[0061] The Hi-CNA dataset combines infrared and optical modalities for change detection tasks. It contains 2 categories, 4080 training samples and 1358 validation samples. It was published in the ISPRS 2024 journal.

[0062] The OSCD dataset is a multispectral modality used for change detection. It contains 2 categories, 496 training samples and 230 validation samples. It was published in the IEEE GRSS 2018 journal.

[0063] The RefSegRS dataset is used for directional image segmentation tasks. It contains 2 categories, 2603 training samples and 1817 validation samples, and is published in the TGRS 2024 journal.

[0064] The RRSIS-D dataset is also used for directional image segmentation. It contains 2 categories, 13,921 training samples and 3,481 validation samples. It was published in the CVPR 2024 journal.

[0065] The DIOR-RSVG dataset is used for visual positioning, contains 1 category, 30,820 training samples and 7,500 validation samples, and is published in the TGRS 2024 journal.

[0066] The OPT-RSVG dataset combines optics and language for visual positioning tasks. It contains 1 category, 39,162 training samples and 9,790 validation samples. It was published in the TGRS 2024 journal.

[0067] The EarthVQA dataset is used for visual question answering tasks, contains 6 categories, 145,368 training samples and 63,225 validation samples, and is published in the AAAI 2024 journal.

[0068] The VRSBench dataset is also used for visual question answering tasks, containing 10 categories, 85,813 training samples and 37,408 verification samples, and was published in the ArXiv 2024 journal.

[0069] The DIOR dataset is used for object detection tasks in optical modality. It contains 20 categories, 11,725 ​​training samples and 11,738 validation samples. It was published in the ISPRS 2020 journal.

[0070] The Drone dataset combines infrared and optical modalities and is also used for object detection tasks. It contains 5 categories, 19,459 training samples and 8,980 verification samples, and is published in the TCSVT 2022 journal.

[0071] The SARDet-100k dataset is used for vehicle detection tasks in the SAR mode. It contains 6 categories, 94,493 training samples and 10,492 validation samples, and was published in the ArXiv 2024 journal.

[0072] The SSDD dataset is also used for vehicle detection in the SAR mode. It contains 1 category, 928 training samples and 232 validation samples, and was published in the BIGSAR DATA 2017 journal.

[0073] The DFC23_T1 dataset combines SAR and optical modalities for instance segmentation tasks. It contains 12 categories, 3,000 training samples, and 720 validation samples. It was used in the IEEE GRSS 2023 Challenge.

[0074] The CVPPA dataset is used for instance segmentation in the optical modality. It contains 1 category, 1407 training samples and 772 verification samples. It was used in the ICCV 2023 challenge.

[0075] The HRSID dataset is used for instance segmentation tasks in the SAR modality. It contains 1 category, 3642 training samples and 1962 validation samples. It was published in the IEEE Access 2020 journal.

[0076] The DIOR-R dataset is used for the rotating box target detection task in the optical modality. It contains 20 categories, 11,725 ​​training samples and 11,738 verification samples, and is published in the TGRS 2022 journal.

[0077] The DOTA v1.0 dataset is used for the rotating box target detection task in the optical modality. It contains 15 categories, 21,046 training samples and 10,833 validation samples. It was published in the CVPR 2018 journal.

[0078] The SSDD+ dataset is used for the rotating box target detection task in the SAR modality. It contains 1 category, 928 training samples and 232 validation samples. It was published in the BIGSARDATA 2017 journal.

[0079] Step S3: Build a basic model algorithm library under a unified framework.

[0080] The basic model algorithm library under the unified framework supports a total of 60+ basic models based on CNN, Transformer, Mamba and other architectures, with model parameter quantities ranging from a minimum of 3M to a maximum of 1.1B, providing remote sensing scientific researchers with a wealth of model selection and comparison benchmarks. Based on the scientific research experience of the team of this invention, the GRSFM library has selected excellent model algorithms with excellent generalization performance and downstream task performance from top journals and conference articles in the past four years. The models include a variety of pre-trained learning algorithms based on fully supervised learning, generative self-supervised learning, comparative self-supervised learning, and specific remote sensing data, providing a broad domain knowledge prior for subsequent practical application scenarios of remote sensing interpretation; Step S31: Basic model algorithm library under a unified framework.

[0081] The basic model algorithm library under the unified framework supports basic models based on CNN, Transformer, Mamba and other architectures. The basic models based on CNN architecture include: FocalNet, InternImage, SMLFR, ConvNeXt, UniRepLKNet, LSKNet; the basic models based on Transformer architecture include: MiT, Swin Transformer, Swin Transformer v2, ScaleMAE, SatlasPretrain, Cross ScaleMAE, SatMAE++, RSP, SpectralGPT, HyperSigma, GFM, Vision Transformer; the basic models based on Mamba architecture include: VMamba. At the same time, each basic model has models of different network scales, such as Tiny, Small, Base, Large, XLarge, Huge, etc., a total of 60+ basic models, with model parameter quantities ranging from a minimum of 3M to a maximum of 1.1B. The basic model information supported by the basic model algorithm library is shown in Table 2. The specific algorithms are described as follows:

[0082] Table 2 ConvNeXt series (CVPR 2022): ConvNeXt is a model based on convolutional neural network (CNN), which is divided into five scales: Tiny, Small, Base, Large and XLarge. The number of parameters ranges from 27.82M to 348.15M, with high expressiveness and efficiency.

[0083] SMLFR-ConvNeXt Base (TGRS 2024): This is a convolutional model designed specifically for remote sensing image tasks. It is based on the Base version of the ConvNeXt architecture, has 87.57M parameters, and is suitable for processing large-scale remote sensing image data.

[0084] MiT series (NIPS 2021): MiT is a variant of the Transformer model, including six versions from B0 to B5, with parameters ranging from 3.32M to 81.44M. This model is optimized for visual tasks and is particularly suitable for training large-scale datasets.

[0085] PVTv2 series (ICCV 2021): This model is an upgraded version of PVT, including six sub-models from B0 to B5, with parameters ranging from 3.41M to 81.44M, focusing on feature extraction and enhancement in visual tasks.

[0086] LSKNet series (ICCV 2023): LSKNet is a new network architecture, mainly used for image classification and semantic segmentation. It is divided into two versions: Tiny and Small, with 3.99M and 13.84M parameters respectively.

[0087] Vision Transformer (ICLR 2021): Vision Transformer is a pioneering model that applies Transformer to computer vision tasks. It includes two versions: Base and Large, with 85.80M and 303.30M parameters respectively.

[0088] Swin Transformer series (ICCV 2021): Swin Transformer is a model based on the window multi-head self-attention mechanism. It is divided into four versions: Tiny, Small, Base and Large. The number of parameters ranges from 27.52M to 195.00M. It focuses on learning hierarchical image representations.

[0089] Swin Transformer v2 series (CVPR 2022): This is an improved version of Swin Transformer that supports larger-scale training and more complex tasks. It is divided into three versions: Tiny, Small, and Base.

[0090] SatlasPretrain series (ICCV 2023): This series is pre-trained based on Swin Transformer v2 and supports Tiny and Base versions for satellite image analysis.

[0091] FocalNet series (NIPS 2022): FocalNet is a model that focuses on the fusion of local and global features. It is divided into five versions: Tiny, Small, Base, Large, and XLarge, with the number of parameters ranging from 29.15M to 364.28M.

[0092] InternImage series (CVPR 2023): This is a model for large-scale visual tasks, divided into six versions: Tiny, Small, Base, Large, XLarge and Huge, with the number of parameters ranging from 28.77M to 1073.12M.

[0093] UniRepLKNet series (CVPR 2024): This is a model focused on unified representation learning, divided into five versions: Tiny, Small, Base, Large, and XLarge, suitable for multimodal image understanding tasks.

[0094] ScaleMAE series (ICCV 2023): The ScaleMAE model based on ViT supports use in the RSFM and timm frameworks, has 303.30M parameters, and is suitable for multi-scale visual tasks.

[0095] Cross ScaleMAE series (NIPS 2024): This model is also a multi-scale model based on ViT, supports application in RSFM and timm frameworks, and has 303.30M parameters.

[0096] SatMAE++ series (CVPR 2024): This is a further optimized ViT-based model that supports pre-training in RSFM and timm and is suitable for processing large-scale remote sensing data.

[0097] SpectralGPT (TPAMI 2024): This is a remote sensing image task model based on ViT, focusing on the extraction and processing of spectral information.

[0098] HyperSigma series (ArXiv 2024): This is a super-resolution model based on ViT, divided into two versions, Base and Large, with 219.15M and 678.44M parameters respectively, suitable for high-precision image processing tasks.

[0099] RSP series (TGRS 2022): RSP is a Swin Transformer-based model designed specifically for remote sensing tasks with 27.52M parameters.

[0100] GFM series (ICCV 2023): GFM is another variant based on Swin Transformer with 86.75M parameters and is suitable for complex image segmentation tasks.

[0101] Mamba series (ArXiv 2024): Mamba is a new model series consisting of three versions: Tiny, Small, and Base, with 29.94M, 49.38M, and 87.53M parameters respectively, for multi-task visual analysis.

[0102] Step S32: Multiple mainstream frameworks of compatible algorithms.

[0103] In the basic model algorithm library under the unified framework, considering that some algorithms are encoded in multiple mainstream frameworks, in order to make it easier for users to compare algorithm performance with one click, the basic model algorithm library of GRSFM supports and is compatible with running these algorithms; In the basic model algorithm library under the unified framework, the ScaleMAE, Cross ScaleMAE, and SatMAE++ series algorithms are compatible with running both the GRSFM and timm frameworks, namely ViT-Large-RSFM and ViT-Large-timm.

[0104] Step S4: Build a high-performance downstream task interpretation algorithm library under a unified framework.

[0105] The downstream task interpretation algorithm library includes basic visual tasks and visual language tasks, including scene classification (visual multimodal), semantic segmentation (visual multimodal), transformation detection (visual multimodal), reference image segmentation (visual language multimodal), visual localization (visual language multimodal), visual question answering (visual language multimodal), object detection (visual multimodal), instance segmentation (visual multimodal), rotated box object detection and other 8+ downstream tasks; The algorithm information specifically supported by the downstream task interpretation algorithm library is described as follows: In the semantic segmentation task, SegFormer (NIPS 2021) is a Transformer-based semantic segmentation model characterized by efficient and accurate feature extraction; UMixFormer (ArXiv 2023) is a Transformer-based semantic segmentation model that uses a hybrid approach for context aggregation to improve segmentation performance in complex scenarios; LightHam (ICLR 2021) introduces a lightweight attention mechanism to provide an efficient semantic segmentation method and reduce the amount of computation; UNet (MICCAI 2015) is a classic convolutional neural network architecture, particularly suitable for medical image processing, with a symmetrical U-shaped structure; UNetv2 (ArXiv 2023) is an improved version of UNet that further optimizes segmentation accuracy and computational efficiency; UNetFormer (ISPRS 2022) is a model that combines UNet and Transformer, using Transformer to improve the modeling capability of long-distance dependencies; UperNet (ECCV 2018) is a segmentation model based on pyramid feature extraction, which is good at processing objects in multi-scale scenes; Semantic FPN (CVPR 2019) uses feature pyramid network (FPN) for semantic segmentation, which is suitable for multi-level feature fusion.

[0106] In the change detection task, UperNet (ECCV 2018) is suitable for capturing scene changes; BIT-CD (TGRS 2021) is a change detection method based on a dual-branch architecture, which detects changes by comparing images at different time points; MambaBCD (TGRS 2024) is a model in the Mamba series, focusing on change detection tasks and providing higher detection accuracy.

[0107] In the referring image segmentation task, LAVT (CVPR 2022) is a Transformer-based reference image segmentation model that uses language to guide the segmentation of specified objects in the image; LGCE (TGRS 2024) is a segmentation model used in remote sensing tasks, which enhances the segmentation performance of specified objects through multi-task learning; RMSIN (TGRS 2024) is a model based on the contextual attention mechanism, which is particularly good at segmenting target objects from complex scenes.

[0108] In the Visual Grounding task, TransVG (ICCV 2021) is a Transformer-based visual positioning model that excels at combining visual information with language information to locate targets; VLTVG (CVPR 2022) is a model that combines vision and language, and improves positioning accuracy by processing multimodal data through Transformer; LQVG (TGRS 2024) is an optimized visual positioning model that focuses on object positioning in low-quality remote sensing images.

[0109] In the Visual Question Answering task, SOBA (AAAI 2024) Visual question answering systems that combine image and text information are good at handling multimodal problems; ALBEF (NIPS 2021) uses contrastive learning for joint visual-text representation to improve the accuracy of multimodal question answering.

[0110] In the object detection task, X-VLM (ICML 2021) is a multimodal visual language model that focuses on object detection tasks and improves detection performance by combining text and visual information; Faster RCNN (NIPS 2015) is a classic two-stage object detector that performs object detection through a region proposal network (RPN); RetinaNet (ICCV 2017) is a single-stage object detector that uses Focal Loss to deal with category imbalance problems and is widely used in practical applications; DINO (ICLR 2023) is a Transformer-based object detection method that uses self-supervised learning to improve feature representation.

[0111] In the instance segmentation task, Mask RCNN (ICCV 2017) added a segmentation branch based on Faster RCNN to perform pixel-level instance segmentation of the target; Cascade Mask RCNN (TPAMI 2019) improved the performance of Mask RCNN through multi-stage optimization, especially in high-quality segmentation tasks; HTC (CVPR 2019) is a multi-task joint learning model that can perform target detection and instance segmentation simultaneously, with excellent performance.

[0112] In the task of oriented object detection, FCOSR (RS 2023) is an improved version of FCOS, focusing on oriented target detection tasks, especially suitable for target detection in complex scenes; OrientedRCNN (ICCV 2021) incorporates the rotation information of the target into the detection framework, improving the detection accuracy of rotated targets; S²ANet (TGRS 2021) is an oriented target detection model designed for remote sensing images, which can accurately handle rotated targets in complex backgrounds.

[0113] Step S5: unified model training.

[0114] Based on the previous steps, before starting the training based on the unified framework, you need to complete the following initialization tasks: data loading, basic model loading, downstream task interpretation algorithm network loading, optimizer configuration, learning rate configuration, loss function configuration, etc. The unified model training process covers the evaluation of multiple tasks, including scene classification, semantic segmentation, transformation detection, reference image segmentation, visual positioning, visual question answering, object detection, instance segmentation, and rotation box object detection. The system will execute different training logics according to different tasks.

[0115] Step S51: Unify model training logic.

[0116] The training logic generally includes: optimizing model parameters by setting the training and validation steps for each epoch, and saving the best model checkpoint based on the evaluation indicators of the validation set (such as mIoU, Top1 Accuracy, etc.). The training process includes forward propagation, backpropagation, loss calculation and recording, as well as regular model evaluation and result recording. For some tasks, video memory management is also performed to ensure the efficiency and stability of the training process.

[0117] Step S52: Training start instruction.

[0118] The remote sensing buildable unified model framework GRSFM supports one-step script startup training. Taking into account the inconsistency of user computing power environment, the training startup script supports custom specified GPU, port and other operating environment configurations;

[0119] As in the above script,<num gpus> Indicates the number of GPUs you want to use, for example, 4 means using 4 GPUs. <port>Refers to the communication port number used by distributed training. Ensure that different training tasks use different ports to avoid conflicts. You can specify an idle port number, such as 10000.

[0120] Step S53: training startup logic.

[0121] Step S531: Generate time stamp.

[0122] Generate a timestamp of the current time in the format of YYYYMMDD_HHMMSS and store it in the variable now. This ensures that each training log file has a unique name to prevent overwriting previous logs.

[0123] Step S532: Define the task type.

[0124] Define a variable task to specify the type of task to be performed. "ris" represents the reference image segmentation task, "seg" represents the semantic segmentation task, "cd" represents the change detection task, and "cls" represents the classification task.

[0125] Step S533: Obtain the configuration file path.

[0126] Determine the configuration file path to use based on the value of the task variable. Assuming the configuration file is located in the configs / directory, the file name corresponds to the task type. For example, the ris task will use configs / ris.yaml as the configuration file.

[0127] Step S534: Obtain the training result saving path.

[0128] Define the save path of the training results, save_path, which is stored in the exps / directory according to the task type. This path indicates that the model name may be swin_tiny_LAVT, 4 GPUs per batch, image size 512x512, and training for 50 epochs.

[0129] Step S535: Create a directory to save training results and log files.

[0130] Create a directory to store training results and log files. If the path does not exist, an intermediate subdirectory will be automatically created by executing the mkdir -p command.

[0131] Step S536: Call the PyTorch distributed training script to start training.

[0132] Call the PyTorch distributed training script to start training. Set the training process by specifying different training parameters. --nproc_per_node=$1 indicates the number of processes started for each node. This parameter is passed in by the user when starting the script and is usually equal to the number of GPUs used. --master_addr=localhost: specifies the address of the master node. Here, it is local (localhost) because it is a single-node training. --master_port=$2 specifies the port of the master node. The port number is provided by the user when starting the script to ensure that multiple processes can communicate on the same node. train.py indicates the file name of the training script and executes the training logic. --config=$config indicates the configuration file for the specified training. Here, the config variable defined earlier is used. --save-path $save_path: indicates the path to save the specified training results and logs. --port$2 is used to specify the port number for communication between processes. 2>&1 | tee $save_path / $now.log redirects the standard output and error output to the log file and also displays it in the terminal.

[0133] Step S54: Setting training details.

[0134] Step S541: Cancel the encoding dimension setting in model_info.py.

[0135] Since each base algorithm model file specifies a specific value for embed_dim, there is no need to pass the embed_dim parameter from model_info.py to builder.py.

[0136] Step S542: Support automatic resetting of input channels and aligning and loading pre-trained weights.

[0137] In order to automatically align the number of input channels to adapt to the new data input format when using a pre-trained model. When the model's in_channels is greater than 3 (usually RGB channels), first extract the corresponding patch_embed1.proj.weight from the pre-trained weights. The first channel of the original weight is expanded through the _align_input_channel function, and then aligned one by one according to the new number of input channels, and the 3 channel information in the original weight is copied to the new weight channel. After the alignment is completed, load these weights into the model. If a pre-trained model is not specified, all weights of the model are initialized. If the input pretrained parameter is not a string or None, a type error will be raised.

[0138] Step S543: Use with_cp to relieve the pressure of video memory usage.

[0139] In order to effectively alleviate the pressure of video memory usage in deep models, set the flag with_cp to enable the checkpoint mechanism. In the forward function of the model, if with_cp is set to True and the input x requires gradient (x.requires_grad=True), the model will use torch.utils.checkpoint for forward calculation. In this way, the model does not store intermediate activation values, thereby saving memory, but these activation values ​​need to be recalculated during backpropagation. If with_cp is False, the model performs forward calculations according to the normal process without using checkpoints.

[0140] Step S544: Supporting selective setting of input channel dimensions for multi-spectral data.

[0141] The need_bands parameter is set in the configuration file to set the dimension of the multispectral input channels that need to be loaded. If need_bands is set, the required spectral bands are extracted from the original image and combined into a new input channel matrix. If only one band is selected, it is dimensionally compressed and normalized; if multiple bands are selected, the multiple bands are concatenated along the last dimension and then normalized. This process ensures that the number of channels of the input data is dynamically adjusted according to the required bands to adapt to the characteristics of different multispectral data.

[0142] Step S545: Support customizing the image size of the input model.

[0143] Considering that the default img_size=224 is not very rigorous when building the model input size, this framework supports passing the actual training size to the corresponding model image size from builder.py.

[0144] Step S546: The same training framework is compatible with multiple visual tasks, including change detection tasks.

[0145] In terms of data loading, this framework GRSFM distinguishes different visual tasks through the vision_task parameter in datasets_info. For semantic segmentation tasks, use the create_seg_dataset class to load a single image, and extract specific bands or directly load RGB images according to the configuration; for change detection tasks, use the create_cd_dataset class to load images at two time points, process them separately, and then compare them. Label loading also depends on the task type, reading mask or label images in the corresponding directory, so as to be compatible with the image loading and processing requirements of multiple tasks.

[0146] In terms of model loading, this framework GRSFM uses the engine_builder function to distinguish different visual tasks (such as semantic segmentation and change detection) according to the vision_task parameter in datasets_info, and calls the train_seg and train_cd functions for training respectively. The training logic corresponding to each task is defined independently, but the shared components in the framework such as models, optimizers, loss functions, etc. remain consistent, so that the same framework can be compatible with multiple visual tasks, and corresponding data loading and model evaluation are performed according to the task type.

[0147] In terms of loss function and optimizer configuration, this framework GRSFM dynamically creates loss functions and optimizers suitable for the task based on the task type and related parameters in the configuration file. For example, the loss_builder function selects CrossEntropyLoss or other loss functions based on the configuration, while the optimizer_builder function selects SGD, AdamW or Adam optimizer based on the configuration. In this way, the framework can adjust the training strategy according to specific task requirements, thereby supporting a variety of visual tasks.

[0148] Step S547: Support basic and multiple basic variants (such as LAVT, LGCE, RMSIN) as visual downstream task models.

[0149] The design of the basic downstream task model is flexible, and each feature extraction stage can select the corresponding fusion method according to different variants. These variants (such as LAVT, LGCE, RMSIN) are implemented as modular components (such as fusion modules) and can be dynamically loaded into different stages of the model. By judging the vlf_ris parameter in the model configuration, the framework inserts different fusion mechanisms (such as LAVT_fusion, LGCE_fusion or RMSIN_fusion) at specific locations to achieve different feature fusion methods. Throughout the process, the model supports both the standard basic process and advanced variants as required by the task. At the same time, RMSIN also uses a specific CIM module in the output stage to further fuse the output features. Through this flexible design, the framework can automatically select and load suitable fusion variants according to the different requirements of the task, and support a variety of visual tasks.

[0150] Step S548: Supporting the use of a unified weight naming rule to store the pre-trained weights that need to be saved.

[0151] During the training process, each time a weight file is saved, a uniformly named file name is generated through the get_save_weight_name function. The file name consists of the key information of the model, including the dataset name, the basic model Backbone type, the Neck structure type, the decoder type, the number of GPUs used, the batch size, the image cropping size, the total number of training rounds, the source of the pre-trained model (such as imagenet), and the model's mIoU score on the validation set. The naming template is "dataset.backbone.neck.decoder.batch.imgsize.epoch.pretrain.miou.pth". The naming format is standardized and unified, which facilitates subsequent model management and version control. In addition, the framework will automatically call the delete_previous_best_weight function to delete the previously stored optimal weights before saving the latest optimal model weights to ensure efficient storage space management.

[0152] Step S6: Unified model verification.

[0153] The unified model verification process covers the evaluation of multiple tasks, including scene classification, semantic segmentation, transformation detection, reference image segmentation, visual localization, visual question answering, object detection, instance segmentation, and rotated box object detection. In the evaluation process of each task, the model is in evaluation mode, inferring the input data and calculating relevant evaluation indicators such as mIOU, mF1, OA, mAcc1, mAcc5, oIoU, etc. In distributed training, the evaluation results of each computing node are synchronized through all_reduce to ensure consistency. Finally, all calculated evaluation indicators are summarized and returned to evaluate the performance of the model.

[0154] Step S61: Verify the start instruction.

[0155] The remote sensing buildability unified model framework GRSFM supports one-step script startup verification. The verification startup script takes into account the inconsistency of the user's computing power environment, supports custom specified GPU, port and other operating environment configurations, supports custom loading of model weights, and can also improve the model effect through "enhanced TTA during testing".

[0156]

[0157]

[0158] As in the above script,<num gpus> Indicates the number of GPUs you want to use, for example, 4 means using 4 GPUs. <port>Refers to the communication port number used by distributed training. Ensure that different training tasks use different ports to avoid conflicts. You can specify an idle port number, such as 10000. Indicates the path where weights need to be loaded. --tta indicates adding a "test-time enhancement" module during inference.

[0159] Step S62: Enhance the module during testing.

[0160] The prediction effect of the model is improved through multi-scale and left-right flipping data enhancement strategies. Specifically, the input image x is first interpolated and adjusted according to different scaling ratios, and then forward propagated through the model's base_forward, and the obtained prediction results are processed using softmax. The results of each scaling are restored to the original size and gradually accumulated to the final result. To further enhance the effect, the input image is also horizontally flipped, and then through the same process, all enhanced results are finally added to obtain the average final prediction. This method helps to improve the robustness of the model under different input scales and directions.

[0161] In another embodiment of the present invention, an electronic device is provided, the electronic device comprising a processor and a memory, the memory being used to store a computer program, the computer program comprising program instructions, and the processor being used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of the general remote sensing data interpretation method.

[0162] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device, and can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that more specific examples (non-exhaustive list) of the computer-readable storage medium here include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0163] Computer readable storage media also include data signals propagated in baseband or as part of a carrier wave, which carry readable program codes. Such propagated data signals can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0164] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0165] The processor may load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the general remote sensing data interpretation method in the above embodiment.

[0166] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / port> < / port>

Claims

1. A universal remote sensing data interpretation system, characterized in that: include: Remote sensing data module, used to store remote sensing data sets; Basic model module, used to integrate basic models; An interpretation module, used to integrate interpretation algorithms; A unified training module, used to train the basic model according to the training instructions; the training instructions at least include loading one or more combinations of remote sensing data sets, loading one or more combinations of basic models, and loading one or more combinations of interpretation algorithms; The model verification module is used to verify and evaluate the basic model after each training.

2. The universal remote sensing data interpretation system according to claim 1, characterized in that: The training instructions also include optimizer configuration, learning rate configuration and loss function configuration.

3. The universal remote sensing data interpretation system according to claim 1, characterized in that: The basic models include a basic model based on a CNN architecture, a basic model based on a Transformer architecture, and a basic model based on a Mamba architecture.

4. The universal remote sensing data interpretation system according to claim 1, characterized in that: The training instructions are configured through a YAML file.

5. The universal remote sensing data interpretation system according to claim 1, characterized in that: The unified training module also performs video memory management.

6. The universal remote sensing data interpretation system according to claim 1, characterized in that: The interpretation algorithms include scene classification, semantic segmentation, change detection, reference image segmentation, visual localization, reference expression understanding and segmentation, visual question answering, object detection, instance segmentation and rotated box object detection.

7. A universal remote sensing data interpretation method, based on the universal remote sensing data interpretation system according to any one of claims 1 to 6, characterized in that: The following steps are involved: Configure training instructions; Based on the training instructions, the basic model is trained; After each training session, each interpretation algorithm was evaluated.

8. The universal remote sensing data interpretation method according to claim 7, characterized in that: Each interpretation algorithm is evaluated, and its evaluation indicators include mIOU, mF1, OA, mAcc1, mAcc5, oIoU, PR@.5, PR@.6, PR@.7, PR@.8 and PR@.

9.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the universal remote sensing data interpretation method according to any one of claims 7 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the universal remote sensing data interpretation method according to any one of claims 7 to 8 is implemented.