Map-guided remote sensing image interpretation methods, systems, terminals, and media

By generating graph structures and training visual language models, the problem of missing spatial relationships in remote sensing image interpretation is solved, improving the accuracy and effectiveness of interpretation.

CN121482614BActive Publication Date: 2026-04-21ZHEJIANG INST OF SURVEYING & MAPPING SCI & TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG INST OF SURVEYING & MAPPING SCI & TECH
Filing Date
2026-01-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing remote sensing image interpretation methods lack the ability to extract and utilize the spatial relationships between ground features, resulting in inaccurate interpretation results that fail to accurately reflect the surface condition.

Method used

By generating graph structures, the visual language model is trained based on the spatial and semantic information of the graph structures, enabling it to learn to capture the spatial distribution patterns of ground features and utilize expert prior knowledge to assist in interpretation.

Benefits of technology

It improves the accuracy of remote sensing image interpretation, avoids inaccurate interpretation caused by missing spatial relationships, and achieves better interpretation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482614B_ABST
    Figure CN121482614B_ABST
Patent Text Reader

Abstract

This application provides a graph structure-guided remote sensing image interpretation method, system, terminal, and medium. The remote sensing image interpretation method includes: interpreting the acquired remote sensing images using a visual language model based on acquired remote sensing imagery and expert prior knowledge; the model training process includes: obtaining graph structures based on vector graphics corresponding to the remote sensing images, rapidly synthesizing a large number of samples for training by randomly modifying the graph structures; training the model based on the synthesized training remote sensing images and their corresponding training graph structures and vector graphics, using the graph structure features obtained from each training graph structure as training prior knowledge, interpreting a large number of image features obtained from each training remote sensing image, and using each training vector graphics as labels for feedback optimization. The visual language model used in this application can acquire spatial and semantic information from remote sensing images, avoiding inaccurate interpretation caused by missing spatial relationships, and has a better interpretation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing and relates to a remote sensing image interpretation technology, particularly a graph structure-guided remote sensing image interpretation method, system, terminal, and medium. Background Technology

[0002] Remote sensing image interpretation refers to the process of identifying, classifying, and delineating features from remote sensing images to accurately obtain information such as the type and spatial distribution of features. As an important means of remote sensing image data processing, remote sensing image interpretation can efficiently and over a large area obtain information about the Earth's surface conditions, and has wide applications in environmental monitoring, disaster assessment, and urban planning.

[0003] Traditional remote sensing image interpretation typically relies on manual intervention, meaning that the interpretation process is manually mediated. This approach is not only inefficient, but its accuracy is also limited by the operator's experience. In recent years, deep learning technology has been widely applied to remote sensing image interpretation. Specifically, convolutional neural networks (CNNs) or Transformers are used to extract features from preprocessed remote sensing images, capturing spectral, texture, and shape characteristics of ground features, and then performing interpretation based on these features. However, this method lacks the ability to extract and utilize spatial relationships between ground features, and the model does not incorporate spatial distribution patterns into the interpretation process. This can lead to interpretation results that deviate from actual ground feature patterns, resulting in missing spatial context information or topological errors. Meanwhile, using deep learning models for interpretation usually requires a large number of samples to train the interpretation model in order to improve the accuracy of the interpretation results. However, in practice, it is difficult to obtain remote sensing images for training. Synthesizing new remote sensing images from existing remote sensing images is inefficient and the accuracy of the synthesized remote sensing images is also low, resulting in poor training effects. This leads to low interpretation accuracy of the interpretation model and further affects the interpretation accuracy of remote sensing images. Summary of the Invention

[0004] The purpose of this application is to provide a graph-structure-guided remote sensing image interpretation method, system, terminal, and medium to solve the problems in the existing technology of remote sensing image interpretation that lacks the extraction and utilization of spatial relationships, and the model does not incorporate the spatial distribution patterns of ground features during the interpretation process, resulting in inaccurate interpretation results.

[0005] In a first aspect, this application provides a graph-structure-guided remote sensing image interpretation method, comprising: acquiring remote sensing images and pre-defined expert prior knowledge; inputting a trained visual language model based on the remote sensing images and the expert prior knowledge to understand and recognize the remote sensing images, thereby interpreting the remote sensing images; wherein, the training process of the visual language model includes: acquiring at least one remote sensing image and corresponding vector graphics for each remote sensing image; synthesizing and expanding based on each remote sensing image and corresponding vector graphics to obtain an expanded sample library; the expanded sample library includes... The method includes several training remote sensing images and corresponding training map structures and training vector graphics. Based on each training remote sensing image, corresponding training image features are obtained, and based on each training map structure, training map structure features are obtained. The training map structure features are used as training prior knowledge, and the training prior knowledge and training image features are input into a visual language model to understand and recognize the training image features. The training vector graphics are used as training labels to provide feedback on the accuracy of the interpretation results of the training image features, thereby training the visual language model.

[0006] In one embodiment of this application, each of the training prior knowledge is individual prior knowledge or combined prior knowledge; the step of inputting each of the training prior knowledge and each of the training image features into a visual language model to understand and recognize each of the training image features includes: obtaining corresponding node types and individual semantic features based on each of the individual prior knowledge; guiding fusion through attribute-aware attention based on each of the training image features, each of the node types, and each of the individual semantic features to perform semantic analysis on each of the training image features and obtain semantically fused image features; obtaining corresponding spatial features and semantic features of each node based on the combined prior knowledge; calculating cosine similarity based on each of the semantically fused image features and each of the node semantic features to ensure that each of the node semantic features corresponds one-to-one with each of the semantically fused image features; and clustering each of the semantically fused images based on the spatial features to perform spatial analysis on each of the semantically fused image features.

[0007] In one embodiment of this application, each of the training prior knowledge is individual prior knowledge or combined prior knowledge; the step of inputting each of the training prior knowledge and each of the training image features into a visual language model to understand and recognize each of the training image features includes: obtaining corresponding individual semantic features based on each of the individual prior knowledge, and visually projecting each of the individual semantic features to obtain corresponding visual markers; based on each of the training image features, combined with each of the visual markers, using a shared attention mechanism to perform semantic analysis on the influencing features to obtain each of the semantic fusion image features; performing self-attention calculation based on each of the semantic fusion image features to obtain each of the enhanced remote sensing features, and based on the combined prior knowledge, obtaining corresponding spatial features and semantic features of each node; and performing multi-head cross-attention calculation based on each of the enhanced remote sensing features, combined with the spatial features and semantic features of each node, to inject spatial relationships into each of the enhanced remote sensing features and perform spatial analysis on each of the semantic fusion image features.

[0008] In one embodiment of this application, the synthesis and expansion of any of the remote sensing images and their corresponding vector images includes: obtaining the corresponding graph structure based on the vector image; obtaining corresponding synthesized graph structures by random modification based on the graph structure; obtaining corresponding synthesized vector images based on each synthesized graph structure; extracting corresponding remote sensing features based on the remote sensing image; extracting corresponding vector features based on the vector image; obtaining features of each synthesized graph structure based on the spatial and semantic information of each synthesized graph structure through a geometric semantic perception module; for each synthesized graph structure, obtaining corresponding synthesized remote sensing images based on the corresponding features of each synthesized graph structure, combined with the remote sensing features and the vector features; using both the remote sensing image and each synthesized remote sensing image as training remote sensing images, and obtaining the graph structures or synthesized graph structures corresponding to each training remote sensing image as corresponding training graph structures, and obtaining the vector images and synthesized vector images corresponding to each training remote sensing image as corresponding training vector images, to construct the expanded sample library.

[0009] In one embodiment of this application, for any of the synthetic map structures, based on the corresponding synthetic map structure features, combined with the remote sensing features and the vector features, to obtain the corresponding synthetic remote sensing images includes: normalizing and aligning the synthetic map structure features, the remote sensing features, and the vector features; enhancing the normalized and aligned synthetic map structure features, the remote sensing features, and the vector features through self-attention calculation, and performing feature fusion through multi-head cross-attention calculation to obtain each synthetic remote sensing feature; and generating each synthetic remote sensing image corresponding to each synthetic map structure feature based on each synthetic remote sensing feature.

[0010] In one embodiment of this application, the step of obtaining graph structure features based on the spatial and semantic information of the graph structure through a geometric semantic perception module includes: extracting the semantic information of the graph structure into semantic features through a preset text encoder; mapping based on the spatial information of the graph structure to obtain each spatial feature; and obtaining the correlation between each semantic feature and each spatial feature through an attention mechanism based on a preset graph neural network to generate the graph structure features.

[0011] In one embodiment of this application, obtaining the corresponding graph structure based on the vector map includes: generating corresponding nodes based on each geographic feature of the vector map; obtaining the bounding box of each node based on the geographic coverage of each geographic feature; obtaining the connection relationship of each node based on the geospatial relationship of each geographic feature; obtaining the spatial information of each node based on the bounding box and connection relationship of each node; obtaining the text description of each node based on the geographic feature information of each geographic feature, as the semantic information of each node; and obtaining the graph structure based on each node and its spatial and semantic information.

[0012] Secondly, this application provides a graph-structure-guided remote sensing image interpretation system, including a data acquisition module and a visual semantic model; the data acquisition module is used to acquire remote sensing images and preset expert prior knowledge; the visual language module includes a trained visual language model, used to understand and recognize the remote sensing images based on the remote sensing images and the expert prior knowledge, thereby interpreting the remote sensing images; wherein, the training process of the visual language model includes: acquiring at least one remote sensing image and each vector graphic corresponding to each remote sensing image; Based on the remote sensing images and their corresponding vector graphics, an expanded sample library is obtained through synthesis and expansion. The expanded sample library includes several training remote sensing images and their corresponding training image structures and vector graphics. Based on each training remote sensing image, corresponding training image features are obtained, and based on each training image structure, training image structure features are obtained. The training image structure features are used as prior knowledge for training, and the prior knowledge and training image features are input into a visual language model to understand and recognize the training image features. The training vector graphics are used as training labels to provide feedback on the accuracy of the interpretation results of the training image features, thereby training the visual language model.

[0013] Thirdly, this application provides a terminal, including: a processor and a memory, wherein the memory and the processor are communicatively connected; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal performs the graph-guided remote sensing image interpretation method as described above.

[0014] Fourthly, this application provides a computer storage medium storing a computer program that, when executed by a processor, implements the graph-guided remote sensing image interpretation method as described above.

[0015] As described above, this application provides a graph structure-guided remote sensing image interpretation method, system, terminal, and medium. By generating a graph structure based on vector graphics, it facilitates the acquisition of spatial and semantic information, and injects training prior knowledge into image features to guide the interpretation of each image feature, thereby training a visual language model that can learn to capture the spatial distribution patterns of ground objects. This effectively improves the interpretation accuracy of the trained visual language model, thus contributing to better interpretation results. Attached Figure Description

[0016] Figure 1 The diagram shown is a flowchart illustrating a remote sensing image interpretation method according to an embodiment of this application.

[0017] Figure 2 The diagram shown is a flowchart illustrating the training process of a visual language model as described in an embodiment of this application.

[0018] Figure 3 The diagram shown is a flowchart illustrating a vector graphic to graph structure as described in an embodiment of this application.

[0019] Figure 4 The diagram shown is a flowchart illustrating an expanded sample library acquisition method as described in an embodiment of this application.

[0020] Figure 5 The diagram shown is a flowchart illustrating a graph structure feature acquisition method according to an embodiment of this application.

[0021] Figure 6 The diagram shown is a flowchart illustrating a synthetic remote sensing image generation method as described in an embodiment of this application.

[0022] Figure 7 The diagram shown is a flowchart illustrating a synthetic remote sensing image interpretation method as described in an embodiment of this application.

[0023] Figure 8 The diagram shows a flowchart illustrating another synthetic remote sensing image interpretation method described in this application embodiment.

[0024] Figure 9 The diagram shown is a structural schematic of a remote sensing image interpretation system according to an embodiment of this application.

[0025] Figure 10 The diagram shown is a structural schematic of a terminal as described in an embodiment of this application.

[0026] Explanation of reference numerals in the attached figures

[0027] 41: Data acquisition module; 42: Visual language module; 421: Visual language model; 50: Terminal; 51: Processor; 52: Memory; 521: Operating system; 522: Application program; 53: User interface; 54: Network interface; 55: Bus system. Detailed Implementation

[0028] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0029] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0030] Existing remote sensing image interpretation methods often lack the extraction and utilization of spatial relationships between ground features, resulting in missing spatial context information or topological errors in the interpretation results. Consequently, the interpretation results of remote sensing images are inaccurate, fail to accurately reflect the surface conditions, and are difficult to meet practical needs.

[0031] To address the technical problems existing in the prior art, the following embodiments of this application provide a graph structure-guided remote sensing image interpretation method, system, terminal, and medium. By generating a graph structure, a visual language model is trained based on the spatial and semantic information of the graph structure. This enables the visual language model to learn and capture the spatial distribution patterns in the remote sensing image, avoiding inaccurate interpretation caused by missing spatial relationships. This approach helps to achieve better interpretation results for remote sensing images.

[0032] The technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0033] like Figure 1 As shown, this embodiment provides a graph-guided remote sensing image interpretation method, including:

[0034] S10 acquires remote sensing images and pre-defined expert prior knowledge.

[0035] In this context, expert prior knowledge is used to represent guiding information when interpreting the remote sensing image. That is, when interpreting a remote sensing image, corresponding expert prior knowledge is used to assist in the understanding, so as to improve the accuracy of the interpretation.

[0036] Furthermore, expert prior knowledge can be categorized into single-type and combined-type. Single-type expert prior knowledge involves only a single geographic element and is typically used to assist in understanding that geographic element, such as supplementing semantic information, defining spatial location or spatial extent, etc. Combined-type expert prior knowledge involves multiple geographic elements and is typically used to understand the whole composed of these geographic elements, such as supplementing semantic information of the whole composed of these geographic elements, or supplementing spatial relationships between these geographic elements, etc.

[0037] For example, if the expert prior knowledge is set as follows: for a certain area of ​​wasteland in a remote sensing image, it is expected to be reclaimed into arable land next year, then the remote sensing image can be interpreted based on this expert prior knowledge to mark the changed arable land area in the remote sensing image and obtain the rate of change of arable land area to show the predicted arable land situation next year; or, if the expert prior knowledge is set as follows: parks usually include vegetation, water bodies and some functional buildings, then the remote sensing image can be interpreted based on this expert prior knowledge to mark the location and extent of the park in the remote sensing image, as well as the location and extent of various geographical elements in the park.

[0038] It should be noted that in this embodiment, the preset expert prior knowledge is information input or set by technicians to guide the interpretation of remote sensing images. Those skilled in the art can set the expert prior knowledge specifically for the content of interest in remote sensing image interpretation based on actual circumstances; this embodiment does not impose specific limitations.

[0039] S20, based on remote sensing imagery and expert prior knowledge, inputs a trained visual language model to understand and recognize remote sensing imagery, thereby realizing the interpretation of remote sensing imagery.

[0040] The visual language model is used to understand and analyze images to obtain digital information, textual information, or to annotate images. For example, the visual language model framework adopts a VLM model and employs a multilayer perceptron output mechanism to output the final interpretation result. Furthermore, the visual language model is trained to improve its ability to acquire semantic and spatial information from images, thereby enhancing its interpretation accuracy for interpreting synthetic remote sensing imagery.

[0041] Specifically, remote sensing images and expert prior knowledge are input into a visual language model so that the visual language model can interpret the remote sensing images, identify various geographic elements in the remote sensing images, and annotate the information of interest based on expert prior knowledge.

[0042] It should be noted that during the training process of the visual language model, prior knowledge is acquired based on the spatial and semantic information of the graph structure. This prior knowledge is then used to train the visual language model, enabling it to capture and learn the spatial distribution patterns among various geographic elements. The resulting visual language model can interpret remote sensing images based on the spatial distribution patterns of various geographic elements, avoiding missing spatial context information or topological errors, and effectively improving the accuracy of the interpretation results.

[0043] To facilitate understanding, the steps and principles of the training process for visual language models will be explained in detail below.

[0044] In some alternative implementations, such as Figure 2 As shown, the training process of a visual language model includes:

[0045] S100: Acquire at least one remote sensing image and the corresponding vector graphics for each remote sensing image.

[0046] In this context, the vector map represents the same geographical area and the same acquisition time as the remote sensing image. For example, the interpreted remote sensing image is vectorized to obtain the corresponding vector map. Specifically, those skilled in the art should know the methods for obtaining vector maps, which will not be specifically explained in this embodiment. It should be noted that the vector map contains geographic element information, namely the spatial and semantic information of each geographic element, including but not limited to the geographic coverage area and other geographic element information. Obtaining the vector map facilitates the acquisition of its spatial and semantic information, thereby assisting the training model in understanding the spatial and semantic information of the remote sensing image. Furthermore, the vector map corresponding to each remote sensing image reflects the accuracy of the model's interpretation of that remote sensing image, enabling the visual language model to be optimized and adjusted, thus achieving the training of the visual language model.

[0047] S200, based on each remote sensing image and its corresponding vector image, synthesizes and expands to obtain an expanded sample library.

[0048] The expanded sample library includes remote sensing imagery, vector graphics, and graph structures for training the visual language model. Remote sensing imagery serves as the interpretation object during model training; vector graphics are used to evaluate the accuracy of model interpretation during training, providing feedback for model optimization; and graph structures serve as prior knowledge for training, enabling the model to acquire spatial and semantic information of various geographic elements and learn the spatial distribution patterns among these elements. It should be noted that although vector graphics also contain spatial and semantic information of various geographic elements, for models interpreting remote sensing imagery, the features extracted from vector graphics are actually image features and cannot understand the spatial and semantic information of the vector graphics. Therefore, vector graphics are converted into graph structures to facilitate model understanding.

[0049] Therefore, it is necessary to obtain the corresponding graph structure based on the vector diagram.

[0050] In some alternative implementations, such as Figure 3 As shown, the ways to convert vector graphics into graph structures include:

[0051] S211, based on the geographic elements in the vector map, generates the corresponding nodes.

[0052] Each geographic element is used to represent a geographic entity in the actual geographic space, such as a school, a school building, a road, a river, etc.

[0053] Specifically, for each geographic feature, a node is generated, which, for example, is set at the geometric center of the corresponding geographic feature to reflect the spatial location of each geographic feature.

[0054] Optionally, users can manually add or modify nodes to meet actual interpretation needs or correct geographic features with errors. For example, users can manually add or modify geographic features in a vector map and use the modified geographic features as nodes for subsequent graph structure generation. For example, for land planned for a school, a "school" geographic feature is added to the corresponding area; or for hillside land planned for reclamation, the corresponding area is changed from "hillside land" to "farmland." It should be noted that those skilled in the art can consider whether to manually add or modify geographic features based on actual needs; this embodiment does not impose specific limitations.

[0055] S212: Based on the geographic coverage of each geographic element, obtain the bounding box of each node; based on the geospatial relationship of each geographic element, obtain the connection relationship of each node; based on the bounding box and connection relationship of each node, obtain the spatial information of each node.

[0056] Specifically, the actual area of ​​each geographic entity in the actual geographic space is taken as the geographic coverage of the corresponding geographic element, and the corresponding bounding box is extracted based on the boundary of the geographic coverage on the vector map, which is taken as the bounding box of the corresponding node. The spatial relationship of each geographic entity in the actual geographic space, such as the adjacency relationship between schools and residential buildings, the inclusion relationship between schools and teaching buildings, etc., is taken as the connection relationship between each node. Optionally, for two nodes with a connection relationship, an edge is established between the two nodes, that is, there is a line between the two nodes, in order to reflect the connection relationship between these nodes.

[0057] Furthermore, based on the bounding boxes of each node, i.e. the spatial extent of each node, and the connection relationships between each node, the spatial information of each node is obtained.

[0058] S213. Based on the geographic element information of each geographic element, obtain the text description of each node as the semantic information of each node.

[0059] Specifically, each geographic element in the vector map contains geographic element information such as annotations, including but not limited to geographic entity type, vector map acquisition time, and sensor type used for vector map acquisition. The geographic element information corresponding to each geographic element is used as the text description of the corresponding node, so as to obtain the semantic information of each node based on the text description of each node.

[0060] Optionally, the semantic information of each node can be expanded. Specifically, the user manually adds geographic feature information to the geographic features in the vector map to increase the information that the geographic feature needs to focus on, and obtains a text description based on the added geographic feature information as the semantic information of the corresponding node. It should be noted that those skilled in the art can consider whether to manually add geographic feature information based on actual needs, and this embodiment does not impose specific limitations here.

[0061] S214. Based on each node and its spatial and semantic information, obtain the corresponding graph structure.

[0062] Specifically, a graph structure is actually a collection of nodes, the spatial extent of each node, text descriptions, and the connections between nodes.

[0063] Based on this, vector graphics are converted into graph structures to facilitate the acquisition of spatial and semantic information, thereby assisting in the interpretation of remote sensing images and improving the accuracy of the interpretation results.

[0064] It should be noted that to obtain a more accurate visual language model, a large amount of remote sensing imagery is usually required for training. Therefore, this embodiment expands the training samples to obtain an expanded sample library for training the visual language model, thereby achieving better interpretation results.

[0065] In some alternative implementations, such as Figure 4 As shown, for any remote sensing image and its corresponding vector image, several composite image structures, composite remote sensing images, and composite vector images are obtained to expand the sample. Specifically, the methods for obtaining the expanded sample library for any remote sensing image and its corresponding vector image include:

[0066] S221, Based on the vector map, obtain the corresponding graph structure; based on the graph structure, obtain the corresponding composite graph structure by random modification; and based on each of the composite graph structures, obtain the corresponding composite vector map.

[0067] Specifically, the original graph structure is obtained based on the original vector image, and then randomly modified to generate new graph structures as each synthetic graph structure, thereby obtaining more graph structure samples and expanding the number of samples used for model training. The specific method and principle of obtaining the corresponding graph structure based on the vector image can be referred to the content of steps S211 to S214 above, and will not be repeated here in this embodiment.

[0068] Random modifications are made to the graph structure, including but not limited to adding or removing nodes, changing the size of bounding boxes, or modifying semantic information, thereby altering the geographic information of the graph structure. Synthetic graph structures, synthetic remote sensing images, and synthetic vector maps with changed geographic information are then synthesized. The training samples of the model are then expanded based on the changed synthetic graph structures, synthetic remote sensing images, and synthetic vector maps.

[0069] Based on this, this embodiment can quickly obtain a large number of synthetic image structure samples by modifying the image structures obtained from each vector image. Since each synthetic image structure possesses relatively complete semantic and spatial information, synthetic vector images can be quickly generated based on these structures, thereby rapidly expanding the sample library. This method is highly efficient and produces high-quality synthetic image structures and vector images. Furthermore, based on each synthetic image structure, synthetic remote sensing images are obtained in subsequent steps. Compared to existing technologies that expand remote sensing images using generative adversarial networks, this method not only has higher generation efficiency but also produces higher-precision, higher-quality remote sensing images due to the relatively complete semantic and spatial information of the synthetic image structures. This facilitates better training results and improves the interpretation accuracy of the visual language model.

[0070] S222: Extract the corresponding remote sensing features based on remote sensing images; extract the corresponding vector features based on vector images; and obtain the structural features of each composite image through the geometric semantic perception module based on the spatial and semantic information of each composite image structure.

[0071] In this context, each remote sensing feature is used to characterize a geographic element extracted from a remote sensing image; each vector feature is used to characterize a geographic element extracted from a vector image. For example, a preset autoencoder, such as a VAE Encoder, is used to extract each remote sensing feature from the remote sensing image; and a preset semantic encoder, such as a Semantic Encoder, is used to extract each vector feature from the vector image.

[0072] Based on this, information such as spectrum and texture is obtained from various remote sensing features of the original remote sensing image, and information such as the outline of ground features is obtained from various vector features of the original vector image, in order to assist in the generation of synthetic remote sensing images, effectively reduce the distortion of synthetic remote sensing images, and improve the quality of synthetic remote sensing images.

[0073] Furthermore, through the geometric semantic perception module, the structural features of each composite image are obtained. Based on the structure of each composite image, spatial and semantic information can be obtained, so as to combine each remote sensing feature and each vector feature to generate a composite remote sensing image.

[0074] In some alternative implementations, such as Figure 5 As shown, the methods for obtaining the structural features of each graph include:

[0075] S2221, through a preset text encoder, extracts the semantic information of the graph structure into various semantic features.

[0076] Specifically, a pre-defined text encoder, such as CLIP (Contrastive Language-Image Pre-Training) encoder, is used to take the semantic information of each node in the graph structure, that is, the text description of each node, as the semantic features.

[0077] S2222, based on the spatial information of the graph structure, maps to obtain various spatial features.

[0078] For example, a Box Embedding encoder is used to map the spatial information of the graph structure, i.e., the spatial extent of each node and the connection relationships between nodes, to obtain spatial features. It should be noted that each spatial feature is either a single spatial feature representing the spatial structure of a single node, or a combined spatial feature representing the connection relationships between multiple nodes.

[0079] It should be noted that the CLIP encoder and the Box Embedding encoder are combined to form the GSAM encoder (Geometric Semantic Aware Encoder), which is used to output graph structure features.

[0080] S2223, based on a pre-defined graph neural network, obtains the correlation between each semantic feature and each spatial feature through an attention mechanism to generate each graph structure feature.

[0081] Specifically, a pre-defined graph neural network, such as GraphFormer, is used to leverage its attention mechanism to perform deep fusion based on the correlation between semantic features and spatial features, in order to form the structural features of each graph.

[0082] For example, each semantic feature is aligned and fused with the spatial features of the corresponding node to obtain the corresponding graph structure features. Further, each graph structure feature is a single graph structure feature that represents the spatial structure and semantic information of a single node, or a combined graph structure feature that integrates the semantic features of multiple nodes and includes the connection relationships between multiple nodes.

[0083] S223, For each composite image structure, based on the corresponding composite image structure features, combined with each remote sensing feature and each vector feature, obtain the corresponding composite remote sensing image.

[0084] Specifically, for each composite image structure, its corresponding composite image structure features are obtained, and combined with each remote sensing feature and each vector feature, each composite remote sensing image is generated.

[0085] Based on the structural features, remote sensing features, and vector features of the composite image, a composite remote sensing image is generated to enhance the semantic and spatial information of the generated image, thereby improving the quality of the composite remote sensing image.

[0086] For example, the synthetic graph structure features, remote sensing features, and vector features are input into a trained image synthesis model to generate a synthetic remote sensing image. Specifically, the image synthesis model includes a GDIT (graph difussion image transformer) module connected to the encoder, a normalization layer, and a linear projection layer, which serve as the decoder. The encoder includes, as previously described, an encoder corresponding to each remote sensing feature (VAE encoder), an encoder corresponding to each vector feature (semantic encoder), and a GSAM encoder that outputs the synthetic graph structure features. Based on the extracted remote sensing features and vector features, feature alignment is achieved through the GDIT module. The aligned remote sensing features, vector features, and synthetic graph structure features are then normalized through the normalization layer. After normalization, the normalized features are input into the linear projection layer for dimensionality reduction through linear and reprojection, ultimately outputting the synthetic remote sensing image. Compared to the traditional method of synthesizing images using the DIT (Diffusion Image Transformer) model, the image synthesis model in this embodiment can achieve semantic space alignment between remote sensing features and vector features, enhancing spatial consistency and improving the quality of synthesized remote sensing images.

[0087] Furthermore, the GDIT module performs feature alignment as follows: each remote sensing feature is cropped into small blocks as a unit, these units are scaled and translated, and then input into a multi-head attention mechanism module to obtain effective features. The obtained features are then scaled again and combined with various vector features. The attention mechanism module then filters out effective features to achieve alignment between the vector features and the remote sensing features. Based on this, the GDIT module is constructed to enhance spatial consistency.

[0088] Furthermore, during the training process of the image synthesis model, it can be trained using a large number of graph structure features, remote sensing features, and vector features. The model parameters are updated based on the noise prediction loss and normalized to obtain the trained model. It should be noted that those skilled in the art should know the specific training methods, steps, and principles of the image synthesis model. This embodiment will not elaborate on them here. For the methods of obtaining each graph structure feature, each remote sensing feature, and each vector feature, please refer to the aforementioned methods of obtaining graph structure features, remote sensing features, and vector features. This embodiment will not repeat them here.

[0089] In some alternative implementations, such as Figure 6 As shown, the methods for generating synthetic remote sensing images include:

[0090] S2231, normalize and align the structural features, remote sensing features and vector features of each composite image.

[0091] Specifically, the GDIT module is used to align the various remote sensing features and vector features, and adaptive normalization techniques are employed, such as AdaLN-Single, to normalize the structural features, remote sensing features, and vector features of the composite image. Since the structural features, remote sensing features, and vector features of the composite image have different modalities and significant differences between them, normalization and alignment of these features are performed to enable the image synthesis model to understand different modalities of regions or geographic elements, thereby improving the accuracy and quality of the generated composite remote sensing image.

[0092] S2232 enhances the structural features, remote sensing features, and vector features of each synthesized image after normalization and alignment through self-attention calculation, and performs feature fusion through multi-head cross-attention calculation to obtain each synthesized remote sensing feature.

[0093] Specifically, the image synthesis model includes a self-attention layer and a multi-head cross-attention layer. Normalized and aligned structural features, remote sensing features, and vector features of each synthesized image are input into the image synthesis model. The self-attention layer enhances the structural features, remote sensing features, and vector features of each synthesized image, and the multi-head cross-attention layer performs feature fusion to obtain each synthesized remote sensing feature.

[0094] S2233, based on each synthetic remote sensing feature, generates each synthetic remote sensing image corresponding to each synthetic map structural feature.

[0095] Specifically, the sampler of the image synthesis model is used to generate corresponding images based on each synthetic remote sensing feature through a predefined update formula, which are then used as synthetic remote sensing images.

[0096] Based on this, this embodiment generates synthetic remote sensing images based on the structural features of the synthetic graph, remote sensing features, and vector features. Compared with existing remote sensing image synthesis technologies such as GAN or DDPM, which may cause problems such as confusion or reduced diversity, this embodiment restores the spatial layout of the real remote sensing image by using the spatial and semantic information in the structural features of the synthetic graph. The synthesized remote sensing image is more accurate, more refined, and more efficient, and can quickly acquire a large number of samples, which is beneficial to improving the training accuracy of the visual language model.

[0097] S224, take both remote sensing images and each synthetic remote sensing image as training remote sensing images, and obtain the corresponding image structure or synthetic image structure of each training remote sensing image as the corresponding training image structure, and obtain the corresponding vector image and synthetic vector image of each training remote sensing image as the corresponding training vector image, so as to construct an expanded sample library.

[0098] Specifically, for each remote sensing image and its corresponding vector image, steps S221 to S223 described above are used to obtain the corresponding composite image structure, composite remote sensing image, and composite vector image. Each composite image structure, along with its corresponding composite remote sensing image and composite vector image, is used as the corresponding training image structure, training remote sensing image, and training vector image. The original remote sensing images, vector images, and image structures generated based on the vector images are also used as the corresponding training remote sensing images, training vector images, and training image structures. All training remote sensing images, their corresponding training image structures, and training vector images are used as samples to form an expanded sample library.

[0099] S300: Based on each training remote sensing image, obtain the corresponding training image features; based on each training image structure, obtain the structural features of each training image; use each training image structural feature as training prior knowledge; input the training prior knowledge and the training image features into the visual language model to understand and recognize each training image feature; and use each training vector image as a training label to provide feedback on the accuracy of the interpretation results of each training image feature, so as to achieve the training of the visual language model.

[0100] The training image features are image patches centered on geographic features and having a preset size. Specifically, the center of each geographic feature is used as the image center, and image patches within a preset size range around the image center are cropped as the corresponding image features. It should be noted that the preset size is determined by the spatial size of the corresponding geographic feature. For example, for a certain cultivated land area in a remote sensing image, the size of the extracted image feature is determined by the spatial size of the cultivated land area. Those skilled in the art should know the specific setting method and principle of the preset size, and this example does not impose specific limitations here.

[0101] Each training influence feature is input into the visual language model, enabling the visual language model to analyze, identify, and learn the geographic elements corresponding to each image feature. This trains the visual language model to identify geographic elements in the image, thereby realizing the interpretation function of the visual language model.

[0102] Furthermore, by using the structural features of each training image as prior knowledge, since the structure of the training image includes spatial and semantic information of each node, it can reflect the actual spatial distribution pattern of ground features. Using the structural features of each image as prior knowledge to guide the interpretation of the features of the training image enables the visual semantic model to learn to capture the real spatial distribution pattern of ground features in remote sensing images, thereby effectively improving the accuracy of the interpretation of the trained visual semantic model.

[0103] Furthermore, each training prior knowledge is either individual prior knowledge or combined prior knowledge. Individual prior knowledge is individual graph structure features, that is, individual prior knowledge represents the spatial structure and semantic information of a single node. Combined prior knowledge is combined graph structure features, that is, combined prior knowledge represents the connection relationship between multiple nodes and the fused semantic features after the fusion of the semantic features of multiple nodes.

[0104] Specifically, based on prior knowledge from each training image, the fusion and recognition of features from each training image are guided to perform semantic and spatial analysis, thereby achieving an understanding of the features of each training image and thus training the visual language model.

[0105] In some alternative implementations, such as Figure 7 As shown, guided by prior knowledge, the features of each training image are understood and recognized, including:

[0106] S301: Based on prior knowledge of each entity, obtain the corresponding node type and entity semantic features; based on the features of each training image, each node type, and each entity semantic features, guide the fusion through attribute-aware attention to perform semantic analysis on the features of each training image and obtain semantically fused image features.

[0107] The node type of each individual prior knowledge is the type of the node corresponding to each individual prior knowledge. For example, if the node corresponding to the individual prior knowledge is a field planted with rapeseed, then the node type corresponding to the individual prior knowledge is field.

[0108] The semantic features of each prior knowledge are the semantic features of the nodes corresponding to each prior knowledge. For example, if the semantic features of a prior knowledge are {Hangzhou, plain, March, rapeseed}, then the semantic features of the prior knowledge are the text “rapeseed is planted in the plain area of ​​Hangzhou in March”.

[0109] Furthermore, based on the features of each training image, the type of each node, and the semantic features of each individual entity, attribute-aware attention is used to guide fusion in order to perform semantic analysis on the features of each training image.

[0110] Specifically, each training image feature is used as a query vector, each node type as a key vector, and each individual semantic feature as a value vector. The similarity between the query vector and the value vector is calculated to establish a one-to-one correspondence between each node and each image feature based on the similarity. Then, normalization is performed based on the calculated similarity to obtain the attention weight distribution. The attention weight distribution includes the weight magnitude corresponding to each element in the vector.

[0111] For example, the similarity between the query vector and any element in the key vector is calculated as follows:

[0112]

[0113] in, For the query vector and the key vector, the first... The similarity of elements, For query vector, The first key in the key vector Each element.

[0114] The attention weight distribution is obtained as follows:

[0115]

[0116] in, The first key vector The weight of each element corresponding to a node. This represents the total number of elements in the key vector.

[0117] The attention weight distribution is applied to the value vector, that is, based on the one-to-one correspondence between each node and each image feature, the image features and each individual semantic features are weighted and fused. The weights of the weighted fusion are the corresponding weights in the attention weight distribution.

[0118] S302: Based on combined prior knowledge, obtain the corresponding spatial features and semantic features of each node; based on the semantic fusion image features and the semantic features of each node, calculate the cosine similarity so that the semantic features of each node correspond one-to-one with the semantic fusion image features; and based on the spatial features, cluster the semantic fusion image features to perform spatial analysis on the semantic fusion image features.

[0119] Among them, the spatial features of the combined prior knowledge are the connection relationships between multiple nodes, and the semantic features of each node of the combined prior knowledge are used to characterize the semantic features corresponding to each node.

[0120] Specifically, the cosine similarity between each semantic fusion image feature and each node semantic feature is calculated to reflect the correspondence between each semantic fusion image feature and each node, so that each node semantic feature corresponds one-to-one with each semantic fusion image feature.

[0121] Based on spatial features, the connection relationships between nodes are obtained, and spatial analysis is performed on each semantic fusion image feature based on the semantic fusion image features corresponding to each node and the connection relationships between nodes. Specifically, the spatial relationships between each semantic fusion image feature are obtained, and each semantic fusion image feature is classified based on a clustering method. For example, semantic fusion image features that are adjacent in location are classified as image features of the same region, so as to realize the spatial analysis of each semantic fusion image feature.

[0122] Based on this, semantic and spatial analysis are performed on each image feature to achieve the interpretation of each training image feature.

[0123] In other alternative implementations, such as Figure 8 As shown, the interpretation of features of each training image, guided by prior knowledge, may also include:

[0124] S301': Based on prior knowledge of each entity, obtain the corresponding semantic features of each entity, and perform visual projection on the semantic features of each entity to obtain the corresponding visual labels; based on the features of each training image, combined with the visual labels, adopt a shared attention mechanism to perform semantic analysis on the image features and obtain the semantic fusion image features.

[0125] Among them, the semantic features of each entity's prior knowledge are the semantic features of the nodes corresponding to each entity's prior knowledge.

[0126] Furthermore, semantic features, i.e., textual information, are projected onto a visual feature, serving as the corresponding visual label. A shared attention mechanism is employed for each training image feature and each visual label to perform interactive fusion, thereby obtaining semantically fused image features. Specifically, feature enhancement with the same weight parameters is applied to each training image feature and each visual label to analyze the differences and changes in each training image feature and each visual label after enhancement, and to adjust the weight parameters accordingly. Ultimately, this aligns and fuses the image features and each visual label, thus obtaining semantically fused image features.

[0127] S302': Self-attention calculation is performed based on the semantic fusion image features to obtain the enhanced remote sensing features. Based on combined prior knowledge, the corresponding spatial features and semantic features of each node are obtained. Based on the enhanced remote sensing features, combined with the spatial features and semantic features of each node, multi-head cross-attention calculation is performed to inject spatial relationships into the enhanced remote sensing features and to perform spatial analysis on the semantic fusion image features.

[0128] Specifically, self-attention computation is performed on each semantically fused image feature to enhance its global semantic meaning. Furthermore, by combining spatial features from prior knowledge with the semantic features of each node, and inputting these features into a multi-head cross-attention layer, interactive fusion is achieved. This ensures that each semantically fused image feature corresponds one-to-one with each node based on its semantic features, and spatial relationships are obtained through spatial features, enabling spatial analysis of each semantically fused image feature.

[0129] Based on this, semantic and spatial analysis can be performed on the features of each training image, thereby enabling the interpretation of the features of each training image.

[0130] It should be noted that the above explanation describes the interpretation principles during the training process of the visual language model. In practical applications, the interpretation process and principles of the visual language model for remote sensing images remain consistent with the above. Specifically, the training prior knowledge becomes expert prior knowledge in the actual interpretation process, the individual prior knowledge becomes individual-type expert prior knowledge in the actual interpretation process, and the combined prior knowledge becomes combined-type expert prior knowledge in the actual interpretation process.

[0131] It should be noted that during the guided interpretation process based on prior knowledge, the weight parameters of each attention mechanism need to be learned and optimized through training. Based on this, the visual semantic model is trained using each training remote sensing image and its corresponding training image structure and training vector image in the expanded sample library.

[0132] Specifically, since each training vector map contains geographic feature information corresponding to each geographic element, this geographic feature information is used as a label for the interpretation result to provide feedback on the accuracy of the interpretation result of each training image feature. For example, feedback can be provided through a loss function to assist in the optimization of the visual semantic model and thus improve the accuracy of the visual semantic model.

[0133] like Figure 9 As shown, this embodiment provides a remote sensing image interpretation system, including a data acquisition module 41 and a visual language module 42.

[0134] Among them, the data acquisition module 41 is used to acquire remote sensing images and preset expert prior knowledge.

[0135] The visual language module 42 includes a trained visual language model 421, which is used to understand and recognize remote sensing images based on remote sensing images and expert prior knowledge, thereby realizing the interpretation of remote sensing images.

[0136] Based on the same technical concept, the remote sensing image interpretation method provided in the embodiments of the present invention can be implemented on the terminal side or the server side.

[0137] like Figure 10 The diagram illustrates an optional hardware structure of a terminal according to an embodiment of the present invention. The terminal 50 can be a mobile phone, computer device, tablet device, personal digital processing device, factory back-end processing device, etc. The terminal 50 includes at least one processor 51, a memory 52, at least one network interface 54, and a user interface 53. The various components in the system are coupled together through a bus system 55. It is understood that the bus system 55 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 55 also includes a power bus, a control bus, and a status signal bus.

[0138] The user interface 53 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0139] It is understood that memory 52 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memory characterized in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable categories of memory.

[0140] In this embodiment of the invention, the memory 52 is used to store various types of data to support the operation of the terminal. Examples of this data include: any executable program for operation on the terminal 50, such as the operating system 521 and application programs 522; the operating system 521 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 522 may contain various applications, such as media players, browsers, etc., for implementing various application services. The remote sensing image interpretation method provided in this embodiment of the invention can be included in the application program 522.

[0141] The methods disclosed in the above embodiments of the present invention can be applied to processor 51, or implemented by processor 51. Processor 51 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 51 or by instructions in the form of software. The processor mentioned above may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 51 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present invention. Processor 51 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.

[0142] In an exemplary embodiment, terminal 50 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.

[0143] This invention also provides a computer-readable storage medium storing a computer program that, when invoked by a processor, implements the remote sensing image interpretation method provided by this invention.

[0144] Computer-readable storage media can be tangible devices capable of holding and storing instructions used by an instruction execution device. Computer-readable storage media can be, for example, (but not limited to) electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, and mechanical encoding devices.

[0145] The computer-readable program represented herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, to an external computer or external storage device. A network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network and forwards these instructions to the computer-readable storage medium in the respective computing / processing device.

[0146] In summary, this application generates graph structures based on vector graphics to obtain spatial and semantic information, thereby capturing spatial relationships and semantic information in the interpretation of remote sensing images, improving the accuracy of the interpretation results, and contributing to better interpretation effects.

[0147] The descriptions of the processes or structures corresponding to the above-mentioned figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0148] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A graph-structure-guided remote sensing image interpretation method, comprising: Acquire remote sensing imagery and pre-defined expert prior knowledge; Based on the remote sensing image and the expert's prior knowledge, a trained visual language model is input to understand and recognize the remote sensing image, thereby realizing the interpretation of the remote sensing image; The training process of the visual language model includes: Acquire at least one of the remote sensing images and each vector image corresponding to each of the remote sensing images; Based on the remote sensing images and their corresponding vector graphics, an expanded sample library is obtained by synthesizing and expanding the images. The expanded sample library includes several training remote sensing images and their corresponding training image structures and training vector graphics. Based on the training remote sensing images, corresponding training image features are obtained, and structural features of each training image are obtained based on the structure of each training image. The structural features of each training image are used as prior knowledge for training. The prior knowledge and the features of each training image are input into the visual language model to understand and recognize the features of each training image. The training vector images are used as training labels to provide feedback on the accuracy of the interpretation results of each training image feature, thereby achieving the training of the visual language model. The process of synthesizing and expanding any of the remote sensing images and their corresponding vector maps includes: obtaining the corresponding graph structure based on the vector map; obtaining corresponding composite graph structures by randomly modifying the graph structure; obtaining corresponding composite vector maps based on each composite graph structure; extracting corresponding remote sensing features based on the remote sensing images; extracting corresponding vector features based on the vector maps; obtaining features of each composite graph structure based on the spatial and semantic information of each composite graph structure through a geometric semantic perception module; for each composite graph structure, obtaining corresponding composite remote sensing images based on the corresponding composite graph structure features, combined with the remote sensing features and the vector features; using both the remote sensing images and the composite remote sensing images as training remote sensing images, and obtaining the corresponding graph structures or composite graph structures of each training remote sensing image as the corresponding training graph structures, and obtaining the corresponding vector maps or composite vector maps of each training remote sensing image as the corresponding training vector maps, to construct the expanded sample library.

2. The remote sensing image interpretation method according to claim 1, characterized in that, The training prior knowledge mentioned above can be single prior knowledge or combined prior knowledge; the step of inputting the training prior knowledge and the training image features into the visual language model to understand and recognize the training image features includes: Based on the prior knowledge of each entity, the corresponding node type and entity semantic features are obtained; based on the training image features, node types, and entity semantic features, attribute-aware attention is used to guide fusion in order to perform semantic analysis on the training image features and obtain semantically fused image features. Based on the combined prior knowledge, the corresponding spatial features and semantic features of each node are obtained; based on the semantic fusion image features and the semantic features of each node, cosine similarity is calculated so that the semantic features of each node correspond one-to-one with the semantic fusion image features; and based on the spatial features, the semantic fusion images are clustered to perform spatial analysis on the semantic fusion image features.

3. The remote sensing image interpretation method according to claim 1, characterized in that, The training prior knowledge mentioned above can be single prior knowledge or combined prior knowledge; the step of inputting the training prior knowledge and the training image features into the visual language model to understand and recognize the training image features includes: Based on the prior knowledge of each entity, obtain the corresponding semantic features of the entity, and perform visual projection on the semantic features of each entity to obtain the corresponding visual markers; based on the training image features of each entity, combined with the visual markers, adopt a shared attention mechanism to perform semantic analysis on the influencing features to obtain the semantic fusion image features of each entity. Self-attention calculation is performed based on the semantic fusion image features to obtain the enhanced remote sensing features. Based on the combined prior knowledge, the corresponding spatial features and semantic features of each node are obtained. Based on the enhanced remote sensing features, multi-head cross-attention calculation is performed in combination with the spatial features and the semantic features of each node to inject spatial relationships into the enhanced remote sensing features and to perform spatial analysis on the semantic fusion image features.

4. The remote sensing image interpretation method according to claim 1, characterized in that, For any of the aforementioned composite map structures, based on the corresponding composite map structure features, and combining the aforementioned remote sensing features and vector features, the corresponding composite remote sensing images are obtained, including: The structural features of each of the synthetic graphs, the remote sensing features, and the vector features are normalized and aligned. The normalized and aligned structural features, remote sensing features, and vector features of each synthetic image are enhanced by self-attention calculation and fused by multi-head cross-attention calculation to obtain each synthetic remote sensing feature. Based on the aforementioned synthetic remote sensing features, each of the aforementioned synthetic remote sensing images corresponding to the aforementioned synthetic map structural features is generated.

5. The remote sensing image interpretation method according to claim 1, characterized in that, The spatial and semantic information based on the graph structure is used to obtain the features of each graph structure through a geometric semantic perception module, including: The semantic information of the graph structure is extracted into semantic features using a preset text encoder. Based on the spatial information of the graph structure, mapping is performed to obtain each of the spatial features; Based on a pre-defined graph neural network, the correlation between each semantic feature and each spatial feature is obtained through an attention mechanism to generate the graph structure features.

6. The remote sensing image interpretation method according to claim 1, characterized in that, The step of obtaining the corresponding graph structure based on the vector image includes: Based on the geographic features of the vector map, generate the corresponding nodes; Based on the geographic coverage of each geographic element, obtain the bounding box of each node; based on the geospatial relationship of each geographic element, obtain the connection relationship of each node; based on the bounding box and connection relationship of each node, obtain the spatial information of each node. Based on the geographic element information of each of the aforementioned geographic elements, the text description of each of the aforementioned nodes is obtained as the semantic information of each of the aforementioned nodes; The graph structure is obtained based on each node and its spatial and semantic information.

7. A graph-structure-guided remote sensing image interpretation system, characterized in that, It includes a data acquisition module and a visual semantics module; The data acquisition module is used to acquire remote sensing images and preset expert prior knowledge; The visual language module includes a trained visual language model, which is used to understand and recognize the remote sensing image based on the remote sensing image and the expert's prior knowledge, thereby realizing the interpretation of the remote sensing image. The training process of the visual language model includes: Acquire at least one of the remote sensing images and each vector image corresponding to each of the remote sensing images; Based on the remote sensing images and their corresponding vector graphics, an expanded sample library is obtained by synthesizing and expanding the images. The expanded sample library includes several training remote sensing images and their corresponding training image structures and training vector graphics. Based on the training remote sensing images, corresponding training image features are obtained, and structural features of each training image are obtained based on the structure of each training image. The structural features of each training image are used as prior knowledge for training. The prior knowledge and the features of each training image are input into the visual language model to understand and recognize the features of each training image. The training vector images are used as training labels to provide feedback on the accuracy of the interpretation results of each training image feature, thereby achieving the training of the visual language model. The process of synthesizing and expanding any of the remote sensing images and their corresponding vector maps includes: obtaining the corresponding graph structure based on the vector map; obtaining corresponding composite graph structures by randomly modifying the graph structure; obtaining corresponding composite vector maps based on each composite graph structure; extracting corresponding remote sensing features based on the remote sensing images; extracting corresponding vector features based on the vector maps; obtaining features of each composite graph structure based on the spatial and semantic information of each composite graph structure through a geometric semantic perception module; for each composite graph structure, obtaining corresponding composite remote sensing images based on the corresponding composite graph structure features, combined with the remote sensing features and the vector features; using both the remote sensing images and the composite remote sensing images as training remote sensing images, and obtaining the corresponding graph structures or composite graph structures of each training remote sensing image as the corresponding training graph structures, and obtaining the corresponding vector maps and composite vector maps of each training remote sensing image as the corresponding training vector maps, to construct the expanded sample library.

8. A terminal, characterized in that, include: A processor and a memory, wherein the memory and the processor are communicatively connected; The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to enable the terminal to perform the graph-guided remote sensing image interpretation method as described in any one of claims 1 to 6.

9. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the graph-guided remote sensing image interpretation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Earthquake disaster area remote sensing image interpretation method based on graph transformation knowledge embedding algorithm

    CN113435268A

  • Method for interpreting remote-sensing image on basis of lifelong learning

    WO2024199249A1