Remote sensing scene graph prediction method and system based on visual language model
Through the method based on the visual language model, remote sensing scene graph prediction prompt words are generated and pre-trained and fine-tuned visual language model is used to solve the problem of missing attribute prediction and relying on a large amount of labeled data in remote sensing scene graph prediction, achieving a scene graph prediction effect with higher accuracy and less dependence on labeled data.
Patent Information
- Application Number
- CN202510029969.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-16
AI Technical Summary
The existing remote sensing scene graph prediction methods lack attribute prediction and rely too much on a large amount of labeled data, making it difficult to obtain detailed semantic information in the image.
Using a method based on visual language model, the remote sensing scene graph prediction prompt words are generated, and the pre-trained and fine-tuned visual language model is used to generate the prediction results in the triple format, and the mask elements are extracted and filled into the scene graph to achieve the generation of the complete scene graph.
This enhances the model's specialized understanding of mask elements in different dimensions, improves the accuracy of scene graph prediction, and reduces the dependence on labeled data, and improves the reliability of the model under the condition of few samples.
Smart Images

Figure CN120012919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scene graph prediction, and in particular to a remote sensing scene graph prediction method and system based on a visual language model. Background Art
[0002] With the development of remote sensing technology, high-resolution remote sensing image interpretation has become a research hotspot in recent years, especially for high spatial resolution remote sensing images. Remote sensing scene graphs are designed to detect objects and the relationship between them. They are a structured and semantic representation method for remote sensing image content. Compared with traditional basic tasks such as target detection and scene classification, remote sensing scene graph prediction is a higher-level image semantic understanding task. Remote sensing scene graphs encode various fine-grained semantic information of remote sensing images. Constructing remote sensing scene graph prediction tasks, using visual-language models, and learning cross-modal detailed semantic alignment can promote intelligent understanding of geographic spatial scenes from perception to cognition, and have important application value in resource exploration, environmental monitoring, natural disaster prevention, military reconnaissance, etc.
[0003] At present, the research on scene graph prediction in the field of computer vision mainly focuses on improving the accuracy of relational predicate prediction. Common methods are divided into two categories. One is to focus on fitting the distribution of relational data, including combining contextual features and introducing prior knowledge, and the other is to focus on softening the imbalanced distribution of relations, usually using methods such as reweighting and resampling to improve the long-tail distribution problem. However, detailed semantics, including objects, their attributes, and the relationships between them, are crucial for accurately understanding visual scenes. Traditional scene graph prediction methods are difficult to obtain detailed semantic information in images. Visual language models can simultaneously process and understand both visual (image) and language (text) modal information, and perform well in complex tasks such as visual question answering, image description generation, and image retrieval. Due to the large number of types of objects in remote sensing images, the complexity of relationships and attributes far exceeds that of small-area natural images, and due to different imaging angles, scene graph prediction methods in the field of computer vision cannot be directly migrated to remote sensing images. In recent years, relevant scholars have begun to pay attention to the research related to remote sensing scene graphs, but they mainly focus on dataset construction and relationship prediction tasks. In terms of relationship prediction, there is a lack of prediction of attribute information, and it usually relies on the constructed dataset, with poor generalization ability. Therefore, the remote sensing field needs a complete and accurate solution that adapts to the needs of multi-dimensional scene graph prediction to meet the needs of multi-field applications. Summary of the invention
[0004] To this end, the present invention provides a remote sensing scene graph prediction method and system based on a visual language model to solve the problems that the existing remote sensing scene graph prediction lacks attribute prediction and over-depends on a large amount of labeled data.
[0005] According to the design scheme provided by the present invention, on the one hand, a remote sensing scene graph prediction method based on a visual language model is provided, comprising:
[0006] Generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task, wherein the mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph, and the prompt words are used to describe the mask and predicted element types required in the scene graph prediction task;
[0007] Inputting the prediction prompt words of the remote sensing scene graph and the target remote sensing image into the target model, and using the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements and relationship elements in the target remote sensing image, wherein the target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset;
[0008] The mask elements corresponding to the prediction prompt words of the remote sensing scene graph are extracted from the triple format prediction output, and the mask elements are filled into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
[0009] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, the corresponding remote sensing scene graph prediction prompt words are generated according to the mask element type required for the scene graph prediction task, including:
[0010] Constructing a dataset consisting of remote sensing images and scene graphs with mask elements corresponding to the remote sensing images;
[0011] Obtaining scene graph triplets under all mask element types according to scene graph constituent element types, wherein the triplets consist of entities, attributes, and relationships;
[0012] Traverse the scene graph triples under each mask element type and extract the mask element part, and generate the corresponding mask element question prompt word according to the mask element type.
[0013] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, a remote sensing image dataset is used to pre-train and fine-tune the visual language model, including:
[0014] Acquire a data set, wherein the data set consists of a plurality of remote sensing images and a scene graph with mask elements corresponding to each remote sensing image;
[0015] The remote sensing images and corresponding scene graphs in the dataset are used as model inputs, and the triplets containing the labeled mask elements in the scene graph are used as model outputs to pre-train the visual language model using the dataset.
[0016] The pre-trained visual language model is fine-tuned using the low-rank adaptation method, and the fine-tuned visual language model is used as the target model.
[0017] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, the process of Lora fine-tuning is expressed as: Wx+ΔWx=Wx+BAx, wherein W is a pre-trained parameter of the visual language model, ΔW is a model parameter that needs to be updated in Lora fine-tuning, x is a model input, and A and B are decomposition parameters of ΔW in Lora fine-tuning.
[0018] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, in the data set, the Open_clip model is used to normalize the annotations of the triples in the scene graph of the data set.
[0019] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, the visual language model adopts an open source large-scale visual language model with visual reasoning ability and text comprehension ability.
[0020] As a remote sensing scene graph prediction method based on a visual language model of the present invention, further, a mask element in the scene graph adopts a Mask mode.
[0021] In another aspect, the present invention further provides a remote sensing scene graph prediction system based on a visual language model, comprising: a prompt generation module, a model prediction module and a mask filling module, wherein:
[0022] A prompt generation module, used to generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task, wherein the mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph, and the prompt words are used to describe the mask and predicted element types required in the scene graph prediction task;
[0023] A model prediction module is used to input the remote sensing scene graph prediction prompt words and the target remote sensing image into the target model, and use the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements and relationship elements in the target remote sensing image. The target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset;
[0024] The mask filling module is used to extract the mask elements corresponding to the remote sensing scene graph prediction prompt words from the triple format prediction output, and fill the mask elements into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
[0025] Beneficial effects of the present invention:
[0026] The present invention generates prompt words from multiple dimensions and angles for different element types in the scene graph, and pre-trains and fine-tunes the visual language model in combination with the knowledge base in the field of remote sensing images to achieve scene graph prediction, enhance the specialized understanding of mask elements of different dimensions by the large model, and after fine-tuning with small-scale remote sensing image annotation data, given a remote sensing image and a scene graph of different masked elements, it is possible to predict masked elements such as entities, attributes, relationships, etc., and reload the predicted elements into the scene graph and output them, thereby improving the model prediction accuracy and reliability under training with fewer data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the remote sensing scene graph prediction process based on the visual language model in the embodiment;
[0028] Figure 2 Schematic diagram of the remote sensing scene graph prediction algorithm in the embodiment. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.
[0030] In view of the problems of lack of attribute prediction and over-reliance on a large amount of labeled data in the existing remote sensing scene graph prediction, the present invention embodiment, see Figure 1 As shown, a remote sensing scene graph prediction method based on a visual language model is provided, comprising:
[0031] S101. Generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task. The mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph. The prompt words are used to describe the mask and predicted element types required in the scene graph prediction task.
[0032] Specifically, the corresponding remote sensing scene graph prediction prompt words are generated according to the mask element type required for the scene graph prediction task, which can be designed to include:
[0033] Constructing a dataset consisting of remote sensing images and scene graphs with mask elements corresponding to the remote sensing images;
[0034] The Open_clip model can be used to normalize the annotations of the triplets in the scene graph of the dataset to pre-train visual language models such as qwen-vl-chat.
[0035] According to the scene graph element type, the scene graph triples under all mask element types are obtained, and the triples are composed of entities, attributes and relationships. For example, when describing a river in a scene graph, its triple composition can be expressed as follows: (river, located in Zhengzhou), (river, passing through farmland), (river, shape, curved), (river, color, blue).
[0036] Traverse the scene graph triples under each mask element type and extract the mask element part, generate the corresponding mask element question prompt word according to the mask element type, and the mask element can adopt the Mask method. For example, the question prompt word for the missing entity element can be described as "What entities are there in this image?", "Please describe the entities in the image."; the question prompt word for the missing relationship element can be described as "What is the relationship between the entities in the image?", "Please give the relationship between the entities in the image"; the question prompt word for the missing attribute element can be described as "What characteristics can be seen from the image of the entity?", "Please describe the characteristics of the entity in the image".
[0037] By traversing all the triples, determine the parts of the mask elements and classify them according to the composition, such as entities, attributes and relationships. According to different classifications, prompts for corresponding questions will be generated. By constructing diverse and reasonable prompts, the "zero sample" and "few sample" methods are used to help the large model better understand the input intent. "Zero sample" prompt example: I am doing a remote sensing scene graph prediction task, that is, identifying the types of objects, attributes and relationships between objects in the image. Please output the main types of objects in the current image. Please output the object relationships and attributes in the form of triples, for example: <river, through, farmland>, <river, color, blue>. "Few sample" prompt example: Input sample image + prompt: I am doing a remote sensing scene graph prediction task. When I input this image, you have to answer that the objects in the current image include: river, farmland; the relationship between them is: <river, through, farmland>, and the attributes are <river, color, blue>. Please answer the image I input below.
[0038] S102, inputting the remote sensing scene graph prediction prompt words and the target remote sensing image into the target model, and using the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements, and relationship elements in the target remote sensing image, wherein the target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset;
[0039] Among them, using remote sensing image datasets to pre-train and fine-tune the visual language model may include:
[0040] Acquire a data set, wherein the data set consists of a plurality of remote sensing images and a scene graph with mask elements corresponding to each remote sensing image;
[0041] The remote sensing images and corresponding scene graphs in the dataset are used as model inputs, and the triplets containing the labeled mask elements in the scene graph are used as model outputs to pre-train the visual language model using the dataset.
[0042] The pre-trained visual language model is fine-tuned using the low-rank adaptation method, and the fine-tuned visual language model is used as the target model.
[0043] Pre-training visual language models such as qwen-vl-chat using pre-labeled data sets makes the model's "attention" more focused on the remote sensing field. After completing the model pre-training, the model will be fine-tuned using Lora. When fine-tuning Lora, a low-rank matrix is set to represent the parameter update ΔW for the pre-trained weight matrix k using a low-rank decomposition. The calculation formula is shown in formula (1).
[0044] Wx+ΔWx=Wx+BAx (1)
[0045] W is the parameter initialized by the pre-trained model, and ΔW is the parameter that needs to be updated. During the training process of Lora, W is fixed, and x represents the input sample. The sample is fixed and only A and B are training parameters.
[0046] S103, extracting mask elements corresponding to the remote sensing scene graph prediction prompt words from the triple format prediction output, and filling the mask elements into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
[0047] like Figure 2 As shown in the figure, the model that has completed pre-training and fine-tuning is a specialized large model with a massive knowledge base focused on remote sensing scene graph judgment. The model can predict mask elements based on the input mask scene graph and remote sensing image, and output the results as triples according to the training format. That is, the model can judge the category of the mask element based on the previously input mask scene graph and extract the element from the triple prediction output. By traversing the scene graph, the mask elements are refilled back into all the mask triplets of the scene graph, and after assembly, they are output as the result.
[0048] Furthermore, based on the above method, an embodiment of the present invention also provides a remote sensing scene graph prediction system based on a visual language model, comprising: a prompt generation module, a model prediction module and a mask filling module, wherein:
[0049] A prompt generation module, used to generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task, wherein the mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph, and the prompt words are used to describe the mask and predicted element types required in the scene graph prediction task;
[0050] A model prediction module is used to input the remote sensing scene graph prediction prompt words and the target remote sensing image into the target model, and use the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements and relationship elements in the target remote sensing image. The target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset;
[0051] The mask filling module is used to extract the mask elements corresponding to the remote sensing scene graph prediction prompt words from the triple format prediction output, and fill the mask elements into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
[0052] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0053] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0054] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0055] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.
[0056] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A remote sensing scene graph prediction method based on a visual language model, characterized in that: Include: Generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task, wherein the mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph, and the prompt words are used to describe the mask and predicted element types required in the scene graph prediction task; Inputting the prediction prompt words of the remote sensing scene graph and the target remote sensing image into the target model, and using the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements and relationship elements in the target remote sensing image, wherein the target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset; The mask elements corresponding to the prediction prompt words of the remote sensing scene graph are extracted from the triple format prediction output, and the mask elements are filled into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
2. The remote sensing scene graph prediction method based on a visual language model according to claim 1, characterized in that: Generate the corresponding remote sensing scene graph prediction prompt words according to the mask element type required by the scene graph prediction task, including: Constructing a dataset consisting of remote sensing images and scene graphs with mask elements corresponding to the remote sensing images; Obtaining scene graph triplets under all mask element types according to scene graph constituent element types, wherein the triplets consist of entities, attributes, and relationships; Traverse the scene graph triples under each mask element type and extract the mask element part, and generate the corresponding mask element question prompt word according to the mask element type.
3. The remote sensing scene graph prediction method based on a visual language model according to claim 1, characterized in that: Use remote sensing image datasets to pre-train and fine-tune the visual language model, including: Acquire a data set, wherein the data set consists of a plurality of remote sensing images and a scene graph with mask elements corresponding to each remote sensing image; The remote sensing images and corresponding scene graphs in the dataset are used as model inputs, and the triplets containing the labeled mask elements in the scene graph are used as model outputs to pre-train the visual language model using the dataset. The pre-trained visual language model is fine-tuned using the low-rank adaptation method, and the fine-tuned visual language model is used as the target model.
4. The remote sensing scene graph prediction method based on a visual language model according to claim 3, characterized in that: The process of Lora fine-tuning is expressed as: Wx+ΔWx=Wx+BAx, where W is the pre-trained parameter of the visual language model, ΔW is the model parameter that needs to be updated in Lora fine-tuning, x is the model input, and A and B are the decomposition parameters of ΔW in Lora fine-tuning.
5. The remote sensing scene graph prediction method based on the visual language model according to claim 2 or 3, characterized in that: In the dataset, the Open_clip model is used to normalize the annotations of the triplets in the scene graph of the dataset.
6. The remote sensing scene graph prediction method based on a visual language model according to claim 1, characterized in that: The visual language model adopts an open source large-scale visual language model with visual reasoning and text understanding capabilities.
7. The remote sensing scene graph prediction method based on a visual language model according to claim 1, characterized in that: The mask element in the scene graph uses the Mask method.
8. A remote sensing scene graph prediction system based on a visual language model, characterized in that: Contains: prompt generation module, model prediction module and mask filling module, among which, A prompt generation module, used to generate corresponding remote sensing scene graph prediction prompt words according to the mask element types required for the scene graph prediction task, wherein the mask element types include entity type elements, attribute type elements and relationship type elements in the scene graph, and the prompt words are used to describe the mask and predicted element types required in the scene graph prediction task; A model prediction module is used to input the remote sensing scene graph prediction prompt words and the target remote sensing image into the target model, and use the target model to generate a prediction result in a triple format consisting of entity elements, attribute elements and relationship elements in the target remote sensing image. The target model is a visual language model pre-trained and fine-tuned using a remote sensing image dataset; The mask filling module is used to extract the mask elements corresponding to the remote sensing scene graph prediction prompt words from the triple format prediction output, and fill the mask elements into the scene graph to obtain a complete scene graph corresponding to the target remote sensing image.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 can be implemented.