CLIP model-based geographic scene graph generation method and system
By introducing CLIP models and segmentation models in geographic image processing, geographic scene maps are generated, and the problems of difficult and low accuracy of geographic image scene map generation are solved, and high-accuracy geographic scene map generation is achieved.
Patent Information
- Application Number
- CN202510053440.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
When the prior art is faced with the generation of scene maps of geographic images, it is difficult and has low prediction accuracy, and cannot effectively improve the interpretation efficiency and accuracy of geographic images.
The geographical scene graph generation method based on the CLIP model is adopted, and the semantic segmentation is performed through the segmentation model, key land objects are selected, the relationship between key land objects is predicted, and triples are generated based on OSM information to finally generate a scene graph.
It realizes the high accuracy generation of geographical scene maps, improves the accuracy and reliability of predicate relationship prediction, and enhances the completeness and accuracy of geographic attribute information.
Smart Images

Figure CN119942341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross-technical field of computer vision and geography, and in particular to a method and system for generating a geographic scene graph based on a CLIP model. Background Art
[0002] Scene graph generation aims to enable computers to recognize objects and their relationships in images like humans. By expressing abstract semantics of images, it helps computers complete high-level visual tasks. It is a hot research direction in the field of computer vision and has been widely used in downstream tasks such as image retrieval, image generation, image description, and visual question answering. In recent years, with the rapid development of remote sensing and mapping equipment such as remote sensing satellites and drones, a large number of high-resolution geographic images have been collected. Geographic scene graph generation is based on the input geographic image, which can automatically generate relationship triples describing the content of the geographic image, which can be regarded as a subgraph of the geographic knowledge graph. It has important theoretical and practical value in improving the data processing and analysis capabilities of geographic information systems, environmental monitoring and management, urban planning and design, etc. Traditional manual judgment and screening methods require huge investment but are inefficient, so new technical means are urgently needed to improve the efficiency and accuracy of geographic image interpretation. Geographic scene graphs can provide knowledge support for tasks such as image interpretation, natural language description, and visual question answering, and further improve the level of intelligent understanding of geographic images.
[0003] At present, there have been a lot of studies on scene graph generation algorithms for natural scenes in the field of computer vision, while the progress of scene graph generation technology for geographic images lags far behind that of natural images. The main reasons include the following aspects: First, there is a lack of relevant datasets that contain both targets and geographic spatial relationships in the field of geography, which makes it difficult to train and evaluate scene graph generation algorithms. Second, geographic images usually use vertical perspectives such as "bird's eye view" to reflect the complex distribution of objects on the earth's surface, which is quite different from natural scenes and cannot be generalized from datasets of natural scenes. Therefore, many natural scene graph generation algorithms do not work well when migrated to geographic images. In addition, the large span and special spatial distribution of entities in geographic images increase the difficulty of scene graph generation. Although remote sensing image scene graph generation methods have also been proposed in the field of geography, they are mainly aimed at large-format remote sensing images, and are usually based on target detection methods. The main body of the object is marked by the target box and the relationship between the main body is predicted using the relationship network. There are problems of insufficient precision and fine-grained information, which affects the accuracy of relationship prediction. Summary of the invention
[0004] In order to at least partially solve the problem of difficulty in scene graph generation and low prediction accuracy for geographic images, the present invention provides a method and system for generating geographic scene graphs based on the CLIP model, which selects key objects after semantic segmentation through a segmentation model under the condition of a given remote sensing image, predicts the relationship between key objects, generates triples in combination with OSM information, and finally generates a scene graph. The present invention realizes the generation of geographic scene graphs through a relatively simple method with high accuracy.
[0005] In order to achieve the above object, the technical solution of the present invention is: The first aspect of the present invention proposes a method for generating a geographic scene graph based on a CLIP model, comprising: Step 1: Collect geographic images to construct a geographic scene graph dataset; Step 2: Input the training image in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain a geographic scene graph corresponding to the training image, so as to optimize the geographic scene graph generation unit according to the geographic scene graph corresponding to the training image; Step 3: According to the geographic scene graph, the precision, recall and average recall of the geographic scene graph generation unit are obtained, and the geographic scene graph generation unit is optimized according to the precision, recall and average recall to obtain the optimal geographic scene graph generation unit, so as to improve the accuracy of generating the geographic scene graph; Step 4: Input the target image into the optimal geographic scene graph generation unit to obtain the predicted triplet and geographic scene graph corresponding to the target image.
[0006] Further, the geographic scene graph generation unit includes a segmentation model, a pixel-level filter, a CLIP model, a multi-layer perception mechanism, a subtraction model, a classification network, and a generation subunit; The segmentation model is used to segment the input image to obtain a ground object type label segmentation result, which is convenient for subsequent identification of ground object types; The pixel-level filter is used to filter out the main types of objects in the training image, and arrange the main object type labels in pairs in the form of subject and object to generate all combinations; The CLIP model is used to extract features from the output of the pixel-level filter to obtain text features of the subject and the object; and to extract features from the image area of the main landform type to obtain image features, so as to remove redundant feature information and thereby ensure the accuracy of the model; The multi-layer perception mechanism is used to process image features and text features to obtain a joint embedding vector to facilitate predicate prediction; The subtraction model is used to subtract the joint embedding vector from the subject text feature and the object text feature to obtain the subtracted feature; The classification network is used to process the subtracted features to obtain predicted predicates, and finally obtain predicted triples; The generating subunit is used to generate a geographic scene graph according to the predicted triples.
[0007] Furthermore, the segmentation model is used to segment the input image to obtain a ground object type label segmentation result, which specifically includes: The segmentation model is used to classify the input image pixel by pixel, identify different object areas in the image, and assign corresponding object type labels to each area according to the color value of the pixel to obtain the object type label segmentation result, which is convenient for extracting image features.
[0008] Furthermore, the extracting features of the image area of the main land object type to obtain image features specifically includes: Mask processing is performed on the area of the non-combined feature type to obtain mask features and joint representation areas. The mask features and joint representation areas are input into the CLIP model to obtain image features of the joint representation area, which facilitates the removal of redundant feature information.
[0009] Furthermore, the multi-layer perceptron is used to process the image features and the text features to obtain a joint embedding vector, specifically including: The image features are added to the text features, and the result of the addition is embedded into the text features of the subject and object using a multi-layer perception mechanism to obtain a joint embedding vector.
[0010] Furthermore, the optimization of the geographic scene graph generation unit according to the precision rate, the recall rate and the average recall rate specifically includes: If the precision, recall and average recall are lower than the preset values, the segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network are optimized and trained using their corresponding loss functions to obtain the optimal segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network, and then the optimal geographic scene graph generation unit is obtained, so that an accurate geographic scene graph can be obtained through the model.
[0011] The second aspect of the present invention proposes a geographic scene graph generation system based on the CLIP model, comprising: A collection module, used to collect geographic images to construct a geographic scene graph dataset; A geographic scene graph generation module is used to input the training image in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain a geographic scene graph corresponding to the training image, so as to optimize the geographic scene graph generation unit according to the geographic scene graph corresponding to the training image; A training module is used to obtain the precision, recall and average recall of the geographic scene graph generation unit according to the geographic scene graph, and optimize the geographic scene graph generation unit according to the precision, recall and average recall to obtain the optimal geographic scene graph generation unit, so as to improve the accuracy of generating the geographic scene graph; The generation module is used to input the target image into the optimal geographic scene graph generation unit to obtain the prediction triplet and geographic scene graph corresponding to the target image.
[0012] The third aspect of the present invention proposes an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements a method for generating a geographic scene graph based on a CLIP model as described in the first aspect above.
[0013] The fourth aspect of the present invention proposes a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a method for generating a geographic scene graph based on a CLIP model as described in the first aspect above.
[0014] Beneficial effects of the present invention: (1) The present invention introduces a segmentation model to perform accurate semantic segmentation on geographic images and screen out key features. Compared with the traditional method that only relies on bounding box features, it has higher accuracy.
[0015] (2) The present invention utilizes the multimodal model CLIP to extract text features and image features, and after feature fusion, applies them to the prediction of predicate relations, effectively combining the semantic information of text and image, thereby improving the accuracy and reliability of predicate relation prediction.
[0016] (3) The present invention supplements the attributes of features by combining open source map data and utilizes external data sources to enhance the integrity and accuracy of feature attribute information, thereby improving the generation quality and practicality of scene graphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 One of the flowcharts of a method for generating a geographic scene graph based on a CLIP model provided in an embodiment of the present invention.
[0018] Figure 2 The present invention provides a second flowchart of a method for generating a geographic scene graph based on a CLIP model according to an embodiment of the present invention.
[0019] Figure 3 An architecture diagram of a geographic scene graph generation system based on the CLIP model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Example 1 like Figure 1 and Figure 2 As shown, a method for generating a geographic scene graph based on a CLIP model includes: Step 1: Collect geographic images to construct a geographic scene graph dataset.
[0022] Step 2: Input the training image in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain the geographic scene graph corresponding to the training image.
[0023] Specifically, the geographic scene graph generation unit includes a segmentation model (pixel-level semantic segmentation model), a pixel-level filter, a CLIP model, a multi-layer perception mechanism, a subtraction model, a classification network, and a generation subunit.
[0024] The segmentation model is used to segment the input image and obtain the object type label segmentation result.
[0025] The pixel-level filter is used to filter out the main types of objects in the training images, and arrange the main object type labels in pairs in the form of subject and object to generate all combinations. During training, the generated combinations need to be manually annotated to obtain manually annotated triplets, which is convenient for evaluating the accuracy of the prediction predicate.
[0026] The CLIP model is used to extract features from the output of the pixel-level filter to obtain text features of the subject and object; and to extract features from image areas of major landform types to obtain image features.
[0027] The multi-layer perception mechanism is used to process image features and text features to obtain a joint embedding vector.
[0028] The subtraction model is used to subtract the joint embedding vector from the subject text features and the object text features to obtain the subtracted features.
[0029] The classification network is used to process the subtracted features to obtain the predicted predicates and finally obtain the predicted triples.
[0030] The generation subunit is used to generate a geographic scene graph based on the predicted triples.
[0031] Step 3: According to the geographic scene graph, the precision, recall and average recall of the geographic scene graph generation unit are obtained, and the geographic scene graph generation unit is optimized according to the precision, recall and average recall to obtain the optimal geographic scene graph generation unit.
[0032] Specifically, if the precision, recall and average recall are lower than the preset values, the segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network are respectively optimized and trained using the corresponding loss functions (existing loss functions) during training to obtain the optimal segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network, and then the optimal geographic scene graph generation unit is obtained.
[0033] Step 4: Input the target image into the optimal geographic scene graph generation unit to obtain the predicted triplet and geographic scene graph corresponding to the target image.
[0034] Specifically, for each ground feature in the predicted triplet, a search bounding box is obtained according to the longitude and latitude of the vertices of the image to which it belongs.
[0035] Preferably, the latitude and longitude of the upper left corner and the latitude and longitude of the lower right corner of the target image are selected to determine a rectangular range as a search boundary box.
[0036] Using the retrieval method based on search bounding box and feature type label, the feature instances corresponding to the subject and object are retrieved in the open source map data OSM, and further extracted to obtain the instance names corresponding to the subject and object.
[0037] Preferably, the bounding box and label retrieval method is selected. If the corresponding feature instance can be retrieved, the feature instance information is obtained in the open source map data to obtain the names of the subject and object in the triple. If the corresponding feature instance cannot be retrieved, stop further information collection and attribute extraction of the current feature, and directly use the subject and object in the predicted triple as graph nodes, feature attributes as attributes of the graph nodes, and predicates as edges, and use the scene graph generator to generate a complete scene graph.
[0038] Use Baidu Encyclopedia crawler to obtain instance information from Baidu Encyclopedia based on instance name.
[0039] By using a large language model (such as the Kimi model), corresponding prompt words are used according to the feature type label to format the attribute extraction of all the information of the instance to obtain the feature attributes of the instance, that is, the feature attributes of the subject and object in the triple.
[0040] Specifically, attribute extraction refers to limiting the attribute types to be extracted in the prompt word according to the feature type label. For example, if the feature type label is road, the attribute types to be extracted are limited to the number of lanes, road width, road function, road speed limit and connected roads.
[0041] The subject and object in the prediction triplet are taken as graph nodes, the attributes of the ground objects are taken as the attributes of the graph nodes, and the predicates are taken as the edges. The scene graph generator is used to generate a complete scene graph.
[0042] Therefore, through the above process, the alignment of the objects in the target image with the open source map data is achieved, and by crawling the encyclopedia data, rich attribute information is obtained. Finally, the predicted triples are combined with the object attributes extracted from the open source map data to generate a complete scene graph.
[0043] The present invention proposes a geographic scene graph generation method based on the CLIP model. The geographic scene graph is generated by combining OSM data through semantic segmentation and multimodal feature fusion. Compared with the traditional remote sensing scene graph generation method based on target detection, the classification of land objects is more complete, the accuracy and reliability of relationship prediction are higher, and the generated geographic scene graph information is richer and more practical, which can provide support for high-level geographic scene understanding tasks.
[0044] Example 2 Based on the above embodiment, an embodiment of the present invention provides a training process of a geographic scene graph generation unit, which specifically includes: For the input training image, the preset pixel-level semantic segmentation model is used to classify the image pixel by pixel, identify different ground object areas in the image, and assign corresponding ground object type labels to each area according to the color value of the pixel to obtain the ground object type label segmentation result.
[0045] Specifically, the training image is a small image generated by cropping a satellite remote sensing image, and the width and height are both 512. The pixel-level semantic segmentation model used in the present invention is a Resnet model, and UPerHead and FCNHead are used as segmentation heads.
[0046] The pixel-level filter calculates the proportion of specific pixels in the entire image, thereby filtering out the main pixel types and mapping them to the object types.
[0047] The main types of land objects are screened out from the training images according to the pixel-level filter, and the labels of these land object types are arranged in pairs in the form of subject and object to generate all possible combinations. Each combination is manually annotated to form a triple representing the predicate relationship between the subject and the object. Each element in the triple represents the subject, predicate and object respectively.
[0048] Understandably, manual labeling of the subject and object combination refers to using professional labeling tools to label the corresponding predicates according to the relationship between the subject and the object in the training image.
[0049] The labels of the feature types are arranged in pairs in the form of subject and object to generate all possible combinations. Each combination is input into the text encoder of the CLIP model, and the text features of the subject and object are obtained through feature extraction.
[0050] According to the combination of feature type labels, the areas of non-feature types in the training image are masked to obtain mask features and joint representation areas. The mask features and joint representation areas are input into the image encoder of the CLIP model, and the image features of the joint representation area are obtained through feature extraction.
[0051] The image features are added to the text features, and the result of the addition is embedded into the text features of the subject and object using a multi-layer perception mechanism to obtain a joint embedding vector. Specifically, the text features and image features need to be adjusted to the same dimension, and then the two can be directly added. The joint embedding vector is input into the text encoder, and the joint embedding features are fitted through a contrastive learning method. The contrastive learning method refers to the contrastive learning of the joint embedding vector and the manually annotated triplets.
[0052] Using the subtraction model, the fitted joint embedding features are subtracted from the text features of the subject and object to obtain the predicate features based on the image features.
[0053] The predicate features based on image features are input into the classification network, and the predicted predicate is calculated through the classification network.
[0054] Finally, the triple generation tool is used to combine the subject, the predicted predicate and the object to obtain the predicted triple.
[0055] Specifically, for each feature type label of the predicted triplet, a search bounding box is obtained according to the longitude and latitude of the vertices of the image to which it belongs.
[0056] Preferably, the latitude and longitude of the upper left corner and the latitude and longitude of the lower right corner of the training image are selected to determine a rectangular range as the search boundary box.
[0057] Using the retrieval method based on search bounding box and feature type label, the corresponding instances of subject and object are retrieved in the open source map data OSM, and further extracted to obtain the instance names corresponding to the subject and object.
[0058] Preferably, the bounding box and label retrieval method is selected. If the corresponding feature instance can be retrieved, the feature instance information is obtained in the open source map data to obtain the names of the subject and object in the triple. If the corresponding feature instance cannot be retrieved, stop further information collection and attribute extraction of the current feature, and directly use the subject and object in the predicted triple as graph nodes, feature attributes as attributes of the graph nodes, and predicates as edges, and use the scene graph generator to generate a complete scene graph.
[0059] Using encyclopedia crawler technology, all the information of the instance is obtained according to the instance name. Using a large language model, according to the feature type label, the corresponding prompt words are used to format the attribute extraction of all the information of the instance to obtain the feature attributes of the instance, that is, the feature attributes of the subject and object in the triple.
[0060] The subject and object in the prediction triplet are taken as graph nodes, the features are taken as attributes of the graph nodes, and the predicates are taken as edges. The Neo4j tool is used to generate a complete geographic scene graph.
[0061] Finally, the precision, recall and average recall of the model are obtained according to the geographic scene graph. If the precision, recall and average recall are lower than the preset values, the segmentation model, CLIP model, multi-layer perception mechanism, contrastive learning model, subtraction model and classification network are trained respectively using their corresponding loss functions to obtain the optimal segmentation model, CLIP model, multi-layer perception mechanism, contrastive learning model, subtraction model and classification network, and then the optimal geographic scene graph generation unit is obtained.
[0062] Example 3 Based on the above embodiments, Figure 3 As shown, an embodiment of the present invention provides a geographic scene graph generation system based on a CLIP model, comprising: The collection module is used to collect geographic images to construct a geographic scene graph dataset.
[0063] The geographic scene graph generation module is used to input the training images in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain the geographic scene graph corresponding to the training images.
[0064] The training module is used to obtain the precision, recall and average recall of the geographic scene graph generation unit according to the geographic scene graph, and optimize the geographic scene graph generation unit according to the precision, recall and average recall to obtain the optimal geographic scene graph generation unit.
[0065] The generation module is used to input the target image into the optimal geographic scene graph generation unit to obtain the prediction triplet and geographic scene graph corresponding to the target image.
[0066] It should be noted that the geographic scene graph generation system based on the CLIP model provided in an embodiment of the present invention is to implement the above-mentioned geographic scene graph generation method based on the CLIP model. Its specific functions can be referred to the above-mentioned method embodiments and will not be repeated here.
[0067] Example 4 Based on the above embodiments, an embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for generating a geographic scene graph based on a CLIP model in the above embodiments is implemented.
[0068] The present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a method for generating a geographic scene graph based on a CLIP model in the above embodiment.
[0069] In summary, the present invention introduces a segmentation model to accurately semantically segment remote sensing images and screen out key features, which has higher accuracy than the traditional method that only relies on bounding box features. The present invention uses the multimodal model CLIP to extract text features and image features, and after feature fusion, applies them to the prediction of predicate relationships, effectively combining the semantic information of text and images, thereby improving the accuracy and reliability of predicate relationship prediction. The present invention combines open source map data to supplement the attributes of features, and uses external data sources to enhance the integrity and accuracy of feature attribute information, thereby improving the generation quality and practicality of scene graphs.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating a geographic scene graph based on a CLIP model, characterized in that: include: Step 1: Collect geographic images to construct a geographic scene graph dataset; Step 2: Input the training image in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain a geographic scene graph corresponding to the training image; Step 3: Obtain the precision, recall and average recall of the geographic scene graph generation unit according to the geographic scene graph, optimize the geographic scene graph generation unit according to the precision, recall and average recall, and obtain the optimal geographic scene graph generation unit; Step 4: Input the target image into the optimal geographic scene graph generation unit to obtain the predicted triplet and geographic scene graph corresponding to the target image.
2. A method for generating a geographic scene graph based on a CLIP model according to claim 1, characterized in that: The geographic scene graph generation unit includes a segmentation model, a pixel-level filter, a CLIP model, a multi-layer perception mechanism, a subtraction model, a classification network and a generation subunit; The segmentation model is used to segment the input image to obtain a ground object type label segmentation result; The pixel-level filter is used to filter out the main types of objects in the training image, and arrange the main object type labels in pairs in the form of subject and object to generate all combinations; The CLIP model is used to extract features from the output of the pixel-level filter to obtain text features of the subject and the object; and to extract features from the image area of the main landform type to obtain image features; The multi-layer perception mechanism is used to process image features and text features to obtain a joint embedding vector; The subtraction model is used to subtract the joint embedding vector from the text features of the subject and the text features of the object to obtain the subtracted features; The classification network is used to process the subtracted features to obtain predicted predicates, and finally obtain predicted triples; The generating subunit is used to generate a geographic scene graph according to the predicted triples.
3. A method for generating a geographic scene graph based on a CLIP model according to claim 2, characterized in that: The segmentation model is used to segment the input image to obtain the ground object type label segmentation result, which specifically includes: The segmentation model is used to classify the input image pixel by pixel, identify different object areas in the image, and assign corresponding object type labels to each area according to the color value of the pixel to obtain the object type label segmentation result.
4. The method for generating a geographic scene graph based on a CLIP model according to claim 2, characterized in that: The extracting of features from the image area of the main ground object type to obtain image features specifically includes: Mask processing is performed on the area of the ground object type not in the combination to obtain mask features and joint representation areas, and the mask features and joint representation areas are input into the CLIP model to obtain image features of the joint representation areas.
5. A method for generating a geographic scene graph based on a CLIP model according to claim 4, characterized in that: The multi-layer perceptron is used to process image features and text features to obtain a joint embedding vector, specifically including: The image features are added to the text features, and the result of the addition is embedded into the text features of the subject and object using a multi-layer perception mechanism to obtain a joint embedding vector.
6. A method for generating a geographic scene graph based on a CLIP model according to claim 2, characterized in that: The optimization of the geographic scene graph generation unit according to the precision rate, the recall rate and the average recall rate specifically includes: If the precision, recall and average recall are lower than the preset values, the segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network are optimized and trained using their corresponding loss functions to obtain the optimal segmentation model, CLIP model, multi-layer perception mechanism, subtraction model and classification network, and then the optimal geographic scene graph generation unit is obtained.
7. A geographic scene graph generation system based on CLIP model, characterized in that: include: A collection module, used to collect geographic images to construct a geographic scene graph dataset; A geographic scene graph generation module is used to input the training image in the geographic scene graph data set into a preset geographic scene graph generation unit to obtain a geographic scene graph corresponding to the training image; A training module, used to obtain the precision, recall and average recall of the geographic scene graph generation unit according to the geographic scene graph, and optimize the geographic scene graph generation unit according to the precision, recall and average recall to obtain the optimal geographic scene graph generation unit; The generation module is used to input the target image into the optimal geographic scene graph generation unit to obtain the prediction triplet and geographic scene graph corresponding to the target image.
8. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for generating a geographic scene graph based on a CLIP model as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a method for generating a geographic scene graph based on a CLIP model as described in any one of claims 1 to 6.