Visual understanding method, device, system, equipment and medium for remote sensing image
By extracting structured information from remote sensing images and generating an image information database, and combining this with a large language model to generate and visualize target prompts, the problems of inaccurate target recognition and insufficient semantic relationships in remote sensing image analysis are solved, thereby improving the accuracy of understanding remote sensing images.
Patent Information
- Application Number
- CN202510045978.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-01-13
Smart Images

Figure CN119474440B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a visual understanding method, device, system, equipment and medium for remote sensing images. BACKGROUND
[0002] With the rapid development of remote sensing technology, especially the wide application of high-resolution satellites and unmanned aerial vehicles and other ground observation equipment, a large amount of remote sensing data is continuously generated. These data are mainly in the form of overhead images and contain rich geographical, environmental and object information, such as city layout, natural resources, transportation network, etc. However, traditional remote sensing image analysis methods usually rely on single modal data such as optical images or SAR (Synthetic Aperture Radar) data. Such single modal methods are prone to information loss and understanding deviation when understanding complex scenes, and are difficult to meet the needs of scene diversification and complex information.
[0003] To solve the above problems, the existing technology mainly uses intelligent question and answer systems (such as large language models, LLM) based on natural language recognition to complete the analysis and understanding of remote sensing images. Large language models have strong language understanding and generation capabilities through large-scale pre-training and can provide high-quality answers in general scenarios. However, there are still some problems in understanding remote sensing images based on large language models, mainly in the following aspects: (1) inaccurate target quantity recognition (2) difficult target orientation and position recognition (3) insufficient semantic relationship and triple recognition.
[0004] At present, there is no effective solution to the problem of low recognition accuracy of remote sensing images in the existing technology. SUMMARY
[0005] Therefore, it is necessary to provide a visual understanding method, device, system, equipment and medium for remote sensing images to solve the above technical problems.
[0006] In a first aspect, the present application provides a visual understanding method for remote sensing images. The method comprises:
[0007] extracting structured information of a remote sensing image to be understood, and generating an image information database corresponding to the remote sensing image to be understood based on the structured information;
[0008] obtaining an initial prompt instruction for the remote sensing image to be understood, determining a search element based on the initial prompt instruction, searching from the image information database based on the search element to obtain corresponding target structured information, and generating a target prompt instruction according to the target structured information and the initial prompt instruction;
[0009] Input the target prompt instruction into the well-trained large language model to obtain an understanding result of the remote sensing image to be understood.
[0010] In one of the embodiments, the structured information of the obtained remote sensing image to be understood is extracted, including:
[0011] Obtain a well-trained image analysis model;
[0012] Input the remote sensing image to be understood into the image analysis model, detect the elements contained in the remote sensing image to be understood through the image analysis model, and generate a structural relationship corresponding to the elements;
[0013] Obtain the structured information corresponding to the remote sensing image to be understood based on the elements and the structural relationship, wherein the elements and the corresponding structural relationship exist in the form of key-value pairs.
[0014] In one of the embodiments, the image information database corresponding to the remote sensing image to be understood is generated based on the structured information, including:
[0015] Obtain a preset text embedding model, and embed and represent the structured information based on the text embedding model to obtain structured vector information corresponding to the structured information, and generate the image information database based on the structured vector information.
[0016] In one of the embodiments, based on the obtained initial prompt instruction for the remote sensing image to be understood, the corresponding target structured information is retrieved from the image information database, including:
[0017] Embed and represent the initial prompt instruction based on the text embedding model to obtain a prompt instruction vector;
[0018] Retrieve the corresponding target structured information from the image information database based on the prompt instruction vector.
[0019] In one of the embodiments, based on the obtained initial prompt instruction for the remote sensing image to be understood, the corresponding target structured information is retrieved from the image information database, including:
[0020] Obtain semantic information of the initial prompt instruction, and calculate a similarity between the semantic information and the structured information;
[0021] According to the similarity calculation result, at least one target structured information is determined from the structured information.
[0022] In one of the embodiments, the target prompt instruction is input into the well-trained large language model to obtain the understanding result of the remote sensing image to be understood, including:
[0023] input the target prompt instruction into the large language model, and generate a visual analysis result corresponding to the target prompt instruction based on the large language model, and perform visual recognition in the to-be-understood remote sensing image based on the visual analysis result.
[0024] In a second aspect, the present application further provides a visual understanding device for remote sensing images. The device comprises:
[0025] The acquisition module is configured to extract structured information of the acquired to-be-understood remote sensing image, and generate an image information database corresponding to the to-be-understood remote sensing image based on the structured information.
[0026] The computing module is configured to acquire an initial prompt instruction for the to-be-understood remote sensing image, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, and generate a target prompt instruction according to the target structured information and the initial prompt instruction.
[0027] The generation module is configured to input the target prompt instruction into the trained large language model, and obtain an understanding result of the to-be-understood remote sensing image.
[0028] In a third aspect, the present application further provides a visual understanding system for remote sensing images. The system comprises an image analysis module, a database construction module, a retrieval module, a question and answer generation module, and a visualization module.
[0029] The image analysis module is configured to extract structured information of the acquired to-be-understood remote sensing image, and send the structured information to the database construction module.
[0030] The database construction module is configured to integrate the structured information, and establish an image information database corresponding to the to-be-understood remote sensing image.
[0031] The retrieval module is configured to acquire an initial prompt instruction, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, combine the target structured information with the initial prompt instruction to obtain a target prompt instruction, and send the target prompt instruction to the question and answer generation module.
[0032] The question and answer generation module is configured to generate a visual analysis result based on the target prompt instruction.
[0033] The visualization module is configured to perform visual recognition in the to-be-understood remote sensing image based on the visual analysis result.
[0034] In a fourth aspect, the present application further provides a computer device. The computer device comprises a memory and an identifier. The memory stores a computer program. When the identifier executes the computer program, the following steps are implemented.
[0035] extract structured information of the remote sensing image to be understood, and generate an image information database corresponding to the remote sensing image to be understood based on the structured information;
[0036] obtain an initial prompt instruction for the remote sensing image to be understood, determine a to-be-retrieved element based on the initial prompt instruction, retrieve from the image information database based on the to-be-retrieved element, obtain corresponding target structured information, and generate a target prompt instruction according to the target structured information and the initial prompt instruction;
[0037] input the target prompt instruction into the well-trained large language model to obtain an understanding result of the remote sensing image to be understood.
[0038] In a fifth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium has a computer program stored thereon, and the computer program is executed by an identifier to implement the following steps:
[0039] extract structured information of the remote sensing image to be understood, and generate an image information database corresponding to the remote sensing image to be understood based on the structured information;
[0040] obtain an initial prompt instruction for the remote sensing image to be understood, determine a to-be-retrieved element based on the initial prompt instruction, retrieve from the image information database based on the to-be-retrieved element, obtain corresponding target structured information, and generate a target prompt instruction according to the target structured information and the initial prompt instruction;
[0041] input the target prompt instruction into the well-trained large language model to obtain an understanding result of the remote sensing image to be understood.
[0042] The above-mentioned remote sensing image-oriented visual understanding method, device, system, equipment and medium first obtain structured information of a remote sensing image to be understood, and generate an image information database according to the structured information; then obtain an initial prompt instruction, determine a to-be-retrieved element according to the initial prompt instruction, retrieve from the image information database based on the to-be-retrieved element, obtain target structured information, generate a target prompt instruction according to the target structured information and the initial prompt instruction; and input the target prompt instruction into a well-trained large language model to obtain a corresponding understanding result. The present application can greatly improve the accuracy of understanding of a remote sensing image with more elements and complex scenes. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 a flowchart of a visual understanding method in one embodiment;
[0044] Figure 2 a data flow diagram of a visual understanding method in one embodiment;
[0045] Figure 3 This is a structural block diagram of a visual understanding device in one embodiment;
[0046] Figure 4 This is a block diagram of a visual understanding system in one embodiment;
[0047] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0049] In one embodiment, such as Figure 1 As shown, a visual understanding method for remote sensing images is provided, including the following steps:
[0050] Step S110: Extract the structured information of the remote sensing image to be understood, and generate an image information database corresponding to the remote sensing image to be understood based on the structured information.
[0051] Specifically, the structured information of the remote sensing image to be understood is extracted. This structured information represents the elements contained in the remote sensing image and the corresponding related information of the elements. That is, the structured information includes various types of elements in the remote sensing image and the structural information corresponding to each element, such as the quantity of each element, its specific location in the image, and the relationship between the element and other elements. In some preferred embodiments, after obtaining the structured information of the remote sensing image to be understood, the format of the structured information can be organized. The following provides a format for structured information:
[0052] "Object Category": {Quantity: The specific number of objects of the current category in the image;}
[0053] Location: The specific position of each object in the current category in the image, such as left, right, bottom right, etc.
[0054] Relationship: The relationship between objects of the current category and other objects.
[0055] This application includes, but is not limited to, the aforementioned format of structured information. It is understood that a single remote sensing image to be understood should be able to extract structured information for multiple object categories.
[0056] generate an image information database corresponding to the remote sensing image to be understood based on the structured information, wherein one remote sensing image to be understood corresponds to one image information database, and the image information database includes all structured information corresponding to the remote sensing image to be understood. In some preferred embodiments, the structured information of each different object category can be divided into different text blocks for storage.
[0057] In step S120, an initial prompt instruction for the remote sensing image to be understood is obtained, the to-be-retrieved element is determined based on the initial prompt instruction, the corresponding target structured information is obtained by retrieving the image information database based on the to-be-retrieved element, and the target prompt instruction is generated according to the target structured information and the initial prompt instruction.
[0058] Specifically, the initial prompt instruction for the remote sensing image to be understood is obtained, wherein the initial prompt instruction can be an image understanding instruction input by a user, such as “query the number and position of all vehicles in the image”, or a template prompt corresponding to a large language model, such as:
[0059] “You are a professional remote sensing image question and answer intelligent assistant. Please answer the user's question based on the structured information in the following image. Please ensure that the answer is accurate and comprehensive. If you do not know, please answer that you do not know, and do not fake information.
[0060] User question: {user question}”.
[0061] After obtaining the initial prompt instruction, the to-be-retrieved element is determined according to the initial prompt instruction. In this embodiment, the semantic information in the initial prompt instruction can be captured based on a preset large language model, so as to determine the to-be-retrieved element in the initial prompt instruction, wherein the large language model can use existing general-purpose question answering systems such as deepseek. After the to-be-retrieved element is determined, the corresponding target structured information is obtained by retrieving the image information database corresponding to the remote sensing image to be understood based on the to-be-retrieved element. In some preferred embodiments, the similarity between the to-be-retrieved element and the structured information in the image information database can be calculated, and several structured information with high similarity can be selected as the target structured information, such as four structured information with the highest similarity can be selected as the target structured information, and the RAG algorithm (Retrieval-Augmented Generation) can be used to complete the retrieval of the target structured information. After the target structured information is retrieved, the target structured information and the initial prompt instruction can be combined based on the large language model to generate the target prompt instruction, such as the target prompt instruction can be:
[0062] "You are a professional remote sensing image question and answer intelligent assistant, please answer the user's question based on the structured information in the following image. Please ensure that the answer is accurate and comprehensive, if you do not know, please answer do not know, do not fake information.
[0063] User question: {user question}
[0064] Structured information: {object 1 target structured information}, {object 2 target structured information}…{object 4 target structured information}".
[0065] The above is a template for a target prompt instruction, and other templates can also be selected for implementation in actual applications. The present application does not make too many limitations on this. Further, the above image information database can be established based on any storage method convenient for retrieving and querying structured information, so the above target structured information can be in the form of text, or in the form of vector, etc.
[0066] Step S130, input the target prompt instruction into the trained large language model to obtain the understanding result of the remote sensing image to be understood.
[0067] Specifically, in actual application, the retrieved target structured information is structured information with high similarity to the element to be retrieved, but it is not necessarily the reply result of the initial prompt instruction. For example, the initial prompt instruction can be to query the number of vehicles in the remote sensing image to be understood, and then the corresponding retrieved target structured information can be the orientation of each vehicle, the relationship between each vehicle and other objects, and the number of vehicles, etc. At this time, the target structured information and the initial prompt instruction can be integrated to obtain the target prompt instruction and input the large language model for understanding, and the large language model can answer the prompt instruction input by the user according to the retrieved target structured information, such as replying the number of vehicles in the remote sensing image to be understood. And further, in some embodiments, the retrieved target structured information can be in the form of vector, and the user is difficult to directly read the meaning of the target structured information in the form of vector, so the large language model is also needed to convert the target prompt instruction, and output the natural language that can be directly read by the user.
[0068] Further, the above understanding result includes but is not limited to the reply result output by the large language model corresponding to the target prompt instruction, and according to the need, the remote sensing image to be understood can be understood based on the target prompt instruction, such as labeling the element to be retrieved in the image, segmenting the region corresponding to the target prompt instruction, etc.
[0069] Through steps S110 to S130, the structured information of the remote sensing image to be understood is extracted. Especially when facing a complex scene (such as the connection of buildings and roads, the interaction of vegetation and water bodies, etc.), various data such as the number, position, and context relationship of elements in the remote sensing image to be understood can be comprehensively extracted through the present application. Further, after obtaining the initial prompt instruction, the corresponding target structured information is retrieved from the image information database, and then the large language model generates the corresponding reply result and the understanding result of the remote sensing image to be understood based on the target structured information and the initial prompt instruction, thereby greatly improving the accuracy of image understanding of images with more elements and complex scenes.
[0070] In some embodiments, the above method further comprises:
[0071] obtaining a trained image analysis model;
[0072] inputting the remote sensing image to be understood into the image analysis model, detecting the elements contained in the remote sensing image to be understood through the image analysis model, and generating the structural relationship corresponding to the elements;
[0073] obtaining the structured information corresponding to the remote sensing image to be understood based on the elements and the structural relationship, wherein the elements and the corresponding structural relationship exist in the form of key-value pairs.
[0074] Specifically, the image analysis model is obtained. In some preferred embodiments, the image analysis model can be an SGG (Scene Graph Generation) model, including but not limited to a PENET model, etc. The training of the image analysis model can be as follows: obtaining an initial image analysis model, and obtaining a remote sensing image training set. The remote sensing image can use an open source remote sensing image relationship dataset STAR, and corresponding annotation boxes, relationship triples, etc. In some preferred embodiments, before training the initial image analysis model based on the remote sensing image training set, the remote sensing image training set can be pre-identified, such as cropping, cleaning, denoising, and deduplication of the images in the training set, etc. Then, the initial image analysis model is trained based on the pre-identified remote sensing image training set to obtain the trained image analysis model.
[0075] The above remote sensing image to be understood is input into the image analysis model, and the elements contained in the remote sensing image to be understood and the structural relationship corresponding to each element are output by the image analysis model. The structural relationship includes but is not limited to the coordinates of the elements, the relationship between the elements and other elements, etc.
[0076] Finally, the above elements and corresponding structural relationships are integrated to obtain the above structured information. In some preferred embodiments, the structured information can be saved in the form of a dictionary, for example, setting the category or keyword of each element as the key, and the corresponding value as the quantity, position, and relationship between other objects of the element, and the like.
[0077] The structured information of the remote sensing image can be more comprehensively and accurately extracted through the embodiment.
[0078] In some embodiments, the above method further comprises:
[0079] A preset text embedding model is obtained, and the structured information is embedded and represented based on the text embedding model to obtain structured vector information corresponding to the structured information, and an image information database is generated based on the structured vector information.
[0080] Specifically, a preset text embedding model is obtained, which can use text2vec-base-Chinese, text2vec-large-Chinese, or the like. The structured information is embedded and represented based on the text embedding model to generate a low-dimensional vector representation as a semantic representation of the structured information, thereby obtaining structured vector information. After obtaining the structured vector information, a high-efficiency vector index can be established using the Milvus library. In some preferred embodiments, when the structured information is saved in the form of key-value, the data corresponding to the key can be converted into a vector, the data corresponding to the value can be directly saved into the database, and the data corresponding to the key and the value can be processed in one-to-one binding, thereby completing the embedding representation of the structured information.
[0081] In the embodiment, the structured information is embedded and represented to obtain a low-dimensional vector representation corresponding to the structured information, thereby more efficiently capturing the semantic information of the query, and reducing the storage space occupied by the structured information.
[0082] In some embodiments, the above method further comprises:
[0083] The initial prompt instruction is embedded and represented based on the text embedding model to obtain a prompt instruction vector.
[0084] The corresponding target structured information is retrieved from the image information database based on the prompt instruction vector.
[0085] Specifically, when the structured information in the image information database is saved in the form of a vector, i.e., in the form of structured vector information, in some preferred embodiments, the obtained initial prompt instruction can also be embedded and represented, and the same text embedding model as that used for embedding and understanding the structured vector is used to ensure that the vector of the initial prompt instruction provided by the user and the vector of the structured data in the database are in the same semantic space. Based on the text embedding model, the initial prompt instruction input by the user can be converted into a dense vector representation, thereby effectively capturing the semantic information of the query. In some preferred embodiments, to improve computational efficiency and facilitate subsequent retrieval and identification, the initial prompt instruction input by the user and the structured information stored in the database can be converted into the same form, such as both being in the form of text or both being in the form of a vector.
[0086] In some embodiments, the above method further comprises:
[0087] obtaining semantic information of the initial prompt instruction and calculating the similarity between the semantic information and the structured information;
[0088] determining at least one target structured information from the structured information according to the similarity calculation result.
[0089] Specifically, the semantic information of the initial prompt instruction is obtained, and the method for extracting the semantic information includes but is not limited to obtaining the semantic features in the initial prompt instruction based on a preset semantic extraction model, such as the text2vec-base-Chinese, text2vec-large-Chinese model. It should be noted that the semantic extraction model used to extract the semantic features of the initial prompt instruction in this embodiment is consistent with the text embedding model described above, i.e., the semantic extraction model used at this time is the text embedding model described above, which converts the initial prompt instruction into a dense vector representation, thereby obtaining the above semantic information. After obtaining the semantic information, the similarity between the semantic information and the structured information in the database is calculated. The similarity between the semantic information of the initial prompt instruction and the structured information can be calculated based on the Milvus index in the database, the key vector with similar semantics to the query in the database is identified, and according to the similarity calculation result, a plurality of target structured information is determined, i.e., the structured information corresponding to the key is determined. In some preferred embodiments, four target structured information can be obtained from the structured information, i.e., four target structured information most similar to the initial prompt instruction can be obtained from the structured information. In actual applications, the number of target structured information obtained can be set by a person skilled in the art. It can be understood that if the number of target structured information obtained is too small, the corresponding results of the prompt instruction may not be comprehensively retrieved, and if the number of target structured information obtained is too large, the calculation speed will decrease. Therefore, in actual applications, the calculation speed and effect can be considered to select the most balanced setting.
[0090] In some embodiments, the method further comprises:
[0091] The target prompt instruction is input into the large language model, and a visual analysis result corresponding to the target prompt instruction is generated based on the large language model, and visual recognition is performed in the to-be-understood remote sensing image based on the visual analysis result.
[0092] Specifically, the target prompt instruction includes an initial prompt instruction and corresponding target structured information. In actual application, the target structured information and the initial prompt instruction can be fused into a prompt that is logically coherent and semantically consistent. The target prompt instruction is input into the large language model, so that a high-quality answer can be generated by the large language model, i.e., the visual analysis result is generated. Then, according to the need, the relevant structured information can be visualized in the original to-be-understood remote sensing image. For example, if the initial prompt instruction input by the user is to query the number of vehicles in the image, the output visual analysis result can be "the number of vehicles is: 36", and the positions of each vehicle in the corresponding to-be-understood remote sensing image are labeled with a bounding box.
[0093] The application also provides a preferred embodiment of a visual understanding method for a remote sensing image, Figure 2 A data flow diagram of the visual understanding method in one preferred embodiment.
[0094] The prior art is mostly an intelligent question and answer system (such as a large language model, LLM) that uses natural language recognition to complete the analysis and recognition of remote sensing images. The large language model has strong language understanding and generation capabilities through large-scale pre-training, and can provide high-quality answers in general scenarios. However, remote sensing images have many problems, such as a large variety of elements in the image, uneven density, often covering a large area, complex semantic relationships, lack of clear context references, etc. Directly understanding remote sensing images through a large language model usually has problems such as low accuracy and incomplete recognition.
[0095] In the present application, first, the user terminal 21 detects the remote sensing image to be understood through the preset image analysis model 22, generates structured information, and establishes an image information database corresponding to the remote sensing image to be understood based on the structured information. Secondly, the user terminal 21 inputs the image information database and the preset initial prompt instruction into the text embedding model 23, and the text embedding model 23 embeds and represents the image information database and the initial prompt instruction, respectively, to obtain an image information database composed of low-dimensional structured vector information and a prompt instruction vector. Then, the user terminal 21 retrieves a plurality of target structured information with the highest similarity in the image information database composed of structured vector information based on the prompt instruction vector, combines the target structured information with the initial prompt instruction to obtain a target prompt instruction. Finally, the target prompt instruction is input into the large language model 24, and the large language model 24 outputs a visual analysis result with high understanding and a visual recognition result, and feeds back the visual analysis result and the visual recognition result to the user terminal 21.
[0096] It should be understood that, although each step in the flowchart involved in each of the above-described embodiments is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above-described embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or stages.
[0097] Based on the same inventive concept, the present application also provides a visual understanding device for implementing the above-mentioned visual understanding method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more visual understanding device embodiments provided below can refer to the limitations of the visual understanding method described above, which will not be repeated here.
[0098] In one embodiment, as shown in Figure 3 A visual understanding device is provided, comprising: an acquisition module 31, a calculation module 32 and a generation module 33, wherein:
[0099] The acquisition module 31 is configured to extract structured information of the acquired remote sensing image to be understood, and generate an image information database corresponding to the remote sensing image to be understood based on the structured information;
[0100] The computing module 32 is configured to obtain an initial prompt instruction for the remote sensing image to be understood, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, and generate a target prompt instruction according to the target structured information and the initial prompt instruction.
[0101] The generating module 33 is configured to input the target prompt instruction into the trained complete large language model to obtain an understanding result of the remote sensing image to be understood.
[0102] The above-mentioned various modules in the visual understanding device can be realized by software, hardware, and a combination thereof in whole or in part. The above-mentioned various modules can be embedded in or independent of an identifier in the computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so as to be called and executed by the identifier to perform the operations corresponding to the above-mentioned various modules.
[0103] In one embodiment, as shown in Figure 4 a visual understanding system is provided, comprising:
[0104] an image analysis module, a database construction module, a retrieval module, a question and answer generation module, and a visualization module;
[0105] The image analysis module is configured to extract structured information of the obtained remote sensing image to be understood, and send the structured information to the database construction module.
[0106] The database construction module is configured to integrate the structured information, and establish an image information database corresponding to the remote sensing image to be understood.
[0107] The retrieval module is configured to obtain an initial prompt instruction, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, combine the target structured information with the initial prompt instruction to obtain a target prompt instruction, and send the target prompt instruction to the question and answer generation module.
[0108] The question and answer generation module is configured to generate a visual analysis result based on the target prompt instruction.
[0109] The visualization module is configured to perform visual recognition in the remote sensing image to be understood based on the visual analysis result.
[0110] Specifically, the image analysis module can extract the structured information of the remote sensing image to be understood, thereby providing help for accurate answering and generation of the subsequent large language model.
[0111] The above-mentioned database construction module can recognize the structured data extracted by the image analysis module into a high-dimensional vector using a text embedding model, and then construct the structured data into a searchable vector database.
[0112] The retrieval module is responsible for implementing retrieval-augmented generation technology, receiving an initial prompt instruction input by a user, then converting the initial prompt instruction into a vector form using the text embedding model, and retrieving the most similar target structured information in the database, thereby providing rich context information for subsequent large model generation.
[0113] The question and answer generation module is used to generate a final visual analysis result, which is an accurate answer corresponding to the initial prompt instruction.
[0114] The visualization module is responsible for visualizing the output result corresponding to the large language on the remote sensing image to be understood, such as labeling the categories of corresponding elements and labeling boxes, etc.
[0115] Through the modular design described above, the system has high flexibility and scalability. Each module can be upgraded or replaced independently without affecting the overall system function, such as updating the large language model, improving the retrieval module, and improving the image analysis module, to adapt to the development of technology and demand.
[0116] In one embodiment, a computer device, which can be a server, has an internal structure diagram as shown in Figure 5 The computer device includes an identifier, a memory, and a network interface connected by a system bus. The identifier of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store visual understanding related data. The network interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the identifier to implement a visual understanding method.
[0117] Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0118] In one embodiment, a computer readable storage medium having a computer program stored thereon is provided. The computer program is executed by the identifier to implement the following steps:
[0119] extracting structured information of a remote sensing image to be understood, and generating an image information database corresponding to the remote sensing image to be understood based on the structured information;
[0120] obtaining an initial prompt instruction for the remote sensing image to be understood, determining a to-be-retrieved element based on the initial prompt instruction, retrieving from the image information database based on the to-be-retrieved element to obtain corresponding target structured information, and generating a target prompt instruction according to the target structured information and the initial prompt instruction;
[0121] inputting the target prompt instruction into the trained complete large language model to obtain an understanding result of the remote sensing image to be understood.
[0122] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are information and data authorized by the user or fully authorized by all parties.
[0123] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The identifier involved in the embodiments provided in the present application can be a general identifier, a central identifier, a graphical identifier, a digital signal identifier, a programmable logic device, a data recognition logic device based on quantum computing, etc., without being limited thereto.
[0124] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0125] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A visual understanding method for remote sensing images, characterized in that, The method comprises: extracting structured information of a remote sensing image to be understood, and generating an image information database corresponding to the remote sensing image to be understood based on the structured information; wherein the structured information comprises the number of each element in the remote sensing image to be understood, the specific position of each element in the image, and the relationship between elements and other elements; obtaining an initial prompt instruction for the remote sensing image to be understood, determining a to-be-retrieved element based on the initial prompt instruction, retrieving the corresponding target structured information from the image information database based on the to-be-retrieved element, and generating a target prompt instruction according to the target structured information and the initial prompt instruction; wherein retrieving the target structured information comprises obtaining semantic information of the initial prompt instruction, and calculating the similarity between the semantic information and the structured information; according to the similarity calculation result, at least one target structured information is determined from the structured information; inputting the target prompt instruction into a well-trained large language model to obtain an understanding result of the remote sensing image to be understood, comprising that the large language model answers the initial prompt instruction according to the target structured information.
2. The method of claim 1, wherein, The extracted structured information of the obtained remote sensing image to be understood comprises: obtaining a well-trained image analysis model; inputting the remote sensing image to be understood into the image analysis model, detecting the elements contained in the remote sensing image to be understood by the image analysis model, and generating the structural relationship corresponding to the elements; based on the elements and the structural relationship, the structured information corresponding to the remote sensing image to be understood is obtained, wherein the elements and the corresponding structural relationship exist in the form of key-value pairs.
3. The method of claim 1, wherein, The generation of the image information database corresponding to the remote sensing image to be understood based on the structured information comprises: obtaining a preset text embedding model, embedding and representing the structured information based on the text embedding model to obtain structured vector information corresponding to the structured information, and generating the image information database based on the structured vector information.
4. The method of claim 3, wherein, The retrieval of the corresponding target structured information from the image information database based on the obtained initial prompt instruction for the remote sensing image to be understood comprises: embedding and representing the initial prompt instruction based on the text embedding model to obtain a prompt instruction vector; retrieving the corresponding target structured information from the image information database based on the prompt instruction vector.
5. The method of claim 1, wherein, The input of the target prompt instruction into the well-trained large language model to obtain the understanding result of the remote sensing image to be understood comprises: inputting the target prompt instruction into the large language model, generating a visual analysis result corresponding to the target prompt instruction based on the large language model, and performing visual recognition in the remote sensing image to be understood based on the visual analysis result.
6. A device for visual understanding of a remote sensing image, characterized in that, The device comprises: An acquisition module is configured to extract structured information of a to-be-understood remote sensing image and generate an image information database corresponding to the to-be-understood remote sensing image based on the structured information; the structured information includes the number of elements in the to-be-understood remote sensing image, the specific position of each element in the image, and the relationship between elements and other elements. A calculation module is configured to acquire an initial prompt instruction for the to-be-understood remote sensing image, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, and generate a target prompt instruction according to the target structured information and the initial prompt instruction; the retrieval of the target structured information includes acquiring semantic information of the initial prompt instruction and calculating the similarity between the semantic information and the structured information; at least one target structured information is determined from the structured information according to the similarity calculation result. A generation module is configured to input the target prompt instruction into a trained large language model to obtain an understanding result of the to-be-understood remote sensing image, including that the large language model answers the initial prompt instruction according to the target structured information.
7. A visual understanding system for remote sensing images, characterized in that, The system includes an image analysis module, a database construction module, a retrieval module, a question and answer generation module, and a visualization module. The image analysis module is configured to extract structured information of a to-be-understood remote sensing image and send the structured information to the database construction module; the structured information includes the number of elements in the to-be-understood remote sensing image, the specific position of each element in the image, and the relationship between elements and other elements. The database construction module is configured to integrate the structured information and establish an image information database corresponding to the to-be-understood remote sensing image. The retrieval module is configured to acquire an initial prompt instruction, determine a to-be-retrieved element based on the initial prompt instruction, retrieve corresponding target structured information from the image information database based on the to-be-retrieved element, combine the target structured information with the initial prompt instruction to obtain a target prompt instruction, and send the target prompt instruction to the question and answer generation module; the retrieval of the target structured information includes acquiring semantic information of the initial prompt instruction and calculating the similarity between the semantic information and the structured information; at least one target structured information is determined from the structured information according to the similarity calculation result. The question and answer generation module is configured to generate a visual analysis result based on the target prompt instruction; including that a large language model answers the initial prompt instruction according to the target structured information. The visualization module is configured to perform visual recognition in the to-be-understood remote sensing image based on the visual analysis result.
8. A computer device comprising a memory and an identifier, the memory storing a computer program, characterized in that, The identifier implements the steps of the method of any one of claims 1 to 5 when executing the computer program.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the recognizer, implements the steps of the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Question and answer method and device, equipment, medium and product
CN117216210A
Interactive target statistical analysis method, device, equipment, medium and product
CN119025571A