Remote sensing intelligent interpretation system based on large language model
Through a remote sensing intelligent interpretation system based on large language models, a variety of remote sensing image processing tools and professional knowledge bases are integrated, which solves the problem that ordinary users find it difficult to interpret remote sensing data independently, and realizes efficient processing and intelligent interpretation of complex tasks.
Patent Information
- Application Number
- CN202510514059.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-25
AI Technical Summary
The existing remote sensing data interpretation methods rely on professional and technical personnel, which are difficult for ordinary users to complete independently, and large language models and visual language models cannot call complex tools for in-depth analysis in complex remote sensing tasks, lack professional field knowledge, and it is difficult to meet actual needs.
Design a remote sensing intelligent interpretation system based on large language models, integrating user input interface, central controller, solution searcher, professional data searcher, tool library and output module, using RAG technology and FAISS algorithm, combining professional knowledge base and a variety of remote sensing image processing tools, supports multiple rounds of dialogue and precise understanding of user intentions.
It enables ordinary users to independently handle complex remote sensing tasks, provide personalized and accurate feedback, improves the automation and intelligence level of remote sensing data processing, and can effectively deal with diversified remote sensing problems.
Smart Images

Figure CN120375201A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agents and remote sensing technology, and particularly relates to a remote sensing intelligent interpretation system based on a large language model. Background Art
[0002] In the current era of rapid technological development, significant progress has been made in agent technology and related research. In the early days, AI agents mainly included symbolic and reactive types. Symbolic agents encapsulate knowledge and perform reasoning by means of logical rules and symbolic representations, while reactive agents focus on the mapping relationship between input and output with the main goal of rapid response. With the evolution of technology, agents based on reinforcement learning and transfer learning have gradually emerged. Agents based on reinforcement learning optimize their own strategies through continuous interaction with the environment to maximize cumulative rewards, and achievements like AlphaGo are typical representatives in this field; agents based on transfer learning accelerate the learning process of new tasks by transferring knowledge from the source domain to the target domain.
[0003] In recent years, the booming development of large language models (LLMs) has injected new vitality into the field of agents, giving rise to agents based on LLMs such as Visual ChatGPT, WorldGPT, HuggingGPT, and TreeGPT. These agents use LLMs as the core, can accurately understand user input, then formulate reasonable plans and solutions, and efficiently complete various complex tasks with the help of a tool chain. For example, Visual ChatGPT has successfully achieved visual question - answering functions through structured prompt interaction with ChatGPT and collaborative cooperation among multiple models.
[0004] However, in the field of remote sensing data interpretation, existing methods have many deficiencies. Traditional remote sensing algorithms are usually designed for specific application scenarios, and their deployment and operation often rely on the operation of professional technicians, which makes it difficult for ordinary users to independently interpret remote sensing data, greatly limiting the wide application of remote sensing data. Although large language models (LLMs) and visual language models (VLMs) have shown certain application potential in remote sensing tasks, currently they are mainly limited to basic visual and language instruction adjustment tasks. When faced with complex remote sensing tasks, the limitations of these models become prominent. They cannot call complex tools for in - depth analysis and lack professional domain knowledge, and it is difficult to give satisfactory results when dealing with tasks that require in - depth professional knowledge.
[0005] In this context, remote sensing agents have shown great potential. Traditional remote sensing algorithms are designed for specific applications, and the deployment process requires the participation of professional technicians. It is very difficult for ordinary users to independently complete the interpretation of remote sensing data, which severely restricts the popularization of remote sensing technology and the expansion of its application scope. At the same time, although large language models (LLMs) and vision language models (VLMs) have been applied in remote sensing tasks, they can currently only handle basic visual and language instruction adjustment tasks. Facing complex remote sensing tasks, they can neither call complex tools for in-depth analysis nor have professional domain knowledge reserves, making it difficult to meet actual needs.
[0006] In this predicament, the emergence of remote sensing agents brings hope for solving these problems. Remote sensing agents have the ability to integrate multiple high-performance remote sensing image processing models, which enables them to flexibly select appropriate tools according to different problems and effectively handle diverse task requirements. Their multi-tool integration feature provides richer means for solving complex remote sensing problems and greatly enhances the possibility of handling complex tasks.
[0007] Remote sensing agents also support multi-round conversations and can understand the conversation context. This interactive advantage helps to better obtain user intentions, provide services that are more in line with user needs, make the remote sensing interpretation process more intelligent and user-friendly, improve the interaction efficiency between users and the remote sensing interpretation system, enhance the user experience, and lay a foundation for wide application. By introducing Retrieval-Augmented Generation (RAG) technology, remote sensing agents can quickly retrieve information from a professional knowledge base. This enables them to obtain professional knowledge support when dealing with problems involving professional data and complex processes, provides knowledge guarantee for solving complex remote sensing problems, and also provides a powerful tool for in-depth remote sensing research.
[0008] It can be seen that in the face of numerous difficulties in existing remote sensing intelligent interpretation methods, the advantages of remote sensing agents in tool integration, interaction methods, and professional knowledge processing fully demonstrate their huge development potential and the feasibility of research, and are expected to bring innovative breakthroughs and changes to the remote sensing field. Summary of the Invention
[0009] Aiming at the deficiencies of the existing technology, the present invention proposes a remote sensing intelligent interpretation system based on a large language model, which includes: a user input interface, a central controller, a solution retriever, a professional data retriever, a tool library, and an output module;
[0010] The user input interface is used to receive user input information;
[0011] The solution retriever is used to retrieve solutions and input the retrieved solutions into the central controller;
[0012] The professional data retriever is used to retrieve professional knowledge related to the user input information and input the retrieval results into the central controller;
[0013] The tool library stores a variety of remote sensing image processing tools;
[0014] The central controller understands and analyzes the user input information, infers the user's intention, generates a task type, and sends the task type to the solution retriever; selects a suitable tool from the tool library according to the solution to process the user input information. If the user's question involves professional knowledge, it generates query keywords and sends them to the professional data retriever, and generates a final response in combination with the professional knowledge retrieval results;
[0015] The output module outputs the final answer and the processed remote sensing image according to the processing result.
[0016] Preferably, the solution retriever is built based on the RAG technology and is used to assist the central controller in selecting suitable computing tools; in the retrieval stage, the FAISS algorithm is used to retrieve the top k solution documents related to the user input information, and the k solution documents are sent to the central controller and added to the prompt template to enhance the central controller's ability to select suitable tools.
[0017] Preferably, the professional data retriever is built based on the RAG technology, uses the FAISS algorithm to retrieve the top k professional knowledge documents related to the user input information, and sends the k professional knowledge documents to the central controller and adds them to the prompt template to ensure high accuracy of the response.
[0018] Preferably, the remote sensing image processing tools in the tool library include disaster assessment, object detection, scene classification, change detection, image segmentation, aircraft type recognition tools, and functional tools based on other models.
[0019] Preferably, the central controller is built based on a large language model.
[0020] Preferably, the process of the output module outputting the final answer or the remote sensing image according to the processing result includes: judging whether the processing result contains image output. If it contains, the image file path is mapped to an accessible Web URL and sent to the front end together with the text output. The front end determines whether to display the image on the Web interface by judging whether the output contains a URL.
[0021] Preferably, it further includes a RAG knowledge base, which consists of a solution database and a professional data knowledge base; the solution database stores a large number of solutions for different remote sensing tasks to help the LLM select appropriate tools; the professional data knowledge base stores rich remote sensing professional knowledge to provide accurate answers for complex professional questions.
[0022] The beneficial effects of the present invention are as follows: The present invention uses the LLM as the central controller to accurately understand and interpret user intentions; integrates multiple high-performance remote sensing image processing models to handle various complex problems such as single or diverse ones, and it also supports multi-round conversations to ensure consistency and context understanding throughout the interaction process, enabling users to interact deeply and obtain more personalized and accurate feedback; in addition, the present invention also adopts RAG technology, which greatly expands its knowledge reserve by integrating professional knowledge bases, can effectively handle specific problems related to remote sensing, and ensures that the system can still provide efficient and accurate solutions in complex tasks; compared with the prior art, the present invention can understand user intentions more accurately, efficiently process complex remote sensing tasks, perform excellently in tasks such as scene classification, visual question answering, and target counting, and significantly improves the automation and intelligence level of remote sensing data processing. Brief Description of the Drawings
[0023] Figure 1 It is the model architecture diagram of the remote sensing intelligent interpretation system based on the large language model in the present invention;
[0024] Figure 2 It is the workflow diagram of the remote sensing intelligent interpretation system based on the large language model in the present invention. Detailed Embodiments
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] To address the challenge of intelligent processing of massive remote sensing data, the present invention proposes a remote sensing intelligent interpretation system (RS-Agent) based on a large language model. RS-Agent is an intelligent tool set integrating a large language model, and its model architecture is as shown in the appendix Figure 1 and includes: a user input interface, a central controller, a solution retriever, a professional data retriever, a tool library, and an output module. Among them, the central controller is built based on the large language model. The system also includes a RAG knowledge base.
[0027] The RAG knowledge base consists of a solution database and a professional materials knowledge base. The solution database stores a large number of solutions for different remote sensing tasks to help the LLM select appropriate tools; the professional materials knowledge base contains rich remote sensing expertise to provide accurate answers to complex professional questions and make up for the lack of knowledge of the LLM in the professional field.
[0028] When the user inputs a query Q and a remote sensing image I, the system starts to process the task. The central controller (LLM) first understands and analyzes the user input, tries to infer the user's intention, and infers the task type r of the query according to the user intention. s To assist the central controller in selecting appropriate solutions, the solution retriever built by the RAG technology in the system retrieves relevant solution guides g s from the solution database according to the task type r s , and these guiding information helps the central controller select appropriate tools from the rich tool space T If the user query involves knowledge retrieval, the central controller calls the professional materials retriever and, according to the keywords k s extracted from the user query by the central controller, obtains knowledge guides g s from the professional materials knowledge base. s Finally, the central controller calls the selected tool according to g k , processes the input data in combination with g and integrates the processing results, and generates the final answer A and the processed remote sensing image through the output module to
[0029] Complete the task processing flow.
[0030] User input interface: used to receive user input information, that is, the user query, including instructions and pictures.
[0031] Central controller M C : As the core control unit of the RS-Agent, the LLM brain plays the role of the "central controller". Qwen 2.5Taking - 32b as an example, it can, through training on a large - scale corpus, understand user inputs in various natural language forms, accurately analyze the context and subtle semantic differences, thereby accurately infer the user's intentions and needs, generate the task type and query keywords, and send the task type and query keywords to the solution retriever and the professional material retriever respectively. When facing complex remote - sensing tasks, the LLM can, according to the retrieved solution guidance, formulate a detailed processing plan, command various tools in the tool space to work together, and ensure the effective execution of the task. When the selected tool involves knowledge retrieval, it means that the user's question involves professional knowledge and additional knowledge is needed to solve the user query Q. At this time, the retrieval result of the professional material retriever, that is, the knowledge guidance g, needs to be combined. k Generate the processing result.
[0032] The Retrieval - Augmented Generation (RAG) technology enhances the output ability of the language model by combining information retrieval with the language model. In the RS - Agent designed in the present invention, the RAG technology provides domain - specific knowledge support for the LLM by constructing a solution database and a professional material knowledge base to improve the accuracy and relevance of the answers.
[0033] Solution retriever: It retrieves solutions according to the user input information and inputs the retrieved solutions into the central controller.
[0034] The solution retriever is built based on the RAG technology and is used to assist the LLM in selecting suitable computing tools. When the RS - Agent obtains a user query, the central controller determines the task type, then uses the task type as a query to send to the solution retriever, and retrieves relevant solutions from the solution database. In the retrieval stage, the Facebook AI Similarity Search (FAISS) algorithm is used to retrieve the top k relevant documents. These documents are processed by a generator (usually a language model) in the generation stage and combined with the query to generate a response. Finally, the top k relevant documents returned by FAISS are added to the prompt template of the LLM to enhance the LLM's ability to select suitable tools. Specifically:
[0035] In the RS - Agent, the solution retriever M s uses the RAG technology to retrieve information and guide tool selection. After receiving the task type sent by the LLM, the retrieval function f r will use the Facebook AI Similarity Search (FAISS) algorithm to efficiently retrieve in the solution database D s and find the top k documents {d1, d2,..., d k} most relevant to the query Q as the solution guidance g sIn the generation phase, the generator f g (usually a language model) combines the query Q and the retrieved documents to generate a response A. In practical applications, the top k relevant documents returned by FAISS are added to the system prompts of LLM in a specific format, such as "You are a helpful assistant, you should take the following content as Guidance.", to enhance LLM's ability to accurately select the appropriate tool to handle specific problems.
[0036] Professional information retriever: It retrieves professional knowledge related to the user input information based on the user input information and inputs the search results to the central controller.
[0037] The professional data retriever is built based on RAG technology. The professional data retriever is designed to address the problem that the GPT series language models may not be able to provide accurate responses when dealing with specialized topics such as remote sensing due to insufficient training data. When the selected tool involves knowledge retrieval, that is, when RS-Agent needs additional knowledge to solve, it will retrieve relevant documents from proprietary knowledge databases (such as fighter-related knowledge databases) as knowledge guidance. Its retrieval process is the same as the solution retriever. The top k relevant documents returned by FAISS will be sent directly to LLM and a prompt template will be added, so that the model can give precise answers to domain-specific questions and ensure high accuracy of the response. Specifically:
[0038] When the selected tool includes knowledge retrieval, which means that additional knowledge is needed to solve the user query Q, the central controller will determine the query keywords and send them to the professional information retriever M k , Professional Data Retrieval Tool M k The same search method as the solution finder (using the FAISS algorithm to search in the professional information knowledge base D k retrieved from the database) to obtain relevant documents as knowledge guidance k . D k The construction method is similar to D s Similar, but different in subsequent processing, the first k relevant documents returned by FAISS will be directly sent to LLM through a prompt template such as "Using the knowledge guidance below, answer the query...", so that the model can provide precise answers to questions in the field of remote sensing and ensure that answer A has high accuracy.
[0039] Tool Library: Integrates many high-performance remote sensing image processing tools (see Appendix Figure 2), covering disaster assessment (effectively identifying damage to buildings and roads before and after a disaster), target detection (such as detection tools based on yolov8_obb, which can accurately detect targets in remote sensing images, such as identifying the number of aircraft in an image), scene classification (using pre-trained visual transformers to classify remote sensing image scenes), change detection (effectively identifying changes in remote sensing images at different times), image segmentation (segmenting different objects in remote sensing images), and other functions. In addition, it also includes aircraft type recognition tools (based on visual transformers, which can identify a variety of military and civilian aircraft types), SAR image-related tools (such as SAR aircraft target detection tools and SAR aircraft type recognition tools), and functional tools based on other models (such as VQA based on GeoChat, scene classification, and other functions). These tools work together to enable RS-Agent to cope with various complex remote sensing tasks.
[0040] Output module: Outputs the final answer or processed remote sensing image based on the processing results of the central controller.
[0041] In the specific implementation, RS-Agent starts with receiving multi-source remote sensing data (visible light, SAR, hyperspectral) from users, using the large language model (LLM) as the brain, obtaining historical interaction records (memory), and participating in task planning and decision-making (Planning, Reasoning) with the assistance of professional domain knowledge base and tool description knowledge base, calling specific tools in the tool library including target detection, target classification, change detection, etc., and finally generating replies based on the results returned by the tools and interacting with users. This shows that RS-Agent can integrate multi-source remote sensing data, with the help of multiple tools and knowledge bases, to achieve intelligent remote sensing task processing and user interaction.
[0042] For example: Figure 2As shown, the user inputs the problem to be solved on the front-end interface. The problem types cover pure text messages, picture-text messages, and messages combining multiple pictures and text. After the user inputs the problem on the front-end and clicks "Send", if the user input contains an image, the image will be uploaded to the specified location on the server and saved. Then the front-end sends the user request to the back-end. After receiving the request, the LLM tries to understand the user's intention based on the content of the user's problem, calls the corresponding tool to process the problem, and integrates the output result into an expression that is easy for the user to understand. Then, the back-end sends the processing result to the front-end. If the result contains an image output, the image file path will be mapped to an accessible Web URL and sent to the front-end together with the text output. The front-end determines whether to display the image on the Web interface by judging whether the output contains a URL. For example: The user opens the Web interface of RS-Agent and inputs the first question in the chat input box: "How many planes are there at the Capital Airport now?" The front-end packages this information as a request and sends it to the back-end. After receiving the request, the LLM understands this problem and gets 'The user wants to know the number of planes at the Capital Airport. I should call the count plane tool', and searches for the remote sensing image of the Capital Airport in the database, detects and counts the planes on it, and finally returns the counting result and the output image with detection frames. The back-end maps the image path to a URL and sends it to the front-end together with the text output.
[0043] After the front-end receives it, it will display in the chat box: "There are 148 planes at the Capital Airport now. (The image located by the URL is displayed below the text message)"; Since this technology supports multi-round conversations, the user then inputs the second question: "What is the situation of the houses damaged after the wildfire? The pictures before and after the disaster are respectively", and uploads the remote sensing images of a certain area before and after the disaster in two image input components respectively. The front-end sends the user's question and the image file paths before and after the disaster to the back-end. After receiving it, the LLM understands the problem and gets 'The user asked about the situation of the houses damaged after the wildfire and provided the image paths before and after the disaster. To evaluate the damage of the buildings, I should use the building_damage_detection tool', and then calls the tool to detect the images before and after the disaster, returns the final answer and sends it to the front-end. After the front-end receives it, it will display in the chat box: "According to the comparative analysis of the images before and after the disaster, there are 92 undamaged buildings and 14 severely damaged buildings. The number of slightly and moderately damaged buildings is 0."
[0044] In summary, the RS-Agent designed in the present invention uses an advanced LLM as the central controller, combines a variety of high-performance remote sensing image processing tools, can autonomously perceive the environment, accurately understand the user's needs and execute corresponding operations, thereby significantly improving the automation and intelligence level of the system. By introducing the Retrieval-Augmented Generation (RAG) technology, the RS-Agent can access and dynamically utilize a vast knowledge base of professional materials in real time, showing excellent performance in dealing with complex and changing remote sensing tasks, and fully meeting the urgent needs of modern society for efficient remote sensing data processing.
[0045] The above-described embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A remote sensing intelligent interpretation system based on a large language model, characterized in that, It includes: a user input interface, a central controller, a solution retriever, a professional data retriever, a tool library, and an output module; The user input interface is used to receive user input information; The solution retriever is used to retrieve solutions and input the retrieved solutions into the central controller; The professional data retriever is used to retrieve professional knowledge related to the user input information and input the retrieval results into the central controller; The tool library stores a variety of remote sensing image processing tools; The central controller understands and analyzes the user input information, infers the user's intention, generates a task type, and sends the task type to the solution retriever; selects a suitable tool from the tool library according to the solution to process the user input information. If professional knowledge is involved, it generates query keywords and sends them to the professional data retriever, and generates a final response in combination with the professional knowledge retrieval results; The output module outputs the final answer or the processed remote sensing image according to the processing result.
2. The remote sensing intelligent interpretation system based on the large language model according to claim 1, characterized in that, The solution retriever is built based on the RAG technology and is used to assist the central controller in selecting suitable computing tools; in the retrieval stage, the FAISS algorithm is used to retrieve the top k solution documents related to the user input information, and these k solution documents are sent to the central controller and added to the prompt template to enhance the central controller's ability to select suitable tools.
3. A remote sensing intelligent interpretation system based on a large language model according to claim 1, characterized in that, The professional data retriever is built based on the RAG technology and uses the FAISS algorithm to retrieve the top k professional knowledge documents related to the user input information, and these k professional knowledge documents are sent to the central controller and added to the prompt template to ensure high accuracy of the response.
4. The remote sensing intelligent interpretation system based on a large language model according to claim 1, characterized in that, The remote sensing image processing tools in the tool library include disaster assessment, object detection, scene classification, change detection, image segmentation, aircraft type recognition tools, and functional tools based on other models.
5. A remote sensing intelligent interpretation system based on a large language model according to claim 1, characterized in that, The central controller is built based on a large language model.
6. The remote sensing intelligent interpretation system based on the large language model according to claim 1, wherein The process by which the output module outputs the final answer or the remote sensing image according to the processing result includes: determining whether the processing result contains image output. If it does, the image file path is mapped to an accessible Web URL and sent to the front end together with the text output. The front end determines whether to display the image on the Web interface by judging whether the output contains a URL.
7. A remote sensing intelligent interpretation system based on a large language model according to claim 1, characterized in that, It also includes a RAG knowledge base, which consists of a solution database and a professional data knowledge base; the solution database stores a large number of solutions for different remote sensing tasks to help the LLM select suitable tools; the professional data knowledge base stores rich remote sensing professional knowledge to provide accurate answers for complex professional questions.
Citation Information
Patent Citations
Remote sensing interpretation agent system based on large language model
CN119005242A
Knowledge question-answering system based on large language model
CN119396975A
Retrieval execution device, retrieval system, retrieval method and retrieval program
JP2008140272A
Methods and systems for evaluating and optimizing large language models and methods for personalized large language models
US20250117665A1
Computer architecture and method for validating and collecting and metadata and data about the internet and electronic commerce environments (data discoverer)
US6151584A
Cited By
Remote sensing large language model precision improvement method and device based on subdivision database and human feedback, and electronic equipment
CN121121753A
Large-model-driven agricultural remote sensing agent and working method
CN121303187A