Remote sensing disaster intelligent question and answer method based on large model agent

Through a method based on large-model agents, a task execution plan is automatically generated and combined with a variety of perception models and tools, the remote sensing disaster intelligent question-and-answer system is automatically processed, solving the problems of multi-task chain planning and execution in the existing system, and achieving efficient and accurate disaster response.

CN119988696APending Publication Date: 2025-05-13BEIHANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510062598.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing remote sensing disaster intelligent question and answer system is difficult to automatically plan and execute multi-task chains with strong correlation, resulting in users needing to manually select and schedule different models, which is time-consuming and labor-intensive and easy to introduce human errors.

Method used

The method based on large-model agents is adopted to automatically generate task execution plans through large-language models, and the remote sensing images are processed in combination with object detection and semantic segmentation models. Special tools are used to analyze and integrate the perception results to generate the final natural language answer.

Benefits of technology

It realizes highly automated disaster response, can automatically generate and execute tasks according to user requests, improves efficiency and accuracy, reduces manual intervention, and enhances the adaptability and real-time nature of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988696A_ABST
    Figure CN119988696A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a remote sensing disaster intelligent question-answering method based on a large model agent, which comprises the following steps: S1, taking request content input by a user as a main body, and taking task definition, format constraint and tool description as cue words of a large language model, the large language model automatically generates a task execution plan according to the cue word; s2, performing target detection and semantic segmentation on the input remote sensing image according to the task execution plan based on a target detection model and a semantic segmentation model to obtain a sensing result; and S3, the intelligent agent calls corresponding tools according to the task execution plan to analyze the perception result, and integrates the outputs of all the tools to obtain a final answer corresponding to the user request content. According to the method, various perception tasks can be integrated, and the interpretability and the flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and more specifically to an intelligent question-answering method for remote sensing disasters based on a large model intelligent agent. Background Art

[0002] Remote sensing disaster intelligent question answering is a task form in the field of remote sensing image processing. It refers to analyzing given images and questions about the image content through machine learning technology, and outputting the correct answer to the given question.

[0003] In the disaster interpretation of remote sensing images, existing methods mainly focus on solving single tasks, such as segmentation, detection or visual question answering (VQA). Although these methods have made some progress, they are often insufficient for comprehensive disaster assessment when facing complex disaster scenes. Existing disaster interpretation models generally output information in the form of visual results, such as segmentation maps or detection boxes. These forms have certain reference value for professionals, but it is difficult to directly provide clear and intuitive guidance when facing the complex needs of non-professional users or disaster relief sites. In addition, although existing visual question answering models can convert disaster information into answers in natural language form, reducing the cost of understanding, they only output results in text form and lack transparency of the intermediate perception process. This black-box interpretation method may reduce the credibility of the results and make it difficult to ensure the accuracy and reality of the data. For scenarios that require quantitative assessment and disaster response planning, its practicality is greatly limited.

[0004] In actual disaster scenarios, it is often necessary to combine multiple perception tasks to obtain comprehensive information. For example, the change detection results of the disaster area may need to be used as the input of the semantic segmentation task, and the segmented data may be used for further disaster level assessment or scene classification. However, existing methods cannot automatically plan and execute these highly correlated task chains. Users need to manually select and schedule different models and organize the results to form a final analysis. This process is not only time-consuming and labor-intensive, but also prone to errors due to human intervention. The above problems limit the practicality and intelligence of existing methods in disaster scenarios. Therefore, a comprehensive solution that can integrate multiple perception tasks and improve interpretability and flexibility is needed to fill this gap. Summary of the invention

[0005] In view of this, the present invention provides an intelligent question-answering method for remote sensing disasters based on a large model intelligent agent, which solves the technical problem that traditional disaster response systems usually rely on manually designed fixed processing procedures and are difficult to flexibly adjust tasks and processing methods according to specific circumstances.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] A remote sensing disaster intelligent question answering method based on a large model intelligent agent comprises the following steps:

[0008] S1. The request content input by the user is used as the main body, together with the task definition, format constraints and tool description as the prompt words of the large language model. The large language model automatically generates a task execution plan based on the prompt words.

[0009] S2, based on the target detection model and the semantic segmentation model, perform target detection and semantic segmentation on the input remote sensing image according to the task execution plan to obtain the perception result;

[0010] S3. The intelligent agent calls the corresponding tools according to the task execution plan to analyze the perception results, and integrates the outputs of all tools to obtain the final answer to the corresponding user request content.

[0011] Furthermore, the user request content and remote sensing images are used as task inputs, and the task output is represented in the form of a binary tuple (I, Q), where I∈R C×H×W is the input remote sensing image, Q is the input request content, C is the number of channels of the remote sensing image, H is the height of the remote sensing image, and W is the width of the remote sensing image;

[0012] The task execution plan, perception results and final answer are taken as task outputs. The task output is expressed in the form of a triple (P, R, A); where P is the task execution plan, expressed as a subtask sequence, P = {p1, p2, ..., p n}, each subtask comes from a predefined subtask set, and the agent selects and plans these subtasks according to the input request content; R represents the result set obtained after executing the subtask, R = {r1, r2, ..., r n}, where each r i For subtask p i The output of Q includes segmentation maps, object detection boxes or other basic outputs of perception tasks; A represents the generated natural language answer, which is related to the context and generated by the request content Q and the perception result R.

[0013] Furthermore, in S1, the task definition part is to describe the work goal of the agent through a natural language text;

[0014] The tool description section contains detailed information about all tools that can be used in the currently executed task. Each tool is defined with a description in natural language, as well as the number and type of input and output parameters.

[0015] The format constraint is a narrative text, requiring the output of the task execution plan to be returned in the format of a JSON array, and each subtask is defined by three properties: tool identifier, input parameters, and output parameters.

[0016] Furthermore, in S1, a large language model is used to infer the prompt words and generate a JSON object that meets the format requirements as a task execution plan; the task execution plan includes multiple subtasks, each subtask corresponds to a specific operation, and contains a tool identifier, input and output, task execution order and dependencies.

[0017] Furthermore, S1 also includes: verifying the generated JSON object, and the specific verification process includes:

[0018] Assume that the string generated by the large language model is S = s1, s2, ..., s k , from the last character s k Start a reverse search. If it does not conform to the JSON format, backtrack step by step until a valid JSON structure is found. If a valid JSON structure cannot be found after the entire string is traversed, it is considered a generation failure. At this time, the temperature parameter in the autoregressive inference is increased to guide the large language model to generate different results until a valid task execution plan is obtained.

[0019] Furthermore, in S1, after the task execution plan is generated, the legitimacy of each subtask is verified. The verification process includes:

[0020] Check each subtask p i The tool identifier ID referenced in i Does it exist:

[0021]

[0022] Among them, ID is the ID set of all valid tools;

[0023] Verify the input and output of each subtask so that each subtask p i With valid input parameters I i and output parameter O i If there are invalid operations in the task execution plan, the corresponding subtasks are ignored, and the legal subtasks are retained to obtain the final legal task execution plan P valid , expressed as: P valid ={p1,p2,…,p k}, where p i It is a legal subtask.

[0024] Furthermore, in S2, the position and category of each target in the input remote sensing image are identified by a target detection model based on a convolutional neural network, and its output is expressed as:

[0025] B={(c i ,(xi ,y i ,w i ,h i ))|i=1,2,...,n}

[0026] Among them, c i ∈C represents the category of the target, C is the set of all possible categories, (x i ,y i ,w i ,h i ) is the position and size of the bounding box, and n represents the total number of categories in the remote sensing image;

[0027] All outputs are encapsulated and converted into a JSON object array, where each JSON object contains the category information and bounding box coordinates of the target. The final JSON object array is represented as follows:

[0028] Detection Results={{"category":c i ,"bbox":(x i ,y i ,w i ,h i )}|i=1,2,...,n}

[0029] Among them, category represents category information, and bbox represents bounding box coordinates.

[0030] Furthermore, in S2, the remote sensing image is semantically segmented by a semantic segmentation model based on a multi-scale convolutional neural network, and a multi-category segmentation map S is output, in which each pixel p∈I is assigned a semantic category label c(p); assuming that the size of the remote sensing image I is H×W, the segmentation map S is an H×W matrix, in which each element s i,j Represents pixel p i,j The category label is expressed as:

[0031] S={s i,j ∣s i,j ∈C,1≤i≤H,1≤j≤W}

[0032] Among them, C represents the set of all categories, s i,j Represents pixel p i,j The corresponding category label;

[0033] Encapsulate the output and convert it into a dictionary format, where the key is the category name c and the value is the binary mask M of the corresponding category c ; Each binary mask M c is a H×W matrix, representing the positions of all pixels in the image belonging to category c, expressed as:

[0034]

[0035] The final output dictionary structure is:

[0036]

[0037] Among them, c i is the category name, is the binary mask matrix corresponding to the category.

[0038] Furthermore, in S3, for the task of digital calculation, the agent calls the counting tool to count the specified objects in the target detection results;

[0039] For dense prediction tasks, the agent calls the segmentation region calculation tool to calculate the total area of ​​the specified category in the semantic segmentation results.

[0040] Furthermore, in S3, for the path optimization task, the agent calls the A* algorithm as a callable tool. The tool accepts the positions of two points P1 = (x1, y1) and P2 = (x2, y2), and a binary mask M as input to determine whether there is a feasible path between the two points. The path finding problem is expressed as:

[0041]

[0042] Among them, M is a H×W binary mask matrix, which represents the location of the obstacle. If M i,j =1 means that the position (i, j) is an obstacle, M i,j =0 means that the location is accessible; after the agent calls the tool, it determines whether it can reach location P2 from location P1.

[0043] It can be seen from the above technical solutions that, compared with the prior art, the present invention has the following beneficial effects:

[0044] 1) High degree of automation: It can automatically generate and execute tasks based on input images and user requests. Users do not need to set specific operation steps in advance, which reduces manual intervention and improves efficiency.

[0045] 2) Accurate perception and analysis: Using a variety of special tools, such as counting tools, path planning tools, regional analysis tools, etc., it can provide accurate perception results in complex disaster scenarios and avoid the error problems that are prone to occur in traditional systems.

[0046] 3) Strong adaptability: It can handle different types of tasks, whether it is image semantic segmentation, object detection, path planning, or regional statistics, all of which can be dealt with through dynamic adjustment of the intelligent agent, greatly improving the flexibility and adaptability of the intelligent agent system.

[0047] Overall, the present invention organically combines task planning, perception and cognition, so that it can automatically generate execution plans according to user needs, and perform precise perception and analysis through a variety of special tools, thereby achieving a more flexible and efficient disaster response. By optimizing the processing flow of disaster response tasks, it not only improves the system's automation level and processing accuracy, but also enhances the system's adaptability and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0049] Figure 1 A flow chart of the remote sensing disaster intelligent question-answering method based on a large model intelligent agent provided by the present invention;

[0050] Figure 2 Schematic diagram of the reasoning process of the method of the present invention on a real image. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] like Figure 1 As shown, the embodiment of the present invention discloses a remote sensing disaster intelligent question answering method based on a large model intelligent agent, comprising the following steps:

[0053] S1. The request content input by the user is used as the main body, together with the task definition, format constraints and tool description as the prompt words of the large language model. The large language model automatically generates a task execution plan based on the prompt words.

[0054] S2, based on the target detection model and the semantic segmentation model, perform target detection and semantic segmentation on the input remote sensing image according to the task execution plan to obtain the perception result;

[0055] S3. The intelligent agent calls the corresponding tools according to the task execution plan to analyze the perception results, and integrates the outputs of all tools to obtain the final answer to the corresponding user request content.

[0056] The focus of this invention is on the efficient coordination and multi-tool integration of modular structures to meet the complex analysis needs of disaster scenarios. Specifically, this invention consists of three main modules, namely:

[0057] 1) Planning module, which generates an automated task execution plan based on the request content entered by the user. The task requirements are converted into a structured JSON format action plan through prompt words including task definition, format constraints, tool description and input instructions. The planning module makes full use of the reasoning ability of the large language model (LLM) to generate a sequentially related task chain based on the tool description, and uses a dictionary to store intermediate results, supporting the reuse of task results and efficient scheduling.

[0058] Task planning is the core part of the system, which generates the most appropriate task execution plan based on the user input request. The planning module is implemented by a planning agent driven by a large language model. Under the input task request, the planning module selects necessary subtasks from the predefined task space and arranges their execution in sequence. The planning process needs to avoid unnecessary redundant subtasks and ensure that the execution order of each subtask meets the logical requirements. This part of the function is implemented through the prompt words of the large model. The prompt words include user input, task space and planning requirements. The user input provides the user's problem context for the large model, and the task space provides the large model with optional predefined task descriptions. The planning requirements limit the output of the large language model to conform to the list form specification in order to automate the processing of subsequent processes.

[0059] 2) The perception module extracts key disaster information from remote sensing images and provides data support for subsequent tasks. The perception module contains a perception model library, which includes an image semantic segmentation model and a target detection model. The semantic segmentation model converts the input image into a segmentation map with different semantics, while the target detection model converts the input image into segmentation boxes of different categories. If the aforementioned planning module plans the perception model in the perception module, the corresponding perception model in the perception module will be executed and the result will be temporarily stored in the memory for subsequent interpretation.

[0060] 3) The cognitive module analyzes the perception results through special tools to solve the problem of insufficient accuracy of existing large language models in numerical tasks and quantitative analysis. Finally, the cognitive module integrates the outputs of all tools and generates the final answer corresponding to the user's request through the summary agent. The cognitive module consists of two parts. The first part is the cognitive tool library, which provides counting, area statistics and pathfinding tools. If the aforementioned planning module plans the tools in the tool library, the corresponding tools in the cognitive module will be executed and the results will be temporarily stored in the memory. The second part is the summary agent, which is driven by the large language model. Its input prompt words include user input, temporary results generated by all previous tools, and summary requirements. The agent generates a response to the user input based on the previous perception and cognitive results.

[0061] The present invention needs to first define the remote sensing disaster intelligent question answering task, taking the user request content and remote sensing images as the task input, and the task output is expressed in the form of a binary (I, Q), where I∈R C×H×W is the input remote sensing image, Q is the input request content, C is the number of channels of the remote sensing image, H is the height of the remote sensing image, and W is the width of the remote sensing image;

[0062] The task execution plan, perception results and final answer are taken as task outputs. The task output is expressed in the form of a triple (P, R, A); where P is the task execution plan, expressed as a subtask sequence, P = {p1, p2, ..., p n}, each subtask comes from a predefined subtask set, and the agent selects and plans these subtasks according to the input request content; R represents the result set obtained after executing the subtask, R = {r1, r2, ..., r n}, where each r i For subtask p i The output of Q includes segmentation maps, object detection boxes or other basic outputs of perception tasks; A represents the generated natural language answer, which is related to the context and generated by the request content Q and the perception result R.

[0063] The overall goal of the mission is to maximize the accuracy of planning (P), perception (R), and recognition (A) to ensure that the mission can effectively and comprehensively interpret and respond to complex disaster scenarios.

[0064] The above steps are further explained below.

[0065] S1. The core function of the planning module is to generate a series of subtask plans that can be automatically executed according to user needs. This module uses the large language model (LLM) to automatically generate plans and ensure that the generated plans meet the format requirements and can be smoothly parsed and executed in the subsequent execution process. The workflow of the planning module is roughly divided into the following steps:

[0066] S11, prompt word composition.

[0067] The user request is the main body, together with the task definition, format constraints and tool description, which constitute the large model prompt words. The user request is the input natural language request, and the final task of the planning module is to generate the corresponding operation plan based on these requests. For example: "Identify the disaster area", "Calculate the scope of the flood" or "Find a path from the current location to the target location".

[0068] The task definition part is to describe the work objectives of the agent through a natural language text.

[0069] The tool description section provides detailed information about all tools that can be used in the currently executed task. Each tool is defined with a natural language description and lists the number and types of input and output parameters.

[0070] The format constraint is a narrative text, requiring the output of the task execution plan to be returned in the format of a JSON array, and each subtask is defined by three properties: tool identifier, input parameters, and output parameters, specifically expressed as:

[0071] T i =(ID i ,I i ,O i )

[0072] Among them, T i Indicates the i-th subtask, ID i is the tool identifier, I i is the input parameter set, O i is the output parameter set.

[0073] S12. Large model reasoning.

[0074] The planning module uses the large language model (LLM) to infer the prompt words and generate a JSON object that meets the format requirements as a task execution plan; the task execution plan includes multiple subtasks, each subtask corresponds to a specific operation and contains a tool identifier, input and output, task execution order and dependencies.

[0075] During the inference process, the prompt word is feature encoded and passed into the large model to obtain an output feature vector sequence Y = (y1, y2, ..., y n ), where Y is the feature sequence output by the model. Through autoregressive generation, the model generates current features based on historical outputs and de-encodes them into text information to generate the final task plan.

[0076] The model automatically selects the appropriate tool by combining the user request Q and the tool description through context learning technology, and generates a sequence of execution steps P = (p1, p2, ..., p m ).

[0077] S13. Output verification.

[0078] In practical applications, the output of a large model may not conform to the expected format, so output verification is required. In order to ensure the validity of the output structure, the rejection sampling method is used to correct the output of the planning module. The specific verification process includes:

[0079] Assume that the string generated by the large language model is S = s1, s2, ..., s k , from the last character s k Start a reverse search. If it does not conform to the JSON format, backtrack step by step until a valid JSON structure is found. If a valid JSON structure cannot be found after the entire string is traversed, it is considered a generation failure. At this time, the temperature parameter in the autoregressive inference is increased to guide the large language model to generate different results until a valid task execution plan is obtained.

[0080] After the task execution plan is generated, the legitimacy of each subtask is verified. The verification process includes:

[0081] Check each subtask p i The tool identifier ID referenced in i Does it exist:

[0082]

[0083] Among them, ID is the ID set of all valid tools;

[0084] Verify the input and output of each subtask so that each subtask p i With valid input parameters I i and output parameter O i If there are invalid operations in the task execution plan, the corresponding subtasks are ignored, and the legal subtasks are retained to obtain the final legal task execution plan P valid , expressed as: P valid ={p1,p2,...,p k}, where p i It is a legal subtask.

[0085] The core function of S2, the perception module, is to link the visual perception task with the large model (LLM), which can extract key information from the input disaster scene image and convert this information into a structured format for subsequent processing. The perception module is mainly composed of the object detection model and the semantic segmentation tool.

[0086] Specifically, the target detection model is implemented based on the convolutional neural network structure. Its input is the remote sensing disaster scene image I, and its output is a set of detection boxes B = {b1, b2, ..., b n}, where each detection box b i Contains the target category information c i and the coordinate information of the bounding box (x i ,y i ,w i ,h i ), respectively representing the upper left corner coordinates and width and height of the target box. The task of the target detection model is to provide the position and category of each target in the image based on the input image I, and its output is expressed as:

[0087] B={(c i ,(x i ,y i ,w i ,h i ))|i=1,2,…,n}

[0088] Among them, c i ∈C represents the category of the target, C is the set of all possible categories, (x i ,y i ,w i ,h i ) is the position and size of the bounding box, and n represents the total number of categories in the remote sensing image;

[0089] In order to connect the detection results with the large model, all outputs are encapsulated and converted into a JSON object array, where each JSON object contains the category information and bounding box coordinates of the target. The final JSON object array is represented as follows:

[0090] Detection Results={{"category":c i ,"bbox":(x i ,y i ,w i ,h i )}|i=1,2,…,n}

[0091] Among them, category represents category information, and bbox represents bounding box coordinates.

[0092] In this way, the large model can directly read and process the detection results and link them with subsequent tasks such as target recognition or path planning.

[0093] The semantic segmentation model is implemented through a multi-scale convolutional neural network. The input is the disaster scene image I, and the output is a multi-category segmentation map S, where each pixel p∈I is assigned a semantic category label c(p); assuming that the size of the remote sensing image I is H×W, the segmentation map S is an H×W matrix, where each element s i,j Represents pixel p i,j The category label is expressed as:

[0094] S={s i,j ∣s i,j ∈C,1≤i≤H,1≤j≤W}

[0095] Among them, C represents the set of all categories, s i,j Represents pixel p i,j Corresponding category labels; the model is trained on a disaster scene dataset and can accurately assign category labels to each pixel in the image.

[0096] In order to connect it with the large model agent, the output is encapsulated and converted into a dictionary format, with the key being the category name c and the value being the binary mask M of the corresponding category. c ; Each binary mask M c is a H×W matrix, representing the positions of all pixels in the image belonging to category c, expressed as:

[0097]

[0098] This format enables the agent to efficiently process and analyze different regions in the image and apply it to subsequent tasks such as path planning or target recognition. The final output dictionary structure is:

[0099]

[0100] Among them, c i is the category name, is the binary mask matrix corresponding to the category.

[0101] S3. The cognitive module mainly consists of two parts: the cognitive tool library and the summary agent.

[0102] S31, Special tool support:

[0103] For numerical calculation tasks (such as counting the area of ​​disaster-affected areas, calculating the number of targets, etc.), large models may make mistakes when processing numerical values ​​because their predictions based on probability distribution are uncertain. In order to handle these tasks, a dedicated counting tool is introduced. In this task, the agent calls the counting tool to count the specified objects in the target detection results. Its input is the target detection result B = {(c i ,(x i ,y i ,w i ,h i ))} and target object type c target , the output is the exact number N of objects of the specified type target , the formula is:

[0104]

[0105] in, is the indicator function, when c i =c target The value is 1 when the target is detected, otherwise it is 0, indicating whether the target detection frame contains objects of the specified type. By counting the target detection results, this tool can accurately count the number of objects of the specified type in the disaster area to avoid inaccurate numbers.

[0106] For dense prediction tasks, large models cannot directly use segmentation maps as input and output. In order to effectively use the segmentation results, a segmentation area calculation tool is introduced. This tool can calculate the total area of ​​a specified category in the semantic segmentation result. Its input is the segmentation map S = {s i,j} and the target category name c target , the output is the total area A of the specified category target , the formula is:

[0107]

[0108] Among them, s i,j ∈C represents the category of pixel (i, j) in the segmentation map, is an indicator function, if s i,j =c target If the value is 1, it is 1, otherwise it is 0. This tool can quantify the spatial distribution of different objects in the disaster area and provide data support for subsequent disaster analysis.

[0109] For the path optimization task, the agent calls the A* algorithm as a callable tool. The tool accepts the positions of two points P1 = (x1, y1) and P2 = (x2, y2), and a binary mask M as input to determine whether there is a feasible path between the two points. The path finding problem is expressed as:

[0110]

[0111] Among them, M is a H×W binary mask matrix, which represents the location of the obstacle. If M i,j =1 means that the position (i, j) is an obstacle, M i,j =0 means that the location is accessible; after the agent calls the tool, it determines whether it can reach location P2 from location P1, providing reliable support for rescue decisions.

[0112] S32. Summarize the agent.

[0113] The summary agent combines the outputs of all previously executed models and tools and generates a final natural language answer based on the user's request. The prompt for the summary agent includes the following parts: task definition, action history, and user request. The task definition requires the agent to provide a final answer based on the user's request. The action history records the steps performed by the agent and their results, helping the agent understand the context of the entire task and reference relevant intermediate results in the final answer. The user request is the user's original question or expectation, describing what the system should answer.

[0114] Next, the performance of the method of the present invention is further verified.

[0115] 1. The effect of the search method proposed in the present invention is verified on the dataset of remote sensing image disaster interpretation and question answering. In the experiment, the existing visual question answering model and the method proposed in the present invention are compared, and the verification results are shown in Table 1.

[0116] Table 1 Comparative experiment on disaster interpretation of remote sensing images and question answering

[0117]

[0118] The method proposed in the present invention is interpretable, so it can output the planning steps, while the existing visual question answering method cannot achieve the interpretability of the intermediate steps, so the corresponding results in the table are empty. Among them, VR represents the legal proportion of the intermediate steps, P represents the accuracy of the intermediate steps, R represents the recall rate of the intermediate steps, match represents the correct rate of the strictly matched answers, and GPTScore represents the correct rate of the answers matched by the large model.

[0119] Experimental results show that in terms of multiple evaluation indicators, the intelligent agent architecture proposed in the present invention can achieve better results than existing question-answering methods, especially when driven by the Tongyi Qianwen 2.5 model, the accuracy can reach above 0.54, reflecting the superiority of the effect of the present invention.

[0120] 2. The effect of the search method proposed in this invention is verified on the complex disaster scene of real remote sensing images. Figure 2As shown in the figure, for the question of the average area of ​​completely destroyed houses, the agent first breaks it down into two parallel tasks: counting the total area and the total number. Then, it performs a division operation based on the total area and the total number and calculates the correct average area. During the reasoning process, the agent can output the intermediate segmentation and detection results. The overall process has strong interpretability, which can reflect the superiority of the present invention over existing methods.

[0121] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0122] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing disaster intelligent question answering method based on a large model agent, characterized in that: The following steps are involved: S1. The request content input by the user is used as the main body, together with the task definition, format constraints and tool description as the prompt words of the large language model. The large language model automatically generates a task execution plan based on the prompt words. S2, based on the target detection model and the semantic segmentation model, perform target detection and semantic segmentation on the input remote sensing image according to the task execution plan to obtain the perception result; S3. The intelligent agent calls the corresponding tools according to the task execution plan to analyze the perception results, and integrates the outputs of all tools to obtain the final answer to the corresponding user request content.

2. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: The user request content and remote sensing images are used as task input, and the task output is represented as a binary tuple (I, Q), where I∈R C ×H×W is the input remote sensing image, Q is the input request content, C is the number of channels of the remote sensing image, H is the height of the remote sensing image, and W is the width of the remote sensing image; The task execution plan, perception results and final answer are taken as task outputs. The task output is expressed in the form of a triple (P, R, A); where P is the task execution plan, expressed as a subtask sequence, P = {p1, p2, ..., p n }, each subtask comes from a predefined subtask set, and the agent selects and plans these subtasks according to the input request content; R represents the result set obtained after executing the subtask, R = {r1, r2, ..., r n }, where each r i For subtask p i The output of Q includes segmentation maps, object detection boxes or other basic outputs of perception tasks; A represents the generated natural language answer, which is related to the context and generated by the request content Q and the perception result R.

3. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S1, the task definition part is to describe the work goal of the agent through a natural language text; The tool description section contains detailed information about all tools that can be used in the currently executed task. Each tool is defined with a description in natural language, as well as the number and type of input and output parameters. The format constraint is a narrative text, requiring the output of the task execution plan to be returned in the format of a JSON array, and each subtask is defined by three properties: tool identifier, input parameters, and output parameters.

4. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S1, a large language model is used to infer the prompt words and generate a JSON object that meets the format requirements as a task execution plan; the task execution plan includes multiple subtasks, each subtask corresponds to a specific operation, and contains a tool identifier, input and output, task execution order and dependencies.

5. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 4 is characterized in that: S1 also includes: verifying the generated JSON object. The specific verification process includes: Assume that the string generated by the large language model is S = s1, s2, ..., s k , from the last character s k Start a reverse search. If it does not conform to the JSON format, backtrack step by step until a valid JSON structure is found. If a valid JSON structure cannot be found after the entire string is traversed, it is considered a generation failure. At this time, the temperature parameter in the autoregressive inference is increased to guide the large language model to generate different results until a valid task execution plan is obtained.

6. The remote sensing disaster intelligent question answering method based on large model agent according to claim 5 is characterized in that: In S1, after the task execution plan is generated, the legitimacy of each subtask is verified. The verification process includes: Check each subtask p i The tool identifier ID referenced in i Does it exist: Among them, ID is the ID set of all valid tools; Verify the input and output of each subtask so that each subtask p i With valid input parameters I i and output parameter O i If there are invalid operations in the task execution plan, the corresponding subtasks are ignored, and the legal subtasks are retained to obtain the final legal task execution plan P valid , expressed as: P valid ={p1,p2,…,p k }, where p i It is a legal subtask.

7. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S2, the location and category of each target in the input remote sensing image are identified by the target detection model based on convolutional neural network, and its output is expressed as: B={(c i ,(x i ,y i ,w i ,h i ))∣i=1,2,…,n} Among them, c i ∈C represents the category of the target, C is the set of all possible categories, (x i ,y i ,w i ,h i ) is the position and size of the bounding box, and n represents the total number of categories in the remote sensing image; All outputs are encapsulated and converted into a JSON object array, where each JSON object contains the category information and bounding box coordinates of the target. The final JSON object array is represented as follows: Detection Results={{"category":c i ,"bbox":(x i ,y i ,w i ,h i )}∣i=1,2,...,n} Among them, category represents category information, and bbox represents bounding box coordinates.

8. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S2, the remote sensing image is semantically segmented by a semantic segmentation model based on a multi-scale convolutional neural network, and a multi-category segmentation map S is output, in which each pixel p∈I is assigned a semantic category label c(p); assuming that the size of the remote sensing image I is H×W, the segmentation map S is an H×W matrix, in which each element s i,j Represents pixel p i,j The category label is expressed as: S={s i,j ∣s i,j ∈C,1≤i≤H,1≤j≤W} Among them, C represents the set of all categories, s i,j Represents pixel p i,j The corresponding category label; Encapsulate the output and convert it into a dictionary format, where the key is the category name c and the value is the binary mask M of the corresponding category c ; Each binary mask M c is a H×W matrix, representing the positions of all pixels in the image belonging to category c, expressed as: The final output dictionary structure is: Among them, c i is the category name, is the binary mask matrix corresponding to the category.

9. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S3, for the task of digital calculation, the agent calls the counting tool to count the specified objects in the target detection results; For dense prediction tasks, the agent calls the segmentation region calculation tool to calculate the total area of ​​the specified category in the semantic segmentation results.

10. The remote sensing disaster intelligent question answering method based on a large model agent according to claim 1 is characterized in that: In S3, for the path optimization task, the agent calls the A* algorithm as a callable tool. The tool accepts the positions of two points P1 = (x1, y1) and P2 = (x2, y2), and a binary mask M as input to determine whether there is a feasible path between the two points. The path finding problem is expressed as: Among them, M is a H×W binary mask matrix, which represents the location of the obstacle. If M i,j =1 means that the position (i, j) is an obstacle, M i,j =0 means that the location is accessible; after the agent calls the tool, it determines whether it can reach location P2 from location P1.

Citation Information

Cited By

  • Large language model assisted remote sensing index calculation method and device and medium

    CN120706559A

  • Space task processing method, electronic equipment and storage medium

    CN120996390A

  • Model question and answer method, device and equipment, storage medium and computer program product

    CN121009173A

  • Multi-modal disaster relief information analysis and task intelligent planning method and device

    CN121563166A