Large model-agent collaborative dynamic extensible remote sensing target interpretation system and method, and medium
The dynamic and scalable remote sensing target interpretation system, which utilizes a large model-agent collaboration, decomposes tasks using a large language model and constructs parallel execution sequences. This solves the problems of long execution time and low intelligence level in remote sensing target interpretation systems under complex scenarios, and achieves efficient multi-source data processing and adaptive capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AEROSPACE INFORMATION RES INST CAS
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing remote sensing target interpretation systems suffer from rigid workflows and low levels of intelligence when dealing with complex, variable, and unstructured remote sensing scenarios, resulting in long execution times and making it difficult to meet the needs of multi-source heterogeneous data fusion and dynamic task execution.
A dynamic and scalable remote sensing target interpretation system with large model-agent collaboration is adopted. The system decomposes the text task into subtasks and obtains the dependencies through a large language model. The planning module constructs serial and parallel task execution sequences, and the execution module performs multimodal remote sensing image processing. By combining the optimal image processing algorithm and dynamic execution strategy, the parallel scheduling and adaptive processing of subtasks are realized.
It improves the processing efficiency of the remote sensing target interpretation system, shortens the total execution time, enhances the system's intelligence and adaptability, and improves its fault tolerance.
Smart Images

Figure CN122024017A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of remote sensing image processing and artificial intelligence, specifically to a dynamic and scalable remote sensing target interpretation system and method based on large model-agent collaboration, and a computer-readable storage medium. Background Technology
[0002] Remote sensing target resolution systems serve as crucial technical support for fields such as Earth observation, environmental monitoring, urban planning, and national defense security. They aim to automatically extract ground targets and their semantic information from multimodal remote sensing data, including visible light, infrared, synthetic aperture radar (SAR), and hyperspectral data.
[0003] In related technologies, most remote sensing target interpretation systems are based on predefined algorithmic flows and fixed working modes. Domain experts design fixed task processing pipelines, combining preprocessing, detection, classification, and segmentation tasks sequentially. While this mode offers some stability for specific tasks, it suffers from rigid workflows and low levels of intelligence when facing complex, dynamic, and unstructured remote sensing scenarios, resulting in long execution times for complex tasks. Furthermore, integrating new algorithms or data types requires modification and redeployment of the underlying code, making it difficult to meet the practical needs of current multi-source heterogeneous data fusion and dynamic task execution. Summary of the Invention
[0004] This application provides a dynamic and scalable remote sensing target interpretation system and method based on large model-agent collaboration, and a computer-readable storage medium; wherein... In a first aspect, embodiments of this application provide a dynamic and scalable remote sensing target interpretation system based on a large model-agent collaboration. This system includes a cognitive module, a planning module, and an execution module. The cognitive module decomposes the text task to be processed using a large language model, obtains the sub-tasks corresponding to the text task, and identifies the direct dependencies between each pair of sub-tasks. The planning module plans the task sequence based on the direct dependencies between sub-tasks, constructing a task execution sequence. This task execution sequence includes a first sub-task for sequential execution and a second sub-task for parallel execution. The execution module processes the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence and outputs the task execution results.
[0005] Secondly, embodiments of this application provide a dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration, applied to a dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration. The system includes a cognitive module, a planning module, and an execution module. The method includes: the cognitive module decomposing the text task to be processed using a large language model, obtaining sub-tasks corresponding to the text task, and obtaining the direct dependencies between every two sub-tasks; the planning module planning the task sequence of the sub-tasks based on the direct dependencies between them, constructing a task execution sequence; wherein the task execution sequence includes a first sub-task for sequential execution and a second sub-task for parallel execution; the execution module performing task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence, and outputting the task execution results.
[0006] Thirdly, embodiments of this application provide a dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the method of the second aspect.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of the second aspect.
[0008] In this embodiment, firstly, the cognitive module decomposes the input text task using a large language model to obtain the corresponding subtasks and the dependencies between them. Then, the planning module plans the subtasks based on these dependencies, thereby obtaining a task execution sequence that includes a first subtask for sequential execution and a second subtask for parallel execution. Finally, the execution module processes the corresponding multimodal remote sensing image based on this task execution sequence. Thus, the large model-agent collaborative dynamic scalable remote sensing target interpretation system provided in this application can not only successfully execute all subtasks but also achieve sequential and parallel scheduling of subtasks, thereby reducing the total execution time of all subtasks and improving the system's processing efficiency.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0011] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0012] Figure 1 This is a schematic diagram of the structure of the dynamic and scalable remote sensing target interpretation system with large model-agent collaboration proposed in an embodiment of this application. Figure 2 This is a schematic diagram of the planning module proposed in an embodiment of this application. Figure 1 ; Figure 3 This is a schematic diagram of a directed acyclic graph proposed in an embodiment of this application; Figure 4 This is a schematic diagram of the planning module proposed in an embodiment of this application. Figure 2 ; Figure 5 This is a flowchart illustrating the dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration proposed in this application. Figure 1 ; Figure 6 This is a flowchart illustrating the dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration proposed in this application. Figure 2 ; Figure 7 This is a flowchart illustrating the dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration proposed in this application. Figure 3 ; Figure 8 This is a schematic diagram of the structure of the dynamic and scalable remote sensing target interpretation system with large model-agent collaboration proposed in an embodiment of this application. Figure 9 This is a flowchart illustrating the multi-task collaboration mechanism proposed in the embodiments of this application; Figure 10 This is a schematic diagram illustrating the generation process of intermediate results proposed in the embodiments of this application; Figure 11 This is a schematic diagram of the exception handling process proposed in the embodiments of this application; Figure 12This is a flowchart illustrating the adaptive matching mechanism for data and tools proposed in an embodiment of this application. Figure 13 This is a schematic diagram of the structure of the decoding system proposed in the embodiments of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0015] In the following description, references to "some embodiments," "this embodiment," "this application embodiment," and examples, etc., describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0016] The descriptions such as "first," "second," and "third" appearing in the embodiments of this application do not have a specific meaning (such as no order, nor do they indicate a special limitation on the number of devices in the embodiments of this application), but are merely for the purpose of clearly describing the embodiments of this application and do not constitute any limitation on the embodiments of this application.
[0017] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies or terms of the embodiments of this application are described below. The following relevant technologies or terms are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.
[0018] Interpretation systems in related technologies use large language models to obtain subtasks based on the input text task. They then directly construct a sequential task execution sequence based on these subtasks, and process the corresponding remote sensing images using this sequence. Such systems are suitable for certain scenarios (e.g., where each subtask has a strict order) and possess a certain degree of stability. However, when processing complex tasks, they require a long execution time, resulting in low system efficiency.
[0019] To address the aforementioned issues, this application provides a dynamic and scalable remote sensing target interpretation system based on a large model-agent collaboration. In this system, a large language model first decomposes the input text task into corresponding subtasks and obtains the direct dependencies between each pair of subtasks. Then, based on these dependencies, a task execution sequence is constructed, including a first subtask for sequential execution and a second subtask for parallel execution. Finally, the remote sensing images corresponding to the subtasks are processed based on the task execution sequence, and the corresponding task execution results are output. Because the system's task execution sequence includes a second subtask that can be executed in parallel, the system executes the second subtask concurrently, thereby shortening the total task execution time and improving system efficiency.
[0020] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0021] Embodiments of this application propose a dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration, such as... Figure 1 As shown, the large model-agent collaborative dynamic scalable remote sensing target interpretation system 100 includes a cognitive module 110, a planning module 120, and an execution module 130.
[0022] The cognitive module 110 is used to decompose the text task to be processed using a large language model, obtain the sub-tasks corresponding to the text task to be processed, and obtain the direct dependency relationship between each two sub-tasks.
[0023] In this embodiment of the application, the cognitive module 110 is mainly responsible for semantic understanding and logical reasoning of the text task, thereby realizing the decomposition of the text task.
[0024] In the embodiments of this application, the large language model is a natural language processing model based on deep learning, which has the capabilities of semantic understanding, logical reasoning, code generation, and knowledge integration.
[0025] In some embodiments, a large language model refers to a pre-trained natural language processing model, and the training set used during training includes text tasks and sub-tasks corresponding to the text tasks.
[0026] For example, large language models include, but are not limited to, the GPT family, LLaMA, and PaLM.
[0027] In some embodiments, the text task to be processed refers to a task instruction in natural language form entered by the user.
[0028] For example, the text task to be processed is: to detect and analyze ships in the port area.
[0029] In some embodiments, a subtask refers to the smallest executable unit that the large language model decomposes from the text task to be processed.
[0030] For example, subtasks include: identifying port areas, detecting ships located in port areas, analyzing the behavior of detected ships, and predicting the future trajectories of detected ships.
[0031] In the embodiments of this application, a direct dependency refers to a direct dependency between two subtasks.
[0032] For example, detecting a vessel located in a port area depends on identifying the port area, analyzing the behavior of the detected vessel depends on detecting the vessel located in the port area, and predicting the future trajectory of the detected vessel depends on detecting the vessel located in the port area.
[0033] In some embodiments, the large language model generates direct dependencies between subtasks while decomposing them into subtasks.
[0034] In some embodiments, the large language model generates task types for subtasks while decomposing them; then, it obtains the direct dependencies between subtasks through predefined direct dependencies between task types.
[0035] The planning module 120 is used to plan the task order of subtasks based on the direct dependencies between subtasks and construct a task execution sequence; wherein the task execution sequence includes a first subtask for serial execution and a second subtask for parallel execution.
[0036] In this embodiment, the planning model 120 is mainly used to plan the sub-tasks generated by the cognitive module 110 in order to generate an orderly and reasonable task execution sequence.
[0037] In the embodiments of this application, serial execution means that tasks must be executed in a certain order.
[0038] In the embodiments of this application, parallel execution means that tasks can be executed simultaneously.
[0039] In the embodiments of this application, the first subtask refers to a subtask for serial execution.
[0040] For example, the first subtask includes: identifying the port area (T1) and detecting the vessel located in the port area (T2).
[0041] In the embodiments of this application, the second subtask refers to a subtask used for parallel execution.
[0042] For example, the second subtask includes: analyzing the behavior of the detected ships (T3) and predicting the future trajectory of the detected ships (T4).
[0043] In the embodiments of this application, the task execution sequence refers to the arrangement and combination of subtasks that exist simultaneously in a serial / parallel manner and have an execution order.
[0044] For example, the task execution sequence is as follows: .
[0045] In some embodiments, such as Figure 2 As shown, the planning module 120 includes a relationship building module 1201, a task acquisition module 1202, and a sequence generation module 1203.
[0046] Among them, the relationship construction module 1201 is used to construct a directed acyclic graph based on the task execution steps of subtasks.
[0047] In this embodiment, the relationship building module 1201 is responsible for modeling the task execution steps between subtasks to generate a data structure that the system can understand.
[0048] In the embodiments of this application, the task execution steps refer to the execution order among subtasks.
[0049] In the embodiments of this application, a Directed Acyclic Graph (DAG) is a data structure in graph theory used to represent direct dependencies between subtasks.
[0050] In some embodiments, each subtask of this application is a node in a directed acyclic graph, and the direct dependencies between tasks are directed edges in the directed acyclic graph.
[0051] For example, such as Figure 3 As shown, the directed acyclic graph consists of four nodes: T1, T2, T3, and T4. T1 represents the identified port area, T2 represents the detected ships within the port area, T3 represents the analysis of the detected ship behavior, and T4 represents the prediction of the detected ship's future trajectory.
[0052] The task acquisition module 1202 is used to acquire the first subtask and the second subtask based on the directed acyclic graph and the graph algorithm. The first subtask is constructed based on subtasks with direct dependencies, and the second subtask is constructed based on subtasks without direct dependencies.
[0053] In this embodiment of the application, the task acquisition module 1202 is responsible for traversing the directed acyclic graph using the corresponding graph algorithm to obtain the first subtask and the second subtask.
[0054] For example, graph algorithms include, but are not limited to, topology algorithms, genetic algorithms, depth-first search algorithms, and breadth-first search algorithms.
[0055] In some embodiments, the first subtask is built based on direct dependencies.
[0056] For example, such as Figure 3 As shown, since there is a directed edge between T1 and T2, the first subtask includes T1 and T2.
[0057] In some embodiments, the second subtask is built on a basis that does not have direct dependencies.
[0058] For example, such as Figure 3 As shown, since T3 and T4 are located at the same level and there are no directed edges, the second subtask includes T3 and T4.
[0059] The sequence generation module 1203 is used to generate a task execution sequence based on the first subtask and the second subtask.
[0060] In this embodiment, the sequence generation module 1203 is responsible for sequentially planning the first subtask and the second subtask to generate a task execution sequence.
[0061] For example, such as Figure 3 As shown, although there is no direct dependency between T3 and T4, both T3 and T4 have a direct dependency on T2 and are located at the same level. Therefore, the execution order of the second subtask, including T3 and T4, is after T2. The generated task execution sequence is as follows: .
[0062] The execution module 130 is used to perform task processing on the multimodal remote sensing images corresponding to the subtasks based on the task execution sequence and output the task execution results.
[0063] In this embodiment of the application, the execution module 130 is mainly used to perform task processing on the multimodal remote sensing images to be processed corresponding to each sub-task in the planned task execution sequence, and output the corresponding task execution results.
[0064] In some embodiments, when the execution module 130 processes the first subtask, it executes them sequentially; when it processes the second subtask, it starts multiple tasks simultaneously to achieve parallel execution of the subtasks in the second subtask.
[0065] In the embodiments of this application, the multimodal remote sensing images to be processed refer to remote sensing images from different sensor platforms, and there is a corresponding relationship with the sub-tasks.
[0066] For example, when the subtask is to identify the port area, the multimodal remote sensing image to be processed is image data captured by a space platform; when the subtask is to predict the future trajectory of detected ships, the multimodal remote sensing image to be processed is a video stream obtained through port monitoring.
[0067] In this embodiment of the application, the task execution result is the corresponding execution result after the subtask is completed.
[0068] In some embodiments, if all subtasks are executed successfully, the task execution result includes the result of combining the execution results of all subtasks.
[0069] For example, the task execution result includes the presence of vessel 1, vessel 2, and vessel 3 in the current port area. Among them, vessel 1 is loading cargo and will remain stationary for a period of time; vessel 2 is refueling and will leave the port after refueling; vessel 3 is entering the port and will be moored in parking area number three for a period of time.
[0070] In some embodiments, the planning module 120 is further configured to obtain a first set of image processing algorithms corresponding to the task type of the subtask based on a first mapping relationship between task type and image processing algorithm.
[0071] In the embodiments of this application, task type refers to the category corresponding to the nature and objectives of the task.
[0072] In some embodiments, the task includes a perception task.
[0073] For example, the types of perception tasks include, but are not limited to, detection tasks, recognition tasks, segmentation tasks, tracking tasks, scene classification tasks, change detection tasks, and 3D reconstruction tasks.
[0074] In some embodiments, the task also includes a cognitive task.
[0075] For example, cognitive tasks include, but are not limited to, scene understanding tasks, relational reasoning tasks, behavior analysis tasks, trajectory prediction tasks, and decision support tasks.
[0076] In some embodiments, an image processing algorithm refers to a collection of image processing algorithms designed for various task types.
[0077] For example, image processing algorithms used for detection tasks include, but are not limited to, YOLO, RetinaNet, and CenterNet; image processing algorithms used for recognition tasks include, but are not limited to, ResNet and SVM.
[0078] In some embodiments, the first mapping relationship is a predefined association relationship used to establish a corresponding relationship between different task types and corresponding image processing algorithms, thereby ensuring that the planning module 120 can quickly identify image processing algorithms suitable for the current task type.
[0079] For example, the first mapping relationship is established in the form of a table, as shown in Table 1: Table 1
[0080] For example, if the task type of the subtask being processed is a detection task, Table 1 can be used to obtain the first image processing algorithm corresponding to this subtask, including YOLO, RetinaNet, and CenterNet.
[0081] In some embodiments, the planning module 120 is further configured to obtain a second set of image processing algorithms corresponding to the modalities of the multimodal remote sensing image based on a second mapping relationship between image modalities and image processing algorithms.
[0082] In the embodiments of this application, image modality refers to the sensor type or data format that acquires multimodal remote sensing images.
[0083] For example, image modalities include, but are not limited to: visible light mode, infrared mode, SAR mode, hyperspectral mode, and multispectral mode.
[0084] In some embodiments of this application, the image processing algorithm refers to a collection of image processing algorithms designed for various image modalities.
[0085] For example, image processing algorithms for processing infrared modal remote sensing images include, but are not limited to, YOLO and CenterNet; image processing algorithms for processing multispectral modal remote sensing images include, but are not limited to, ResNet and SVM.
[0086] In some embodiments, the second mapping relationship is a predefined association relationship used to establish a corresponding relationship between different image modalities and corresponding image processing algorithms, thereby ensuring that the planning module 120 can quickly identify the image processing algorithm suitable for the current remote sensing image modality.
[0087] For example, the second mapping relationship is established in the form of a table, as shown in Table 2: Table 2
[0088] For example, when the mode of the multimodal remote sensing image to be processed corresponding to the current subtask is infrared mode, the second image processing algorithm corresponding to the multimodal remote sensing image can be obtained through Table 2, including YOLO and CenterNet.
[0089] In some embodiments, the planning module 120 is further configured to perform algorithm evaluation based on the first image processing algorithm set and the second image processing algorithm set to obtain the optimal image processing algorithm.
[0090] In the embodiments of this application, algorithm evaluation refers to evaluating the performance of the corresponding image processing algorithm to determine the image processing algorithm with the best performance.
[0091] In some embodiments, such as Figure 4 As shown, the planning module 120 also includes a first algorithm screening module 1204, a second algorithm screening module 1205, a third algorithm screening module 1206, and an algorithm evaluation module 1207.
[0092] In some embodiments, the first algorithm filtering module 1204 is used to obtain a first set of image processing algorithms corresponding to the task type of the subtask based on a first mapping relationship between task type and image processing algorithm.
[0093] In some embodiments, the second algorithm filtering module 1205 is used to obtain a second set of image processing algorithms corresponding to the modalities of the multimodal remote sensing image based on a second mapping relationship between image modalities and image processing algorithms.
[0094] In some embodiments, the third algorithm filtering module 1206 is further configured to determine the same candidate image processing algorithms in the first image processing algorithm set and the second image processing algorithm set.
[0095] In this embodiment of the application, the same candidate image processing algorithm is obtained by performing intersection processing on the first image processing algorithm set and the second image processing algorithm set.
[0096] In the embodiments of this application, intersection processing refers to performing an intersection operation on two sets.
[0097] In some embodiments, the candidate image processing algorithm set is the set obtained by intersecting the first image processing algorithm set and the second image processing algorithm set. The algorithms in this set can meet the requirements of the current subtask and can correctly process multimodal remote sensing images.
[0098] For example, when the first image processing algorithm set includes YOLO, RetinaNet, and CenterNet, and the second image processing algorithm includes YOLO and CenterNet, the same candidate image processing algorithms include YOLO and CenterNet. In some embodiments, the algorithm evaluation module 1207 is used to determine the performance evaluation results of the candidate image processing algorithms based on the utility function, and to select the candidate image processing algorithm with the best performance evaluation result as the optimal image processing algorithm.
[0099] In some embodiments, the performance evaluation result is a quantitative value of the performance of the image processing algorithm.
[0100] In the embodiments of this application, the utility function is used to quantify the performance of the image processing algorithm.
[0101] For example, the utility function U is constructed using the following formula: (1) in, This represents the j-th image processing algorithm; , , These represent the normalized accuracy score, speed score, and resource efficiency score, respectively, which are obtained by verifying the processing accuracy, processing speed, and resource utilization of the j-th image processing algorithm through a pre-constructed validation set. , , These are the weighting coefficients.
[0102] In some embodiments, after quantifying the performance of each image processing algorithm among the candidate image processing algorithms through a utility function, the corresponding performance evaluation result is obtained, and the image processing algorithm with the best performance evaluation result is taken as the optimal image processing algorithm.
[0103] In some embodiments, the execution module 130 is further configured to perform task processing on the multimodal remote sensing image to be processed corresponding to the subtask based on the task execution sequence and the optimal image processing algorithm, and output the task execution result.
[0104] In this embodiment, the system selects the optimal image processing algorithm for each subtask in the task execution sequence. It should be noted that this application does not limit the timing of the selection of the optimal image processing algorithm. The optimal image processing algorithm for each subtask can be selected directly after the subtasks are generated; it can also be selected after the task execution sequence is generated; or it can be selected while processing the subtask.
[0105] In some embodiments, after obtaining the optimal image processing algorithm corresponding to the current subtask, the execution module 130 can perform task processing on the remote sensing image to be processed corresponding to the current subtask based on the current subtask in the task execution sequence and through the optimal image processing algorithm.
[0106] In the embodiments of this application, by introducing a first mapping relationship between task type and image processing algorithm and a second mapping relationship between image modality and image processing algorithm, the system of this application can dynamically select an optimal image processing algorithm based on the task type of the sub-task and the modality of the remote sensing image corresponding to the sub-task, thereby making the system of this application more efficient and intelligent.
[0107] In some embodiments, the execution module 130 is further configured to determine the error information of the third subtask based on the task execution result after outputting the task execution result, and send the error information of the third subtask to the cognitive module; wherein the third subtask is a subtask in which the task processing failed.
[0108] In this embodiment of the application, when a subtask fails to process, the execution module 130 will directly output the task execution result including error information.
[0109] In some embodiments, error messages refer to information generated when a subtask fails to execute successfully.
[0110] For example, error messages may include, but are not limited to, the error type, the time of the error, and a message indicating the cause of the error.
[0111] In some embodiments, the cognitive module 110 is further configured to analyze the error information of the third subtask, obtain the error analysis result, and send the error analysis result to the planning module.
[0112] In this embodiment of the application, the error analysis result is the extraction of error information to determine the cause of the error.
[0113] In some embodiments, error information is extracted using a large language model to obtain error analysis results.
[0114] For example, error analysis results include, but are not limited to, task execution timeout, insufficient memory, insufficient video memory, and algorithm non-convergence.
[0115] In some embodiments, the planning module 120 is further configured to obtain the optimal candidate execution strategy corresponding to the error analysis result based on the third mapping relationship between the analysis result and the execution strategy, and send the optimal candidate execution strategy to the execution module; wherein the optimal candidate execution strategy includes at least one or more of the following: extending the execution time threshold of the third subtask, reducing the sampling number of the multimodal remote sensing images corresponding to the third subtask, and updating the optimal image processing algorithm of the third subtask.
[0116] In this embodiment of the application, the execution strategy refers to the strategy used for the successful execution of the third subtask.
[0117] In some embodiments, task execution timeouts are avoided by extending the execution time threshold of subtasks, memory shortages are avoided by reducing the number of samples in multimodal remote sensing images, and insufficient video memory and algorithm non-convergence are avoided by changing the optimal image processing algorithm. The updating of the optimal image processing algorithm is based on image processing algorithms from a third set of image processing algorithms.
[0118] In some embodiments, the optimal candidate execution strategy refers to the best strategy for resolving the situation where the current third subtask fails to execute.
[0119] In some embodiments, the third mapping relationship is a predefined association relationship used to establish a corresponding relationship between different analysis results and corresponding execution strategies, thereby ensuring that the planning module 120 can quickly identify the execution strategy suitable for the current error analysis results.
[0120] For example, the third mapping relationship is established in the form of a table, as shown in Table 3: Table 3
[0121] For example, when the error analysis result indicates a task execution timeout, the execution strategies obtained from Table 3 include increasing the task execution time threshold by 1 second, 2 seconds, and 3 seconds. When multiple strategies exist, any one can be selected as the optimal candidate execution strategy; alternatively, the posterior probability of each strategy can be calculated, and the strategy with the highest posterior probability can be selected as the optimal candidate execution strategy.
[0122] In some embodiments, the execution module 130 is further configured to re-execute the third subtask based on the optimal candidate execution strategy.
[0123] In some embodiments, after obtaining the optimal candidate execution strategy, the execution module 130 re-executes the failed third subtask based on the optimal candidate execution strategy to resolve the situation where the subtask failed to execute, thereby ensuring that subsequent subtasks can be executed smoothly.
[0124] In this embodiment of the application, the third mapping relationship enables the generation of a corresponding execution strategy based on the error analysis results, and the execution of subtasks can be restored based on the execution strategy, thereby improving the system's adaptability and fault tolerance.
[0125] In summary, the large-scale model-agent collaborative dynamic scalable remote sensing target interpretation system provided in this application achieves a task execution sequence that simultaneously includes a first subtask executed serially and a second subtask executed in parallel by the planning module, which plans the task order of the subtasks generated by the cognition module. This allows the execution module to execute the second subtask in parallel on top of the first subtask being executed serially, thereby shortening the total execution time of the system in processing subtasks and improving the system's execution efficiency. Furthermore, the dynamic matching mechanism of the optimal image processing algorithm can further improve the efficiency of task processing; and the rescheduling mechanism for task execution failures, based on a dynamically determined execution strategy, can recover the execution of failed subtasks, improving the system's stability and fault tolerance.
[0126] Based on the above embodiments, this application also provides a dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration. This method is applied to a dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration, and the system includes a cognitive module, a planning module, and an execution module. Figure 5 As shown, the method may include the following steps: Step 101: The cognitive module decomposes the text task to be processed using a large language model, obtains the subtasks corresponding to the text task to be processed, and obtains the direct dependency relationship between every two subtasks.
[0127] In this embodiment of the application, the large language model refers to a pre-trained natural language processing model that has the capabilities of semantic understanding, logical reasoning, code generation, and knowledge integration.
[0128] In this embodiment of the application, the text task to be processed refers to the task instruction in natural language form input by the user.
[0129] In some embodiments, a subtask refers to the smallest executable unit that the large language model decomposes from the text task to be processed.
[0130] In the embodiments of this application, a direct dependency refers to a direct dependency between two subtasks.
[0131] In some embodiments, the large language model generates direct dependencies between subtasks while decomposing them into subtasks.
[0132] In some embodiments, the large language model generates task types for subtasks while decomposing them; then, it obtains the direct dependencies between subtasks through predefined direct dependencies between task types.
[0133] Step 102: The planning module plans the task order of subtasks based on the direct dependencies between subtasks and constructs a task execution sequence; wherein, the task execution sequence includes a first subtask for serial execution and a second subtask for parallel execution.
[0134] In this embodiment, after obtaining the subtasks and direct dependencies, the cognitive module sends the subtasks and direct dependencies to the planning module, which can further enable the planning module to plan the task order of the subtasks based on the direct dependencies between the subtasks and construct the task execution sequence.
[0135] In the embodiments of this application, the task execution sequence refers to the arrangement and combination of subtasks that exist simultaneously in a serial / parallel manner and have an execution order.
[0136] In some embodiments, the process of constructing a task execution sequence may include the following steps: Step 1021: Construct a directed acyclic graph based on the task execution steps of the subtasks.
[0137] In some embodiments, each subtask of this application is treated as a node in a directed acyclic graph, and the direct dependencies between tasks are treated as directed edges in the directed acyclic graph, thereby constructing the corresponding directed acyclic graph.
[0138] Step 1022: Based on the directed acyclic graph, obtain the first subtask and the second subtask using a graph algorithm; the first subtask is constructed based on subtasks with direct dependencies; the second subtask is constructed based on subtasks without direct dependencies.
[0139] In this embodiment of the application, the first subtask and the second subtask can be obtained by traversing the directed acyclic graph using a graph algorithm.
[0140] Step 1023: Generate a task execution sequence based on the first subtask and the second subtask.
[0141] Step 103: The execution module performs task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence and outputs the task execution results.
[0142] In this embodiment, after obtaining the task execution sequence, the planning module sends the task execution sequence to the execution module, which can further enable the execution module to perform task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence and output the task execution results.
[0143] In some embodiments, when processing the first subtask, they are executed sequentially; when processing the second subtask, multiple tasks are started simultaneously to achieve parallel execution of the subtasks in the second subtask.
[0144] In the embodiments of this application, the multimodal remote sensing images to be processed refer to remote sensing images from different sensor platforms, and there is a corresponding relationship with the sub-tasks.
[0145] In some embodiments, such as Figure 6 As shown, the method also includes the following steps: Step 201: The planning module obtains the first set of image processing algorithms corresponding to the task type of the subtask based on the first mapping relationship between task type and image processing algorithm.
[0146] In this embodiment of the application, after obtaining the sub-task, the planning module can further obtain the first set of image processing algorithms corresponding to the task type of the sub-task based on the first mapping relationship between the task type and the image processing algorithm.
[0147] In some embodiments, task type refers to a category based on the nature and objectives of the task.
[0148] In some embodiments, the first mapping relationship is a predefined association relationship used to establish a corresponding relationship between different task types and corresponding image processing algorithms.
[0149] Step 202: The planning module obtains the second set of image processing algorithms corresponding to the modalities of the multimodal remote sensing image based on the second mapping relationship between image modalities and image processing algorithms.
[0150] In this embodiment of the application, after obtaining the multimodal remote sensing image, the planning module can further obtain a second set of image processing algorithms corresponding to the modalities of the multimodal remote sensing image based on the second mapping relationship between image modalities and image processing algorithms.
[0151] In some embodiments, the second mapping relationship is a predefined association relationship used to establish a corresponding relationship between different image modalities and corresponding image processing algorithms.
[0152] Step 203: The planning module evaluates the algorithms based on the first image processing algorithm set and the second image processing algorithm set to obtain the optimal image processing algorithm.
[0153] In this embodiment of the application, after obtaining the first image processing algorithm set and the second image processing algorithm set, the planning module can further evaluate the algorithm based on the first image processing algorithm set and the second image processing algorithm set to obtain the optimal image processing algorithm.
[0154] In the embodiments of this application, algorithm evaluation refers to evaluating the performance of the corresponding image processing algorithm to determine the image processing algorithm with the best performance.
[0155] In some embodiments, algorithm evaluation is based on a utility function used to quantify the performance of each image processing algorithm in the intersection of a first set of image processing algorithms and a second set of image processing algorithms.
[0156] Step 205: The execution module performs task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence and the optimal image processing algorithm, and outputs the task execution results.
[0157] In this embodiment of the application, after obtaining the optimal image processing algorithm, the execution module can further perform task processing on the multimodal remote sensing images to be processed corresponding to the sub-tasks based on the task execution sequence and the optimal image processing algorithm, and output the task execution results.
[0158] In some embodiments, each subtask has its corresponding optimal image processing algorithm.
[0159] In some embodiments, after obtaining the optimal image processing algorithm corresponding to the current subtask, the remote sensing image to be processed corresponding to the current subtask can be processed based on the current subtask in the task execution sequence using the optimal image processing algorithm.
[0160] In some embodiments, such as Figure 7 As shown, the method also includes the following steps: Step 301: After outputting the task execution result, the execution module determines the error information of the third subtask based on the task execution result and sends the error information of the third subtask to the cognition module; wherein, the third subtask is the subtask that failed to process the task.
[0161] In this embodiment, when a subtask fails, the execution module directly outputs the task execution result, including error information. Step 302: The cognition module analyzes the error information of the third subtask, obtains the error analysis result, and sends the result to the planning module.
[0162] In this embodiment of the application, after receiving the error information, the cognitive module can further analyze the error information of the third subtask, obtain the error analysis result, and send the error analysis result to the planning module.
[0163] In some embodiments, error analysis results are the extraction of error information to determine the cause of the error.
[0164] In some embodiments, error information is extracted using a large language model to obtain error analysis results.
[0165] Step 303: Based on the third mapping relationship between the analysis results and the execution strategy, the planning module obtains the optimal candidate execution strategy corresponding to the error analysis results and sends the optimal candidate execution strategy to the execution module; wherein, the optimal candidate execution strategy includes at least one or more of the following: extending the execution time threshold of the third subtask, reducing the sampling number of the multimodal remote sensing images corresponding to the third subtask, and updating the optimal image processing algorithm of the third subtask.
[0166] In this embodiment of the application, after receiving the error analysis results, the planning module can further obtain the optimal candidate execution strategy corresponding to the error analysis results based on the third mapping relationship between the analysis results and the execution strategy, and send the optimal candidate execution strategy to the execution module.
[0167] In this embodiment of the application, the execution strategy refers to the strategy used for the successful execution of the third subtask.
[0168] In some embodiments, the optimal candidate execution strategy refers to the best strategy for resolving the situation where the current third subtask fails to execute.
[0169] In some embodiments, the third mapping relationship is a predefined association relationship used to establish a corresponding relationship between different analysis results and the corresponding execution strategies.
[0170] Step 304: The execution module re-executes the third subtask based on the optimal candidate execution strategy.
[0171] In this embodiment of the application, after receiving the optimal candidate execution strategy, the execution module can further re-execute the third sub-task based on the optimal candidate execution strategy.
[0172] In summary, the large-model-agent collaborative dynamic scalable remote sensing target interpretation method provided in this application, by planning the task sequence of the generated subtasks, obtains a task execution sequence that simultaneously includes a first subtask executed serially and a second subtask executed in parallel. This allows the second subtask to be executed in parallel on the basis of the serial execution of the first subtask, thereby shortening the total execution time of processing subtasks and improving the execution efficiency of the system to which this method is applied. Furthermore, the dynamic matching mechanism of the optimal image processing algorithm can further improve the efficiency of task processing; and the rescheduling mechanism when task execution fails, based on a dynamically determined execution strategy, restores the execution of failed subtasks, improving the stability and fault tolerance of the system.
[0173] Based on the above embodiments, another embodiment of this application proposes a dynamic and scalable remote sensing target interpretation system with large model-agent collaboration. The dynamic and scalable remote sensing target interpretation system with large model-agent collaboration proposed in this application takes the cognitive reasoning ability of a large language model as its core, and combines it with the task perception, planning and execution capabilities of an agent. Through the large model-agent collaboration mechanism, it realizes the automatic decomposition, dependency modeling and hybrid scheduling of complex tasks, thereby improving the system's autonomy, flexibility and resource utilization efficiency in complex task scenarios.
[0174] The following is an exemplary description of possible implementations of the dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration proposed in the embodiments of this application.
[0175] With the rapid development of remote sensing technology and the diversification of observation platforms, the scale and complexity of remote sensing data continue to grow, exhibiting characteristics of multi-source, heterogeneous, and massive volume. Remote sensing target interpretation systems, as crucial technical support for fields such as Earth observation, environmental monitoring, urban planning, and national defense, aim to automatically extract ground targets and their semantic information from multimodal remote sensing data, including visible light, infrared, SAR, and hyperspectral data. However, traditional remote sensing target interpretation systems are mostly based on predefined algorithm flows and fixed working modes, with domain experts designing fixed task processing pipelines that combine preprocessing, detection, classification, and segmentation modules in a serial manner. While this mode offers some stability when handling specific tasks, it suffers from problems such as rigid workflows, low intelligence, limited scalability, and poor human-computer interaction when facing complex, variable, and unstructured remote sensing scenarios. Furthermore, the system requires modification and redeployment of the underlying code when integrating new algorithms or data types, making it difficult to meet the practical needs of current multi-source heterogeneous data fusion and dynamic task execution. In recent years, breakthroughs in Large Language Models (LLMs) have injected new momentum into the development of intelligent remote sensing systems. LLMs, represented by the GPT series, possess powerful capabilities in natural language understanding, logical reasoning, code generation, and knowledge integration, and are gradually being introduced into the remote sensing field. GeoChat explored the application of multimodal LLMs in tasks such as remote sensing image description and visual question answering. EarthGPT achieves unified understanding and reasoning of multi-sensor, multi-task remote sensing images through visual enhancement perception mechanisms, cross-modal mutual understanding methods, and unified instruction tuning strategies. Meanwhile, agents emphasize achieving autonomous decision-making and environmental interaction through a closed loop of perception, planning, action, and learning. Against this backdrop, a technical approach integrating LLMs and autonomous agents has emerged. Tree-GPT, through a modular expert system, is applied to forestry remote sensing data analysis; Change-Agent uses a multi-level change interpretation model as the "eyes" and a large model as the "brain" to support dual-temporal image change analysis. However, these methods suffer from problems such as insufficient architectural flexibility, limited task adaptability, low resource scheduling efficiency, and poor tool and data scalability during execution, making it difficult to achieve efficient and stable operation in real-world complex scenarios.
[0176] Most remote sensing target interpretation systems adopt an "algorithm-driven" model based on predefined algorithm flows. Multimodal large language models, such as GeoChat and EarthGPT, attempt to achieve automatic description, question answering, and cross-modal reasoning of remote sensing images through visual-language fusion. However, their core still relies on the unified encoding and static task modeling of a single model, lacking task decomposition and dynamic scheduling capabilities, and making it difficult to cope with complex and unstructured scenarios. Another type of system, represented by Tree-GPT and Change-Agent, has made preliminary explorations in the direction of remote sensing intelligent agents, using modular expert systems or change interpretation models to achieve autonomous task planning and environmental interaction. However, the overall architecture lacks flexibility, has limited adaptability to multi-source data and multi-task, and the mechanisms for task collaboration, resource scheduling, and anomaly recovery are still imperfect.
[0177] The large-scale model-agent collaborative dynamic scalable remote sensing target interpretation system proposed in this application adopts a layered architecture design, such as... Figure 8 As shown, from top to bottom, it includes five core layers: user interaction layer 210, large model intelligent agent layer 220, scalable interpretation task tool layer 230, scalable multi-source heterogeneous data layer 240, and infrastructure layer 250.
[0178] The user interaction layer 210 provides user-facing service interfaces, including a user input module 2101 and a system output module 2102. These two modules support natural language input and structured system output, respectively. Through a user-friendly human-computer interaction interface, the user input module 2102 converts the user's input natural language instructions (text tasks to be processed) into an instruction format that the system can understand. The output results (task execution results) of the system output module are in a structured natural language form, including various forms such as analysis conclusions, data statistics, and visualization charts.
[0179] The large model-agent layer 220 is the cognitive center and task controller of the system. This layer uses a large language model as its cognitive core and intelligent agents as its execution entities, responsible for semantic understanding, intelligent planning, and dynamic scheduling of tasks. Through natural language parsing of user commands, it automatically completes the decomposition of complex tasks, dependency modeling, and execution strategy generation, realizing the transformation from "algorithm-driven" to "cognitive-driven." Specifically, the large model-agent layer 220 adopts a modular architecture design, including five core functional modules: perception module 2201, cognition module 2202, planning module 2203, execution module 2204, and memory module 2205, which work together to achieve intelligent processing of remote sensing target interpretation.
[0180] The perception module 2201 is responsible for receiving and preprocessing various types of information from both inside and outside the system. Its main responsibilities include: receiving and parsing natural language commands from the user input module 2101 in the user interaction layer 210, identifying the user's core needs, constraints, and implicit intentions; perceiving the overall state of the current system, including available remote sensing data, the functionality, performance, and load of tools in the scalable interpretation task tool layer; and receiving real-time feedback from the execution module 2204, such as intermediate results of tool execution, success or failure status, and error messages. Through the comprehensive perception of this multimodal and multi-source information, it provides a comprehensive and accurate basis for subsequent cognition and planning.
[0181] The cognitive module 2202 leverages the powerful reasoning capabilities of a large language model to achieve a deep understanding of complex remote sensing interpretation tasks. Based on the information input from the perception module 2201 (including error messages and natural language instructions), this module performs deep semantic understanding, identifies task types and difficulty levels, understands the matching relationship between data features and task requirements, and combines domain knowledge for reasoning analysis. It can associate the current task with historical experience and domain knowledge, outputting subtasks, direct dependencies, and error analysis results to provide intelligent support for task planning.
[0182] The planning module 2203 is primarily responsible for task planning for subtasks and formulating optimal execution strategies (task execution sequences). Based on the output of the cognition module 2202, the planning module can identify the dependencies between subtasks, automatically construct serial task chains and parallel task groups, and generate dynamically adjustable execution strategies (task execution sequences).
[0183] The execution module 2204 is the action unit for the intelligent agent to interact with the external tool environment, responsible for translating execution strategies into specific tool calls and execution actions. This module has tool call capabilities, dynamically calling corresponding algorithm tools (image processing algorithms) in the extensible interpretation task tool layer 230 through the MCP (Model Context Protocol) service interface to achieve automated task execution. Specifically, it initiates a call request to the interface layer 2301; after responding to the tool call request, the interface layer 2301 retrieves the corresponding algorithm tool (e.g., a perception algorithm) from the tool layer 2302 and returns it to the execution module 2204. It also obtains the remote sensing image corresponding to the current subtask from the extensible multi-source heterogeneous data layer 240 through the interface layer 2301, then processes the current subtask based on the called algorithm tool and the obtained remote sensing image, outputs the corresponding processing result (task execution result), and sends the processing result to the system output module 2102. During execution, this module is also responsible for monitoring the task status, handling execution anomalies, and feeding back the execution results to other modules.
[0184] The memory module 2205 is responsible for storing historical remote sensing interpretation task interaction information, learned professional knowledge, and temporary task information, including receiving natural language instructions output by the user input module 2101 and sub-tasks output by the cognitive module 2202. The memory mechanism of the intelligent agent system corresponds to the human memory model: sensory memory serves as the raw input (including multimodal information such as remote sensing data and user instructions), short-term memory is the contextual information of the current task (limited by the context window length), and long-term memory is an external knowledge base (accessible through rapid retrieval). Through the memory module 2205, the system can quickly recall historical or external experience when processing similar tasks, improving processing efficiency and accuracy.
[0185] The extensible interpretation task tool layer comprises an interface layer 2301 and a tool layer 2302. The interface layer 2301 includes a tool registration interface for registering tools, a tool invocation interface for calling tools, and a formatting interface for formatting the system's output. The tool layer 2302 integrates a rich set of remote sensing processing algorithms (image processing algorithms). The tool layer 2302 organically merges traditional perception tasks with cognitive tasks, integrating perception task tools 23021 for perception tasks and cognitive task tools 23022 for cognitive tasks. This enables the system not only to identify "what" but also to understand "why" and "how," overcoming the limitations of traditional remote sensing processing systems that primarily focus on perception-level tasks. Perception task tool 23021 incorporates various perception algorithms, including detection, recognition, segmentation, tracking, scene classification, change detection, and 3D reconstruction algorithms; cognitive task tool 23022 incorporates various cognitive algorithms, including scene understanding, relationship reasoning, behavior analysis, trajectory prediction, and decision support algorithms. The execution module 2204 can call the remote sensing processing algorithm tool in the tool layer 2302 by calling the interface in the interface layer 2301.
[0186] The scalable multi-source heterogeneous data layer 240 is responsible for the standardized processing and unified management of multi-source heterogeneous remote sensing data. It supports remote sensing data from different acquisition methods, such as ground vehicle platforms, low-altitude UAV platforms, airborne platforms, and aerospace satellite platforms. It covers various sensor data types, including visible light, infrared, SAR, hyperspectral, multispectral, and panchromatic, as well as different data formats such as images and videos. Through data preprocessing, format conversion, and quality control, it performs unified processing on data from different sources and of different types, providing standardized data services for upper-layer applications. For example, the processed remote sensing images are sent to the execution module 2204 through the interface layer 2301.
[0187] The infrastructure layer 250 provides the underlying support for the entire system, including hardware resources such as ground vehicle platforms, low-altitude UAV platforms, aircraft platforms, aerospace satellite platforms, cloud computing platforms, distributed storage systems, and GPU clusters, as well as software components such as various databases and middleware, ensuring the stable operation and efficient processing of the system. Through these platforms, hardware resources, and software components, corresponding remote sensing images can be acquired and transmitted to the scalable multi-source heterogeneous data layer 240.
[0188] This application's system, based on a large model-agent layer, utilizes a multi-task collaborative mechanism to transform macroscopic and fuzzy user natural language commands into a dynamic orchestration process of precise, ordered, and executable remote sensing target interpretation operations. This overcomes the rigidity and inefficiency of traditional fixed workflows when facing complex and variable tasks, achieving a shift from traditional fixed workflows to dynamic adaptive task orchestration. The core concept of this mechanism is to treat complex remote sensing target interpretation tasks as a decomposable, recombinable, and optimizable dynamic system. Through intelligent task decomposition, dependency analysis, parallel scheduling, and collaborative optimization, it maximizes the system's processing efficiency and interpretation accuracy.
[0189] like Figure 9 As shown, the multi-task collaboration mechanism includes the following steps: Step 401, input complex task (text task to be processed).
[0190] For example, a complex task is "to analyze the dynamics of ship targets in the current port area".
[0191] Step 402: Decompose the complex task.
[0192] In this embodiment of the application, step 402 may include the following steps: Step 4021: Identify subtasks.
[0193] Step 4022: Analyze direct dependencies.
[0194] Step 403: Construct the execution strategy (task execution sequence).
[0195] Step 404: Coordinate and schedule multiple tasks.
[0196] In this embodiment of the application, step 404 may include the following steps: Step 4041: Generate intermediate results.
[0197] For example, such as Figure 10 As shown, it includes the following steps: Step 40411, Port Area Identification.
[0198] Step 40412, Ship target detection.
[0199] Step 40413, Ship type identification.
[0200] Step 40414, Ship distribution analysis.
[0201] Step 40415, summarize the number of ships.
[0202] Step 40416: Obtain historical ship data.
[0203] Step 40417: Generate intermediate results In this embodiment, steps 40413 to 40415 are executed in parallel; and during the parallel execution of these three steps, intermediate results are generated for other steps to share (e.g., step 40414). These intermediate results can also be used to predict the future trend of the port, and the predicted results are also used to obtain the interpretation results (task execution results).
[0204] Step 4042, exception handling.
[0205] For example, such as Figure 11 As shown, it includes the following steps: Step 40421, abnormal diagnosis.
[0206] Step 40422, task rescheduling.
[0207] In terms of task decomposition, the system first uses the cognitive module to perform semantic parsing and domain knowledge reasoning on the complex task input by the user, decomposing the original task T (the text task to be processed) into several sets of subtasks. Let the task decomposition function be... Each subtask It includes attributes such as task type, input data requirements, output format specifications, and execution constraints. The decomposition process follows the principle of task atomicity, ensuring that each subtask corresponds to a specific algorithm or tool combination in the tool layer, while maintaining the logical integrity between tasks.
[0208] When analyzing the dependencies between subtasks, a Directed Acyclic Graph (DAG) model is used to describe the execution dependencies between subtasks. Let the dependency graph be... vertex set Represents all subtasks, edge set This represents the dependency relationship between tasks. A directed edge. This indicates the task. Execution depends on the task For example, "ship target detection and identification" must be completed before "ship trajectory tracking" can proceed. The system determines the execution order of tasks through algorithms such as topology sorting and identifies task groups that can be executed in parallel. The strength of the dependency between tasks can be represented by a weight function. This indicates that the weight values reflect the tightness of the dependency relationship and the importance of data transfer.
[0209] Based on the dependency analysis results, the planning module constructs a hybrid execution strategy (task execution sequence) that combines serial task chains and parallel task groups. For subtasks without dependencies, the system organizes them into parallel task groups. ,in To fully utilize the system's computing resources, for task sequences with dependencies, the system constructs a serial task chain. This ensures the correct delivery of data. Then, a hybrid execution strategy is built based on parallel task groups and serial task chains.
[0210] The system in this application also establishes a shared data (intermediate result) cache, result transmission channel, and exception handling mechanism through intermediate result sharing and inter-task collaborative optimization mechanisms to achieve efficient data exchange between subtasks. The system maintains a global state space to record the metadata of intermediate results generated by subtasks during execution. These represent the ID of the subtask that generated the intermediate result, the data type of the intermediate result, the specific data value of the intermediate result, and the timestamp of the intermediate result, respectively. The intermediate result refers to temporary data generated during task execution, which may be used by other tasks.
[0211] When the execution module monitors a certain subtask If execution fails, exception handling is performed. The error message is as follows. The error message is then relayed to the perception module. The data is sent to the cognitive module, which then processes it using a large language model. The system analyzes and diagnoses the reasons for failure (such as task execution timeout, insufficient memory, insufficient video memory, etc.) and sends the analysis results to the planning module. The planning module stores a mapping relationship between failure reasons and strategies. After obtaining the failure reason, the planning module uses this mapping relationship to obtain the corresponding optional strategy. Employing a probabilistic reasoning-based exception handling and rescheduling method, it selects the optional strategy with the highest posterior probability of successfully repairing and completing the task using the following formula. (Optimal candidate execution strategy): (2) in, Represent each optional strategy In a given exception The posterior probability can be obtained using the following formula: (3) in, Let be the likelihood probability, representing the possible strategies to be adopted in history. Post-abnormality The probability of a successful repair is calculated from historical execution logs. Let be the prior probability, representing the possible strategies. The initial confidence level; This is a normalization constant and does not affect policy comparison; optional policies. This includes retrying the task using different parameters (extending the timeout threshold, sampling smaller batches of remote sensing images). The system may invoke a similar backup tool or request user intervention if it cannot resolve the issue automatically. Finally, after all key subtasks have been successfully executed, the system enters the result aggregation and comprehensive analysis phase. The execution module aggregates the final outputs of all subtasks into the memory module. The cognition module then uses logical reasoning and natural language generation to correlate, integrate, and refine these scattered results.
[0212] Taking the interpretation of remote sensing images of port areas as a typical complex task scenario, the large model-agent decomposes the natural language instructions input by the user into 8 sub-tasks (T1-T8) with clear dependencies. Two different execution processes are designed: "traditional fixed workflow" and "dynamic collaborative workflow": (1) Traditional fixed workflow: This process simulates traditional remote sensing information processing software, and executes the 8 sub-tasks linearly in a strict logical order. Its execution path is as follows: In this mode, there is no parallel scheduling between tasks, and the exception handling mechanism is set to "manual intervention to try retry after failure, and the entire workflow is terminated if manual intervention fails". This mechanism represents a common, non-adaptive fault-tolerant strategy in related technologies. (2) Dynamic collaborative workflow: This process adopts the system proposed in this application, and performs task scheduling based on DAG and task dependency relationship. According to the task dependency relationship, the system generates and executes a hybrid scheduling strategy: .in, and The system comprises two parallel task groups, enabling concurrent execution of tasks within each group that have no direct dependencies. Furthermore, the process includes a built-in exception handling module to diagnose and reschedule failed subtasks. Table 4 shows a performance comparison between the traditional fixed workflow and the dynamic collaborative workflow of this application. Table 4
[0213] From a macro-level system perspective, the multi-task collaborative mechanism constructed in this application maintains high efficiency while avoiding redundant calculations through an intermediate result sharing mechanism. T5 shares the detection and identification results of T2 and T3, reducing the overhead of redundant data processing. Regarding parallel execution efficiency verification, the system uses precise timestamp recording to measure the start time, execution duration, and completion time of each subtask. The total execution time in the traditional serial mode is 342 seconds, while the collaborative mechanism reduces the total execution time to 198 seconds through parallel scheduling, improving efficiency by 72.7%. In abnormal situations such as insufficient GPU resources, tool call timeouts, and incorrect data formats, the multi-task collaborative mechanism module proposed in this application quickly identifies the abnormality type through an anomaly diagnosis module and automatically calls backup tools, significantly reducing the anomaly recovery time from 156 seconds to 33 seconds. Simultaneously, the collaborative mechanism, through parallel task scheduling, makes fuller use of system resources, increasing the average CPU utilization to 67.8% and GPU utilization to 71.3%, improving resource efficiency by approximately two times. In summary, the multi-task collaboration mechanism, through task decomposition, dependency analysis, and parallel scheduling driven by a large model-agent collaboration, can significantly improve the system's execution efficiency and resource utilization while ensuring interpretation accuracy.
[0214] In related technologies, the binding relationship between remote sensing data (remote sensing images to be processed) and processing tools is usually static and fixed, which greatly limits the flexibility and scalability of the system. This application designs and implements a data-tool adaptive matching mechanism, such as... Figure 12 As shown, the mechanism includes the following steps: Step 501: Input remote sensing data.
[0215] Step 502: Analyze the characteristics of remote sensing data.
[0216] In some embodiments, the analyzed remote sensing data features include the modalities of the remote sensing data.
[0217] Step 503: Automatically find tools.
[0218] In some implementations, tools are used to automatically find the modal that corresponds to the current remote sensing data.
[0219] Step 504, optimal matching of data and tools.
[0220] In some embodiments, the identified tools are evaluated for performance, and the best-performing tool is used to process the remote sensing image.
[0221] For example, the YOLO algorithm can be used for target detection tasks and visible light remote sensing images; the ResNet algorithm can be used for type recognition tasks and SAR remote sensing images; the Transfoemer algorithm can be used for corresponding scene analysis tasks and multispectral remote sensing images; and the U-Net algorithm can be used for facility recognition tasks and panchromatic remote sensing images.
[0222] The core concept of this mechanism lies in decoupling the "data-task-tool" triad. By quantitatively describing and modeling data characteristics, task requirements, and tool capabilities, a dynamic and intelligent ternary matching framework is constructed to achieve on-demand optimal allocation of interpretation resources. To achieve accurate matching, the mechanism first rapidly extracts metadata and basic attributes from the input data to form a structured data feature vector. For a given remote sensing dataset... Its eigenvectors Defined as:
[0223] in, This indicates the data source platform (e.g., satellite, drone, etc.). Indicates the sensor type (visible light, SAR, multispectral, etc.). Indicates the data storage format. Indicates spatial resolution. Indicates band information, Indicates geographical spatial range, This indicates the data acquisition timestamp. The feature extraction process employs a lightweight analysis method, primarily reading the data file header information and metadata tags to avoid in-depth analysis of the data content and ensure processing efficiency.
[0224] The system provides each tool in the tool library Establish standardized capability description documents It uses key-value pairs for storage, that is, it establishes a tool T and a capability description file. Mapping relationship (second mapping relationship), capability description file The core content includes:
[0225] in, This indicates the core functionality of the tool, which directly corresponds to the subtask type. This indicates the modality of remote sensing data to which the tool is applicable. This indicates detailed input specifications, including data format, size range, etc. Indicates the output format. This represents performance metrics, including accuracy, speed, and resource consumption. In this way, tool capabilities are made explicit and structured, enabling large model-agent systems to "understand" the purpose and limitations of each tool. When new algorithmic tools are introduced, only the capability description file needs to be configured. Once registered in the tool library, it can be automatically discovered and invoked.
[0226] The system also establishes a primary mapping relationship between each tool and the task type of its subtasks.
[0227] The system uses a matching algorithm to select the optimal tool for the current subtask. This algorithm employs a two-stage selection strategy: a candidate selection stage and an optimization selection stage. The candidate selection stage performs rapid filtering based on hard constraints. For each subtask... (type is) ) and data (The remote sensing image to be processed) The system filters the candidate toolset (third image processing algorithm set) using the following formula: (4) The Compatible function is used to determine the compatibility between the data sensor type and the modalities supported by the tool. The logic corresponding to formula (4) is as follows: through the first mapping relationship, obtain the first tool set (first image processing algorithm set) corresponding to the task type of the current subtask; through the second mapping relationship, obtain the second tool set (second image processing algorithm set) corresponding to the modality of the remote sensing image corresponding to the current subtask; perform the intersection operation on the first tool set and the second tool set to obtain the candidate tool set.
[0228] In the optimization selection phase, the tools in the candidate toolset are scored using a utility function to select the optimal tool (optimal image processing algorithm). The utility function U is shown in formula (1), and its weight coefficients are dynamically adjusted by the large model-agent based on the implicit intent of the user's instructions (such as "quick overview" or "requires the highest accuracy") or the current system load. For example, if the user's instructions emphasize "accurate recognition," then... It will be assigned a higher value.
[0229] Then, the tool with the highest score is selected as the optimal tool using the following formula. (Optimal image processing algorithm): (5) After determining the optimal tool, the system performs data adaptation to ensure seamless integration. The adaptation process includes: (1) Format conversion: When When needed, it automatically calls a format conversion tool (such as the GDAL library) to perform the conversion.
[0230] (2) Preprocessing adaptation: Perform necessary operations according to the tool input specifications, including size adjustment, normalization, etc. The preprocessing pipeline is defined as: (6) in, Indicates the first One preprocessing operation.
[0231] (3) Interface standardization: unify the data transmission interface to ensure that the tool can correctly read and process the data.
[0232] The data-tool adaptive matching mechanism effectively extends the cognitive planning capabilities of the large model-agent to the specific resource scheduling level through standardized modeling of data and tools, and a two-stage adaptive matching algorithm. The mechanism's dynamism is reflected in its real-time, on-demand, and context-aware matching decisions; its scalability is demonstrated by its plug-and-play support for new data and tools, working closely with the multi-task collaboration mechanism to jointly constitute the system's adaptive scalability.
[0233] From a micro-level system perspective, the effectiveness of the data-tool adaptive matching mechanism module determines the system's flexibility and intelligence level when facing multi-source heterogeneous data and an ever-evolving tool library. Building upon the port area target interpretation using a "multi-task collaborative mechanism," this application further introduces new data types and tool algorithms. To further verify the system's scalability, this application, in addition to the existing visible light satellite data, introduces multi-source heterogeneous data such as SAR satellite data (Sentinel-1, 10m resolution), multispectral satellite data (Landsat-8, 30m resolution), and high-resolution UAV data (0.1m resolution). Simultaneously, the tool library adds new tools such as the Transformer-based target detection algorithm DETR, a diffusion model-based image enhancement tool, and a CFAR detector specifically for SAR data. Based on these tools, the performance of the data-tool adaptive matching mechanism in this application is compared with that of fixed matching methods in related technologies, as shown in Table 5. Table 5
[0234] As shown in Table 5, in terms of matching accuracy and reliability, the data-tool adaptive matching mechanism shows a significant improvement of 28.6% in matching accuracy compared to the traditional fixed matching mechanism. This indicates that the adaptive mechanism can more accurately select appropriate tools based on data characteristics and task context, thus laying the foundation for subsequent target interpretation. The difference between the two mechanisms is particularly pronounced in terms of system scalability. When integrating new tools, the adaptive mechanism can integrate code modification, compilation, and testing, reducing the time from 4.0 hours to 0.1 hours (approximately 6 minutes), achieving a leap in integration efficiency. When dealing with new data types, the adaptive mechanism achieves a 90% success rate in adapting to new data types, far exceeding the 50% of the traditional mechanism. This is due to its ability to analyze the essential characteristics of the data, rather than relying on preset, limited rules. This adaptive advantage is further amplified when handling complex tasks involving multiple data sources, increasing the success rate of heterogeneous data processing from 40% to 80%.
[0235] In summary, the adaptive matching mechanism, based on collaborative decision-making between the large model and the agent, improves resource utilization efficiency from 40% to 60%. It can make more reasonable resource scheduling according to task requirements (such as accuracy priority or speed priority) and tool characteristics (such as resource consumption), avoiding resource competition and waste, and improving the overall efficiency of the system. It is a key technical support for building the next generation of intelligent and evolvable remote sensing target interpretation systems.
[0236] This application introduces a multi-task semantic parsing and dependency modeling mechanism within a large model-agent collaborative framework. By constructing a directed acyclic graph of task dependencies and combining it with a dynamic scheduling strategy for agents, it achieves automatic decomposition and parallel execution of complex remote sensing tasks, significantly improving the system's task processing efficiency. This mechanism enables intermediate result sharing and anomaly self-recovery during task execution, ensuring stable system operation under high load and multi-task conditions. Compared to existing state-of-the-art technologies, the overall task execution time is reduced by approximately 70%, CPU and GPU resource utilization is nearly doubled, and the task success rate is increased to 90%.
[0237] The data-tool adaptive matching mechanism proposed in this application achieves adaptive matching and optimal combination of multi-source heterogeneous remote sensing data and multi-type algorithm tools by constructing a semantic embedding model of "data features - task requirements - tool capabilities," thereby significantly reducing the costs of manual configuration and tool switching. The system can automatically select the optimal algorithm tool and schedule resources based on task semantics, eliminating the need for reconfiguration or deployment when new algorithms or data types are introduced. Compared with existing best-practice solutions, the tool selection accuracy is improved by approximately 20%, the tool invocation success rate is increased from 82% to 95%, and the average configuration time is reduced from 4 hours to 0.1 hours, significantly enhancing the system's dynamic scalability.
[0238] This application integrates the cognitive reasoning capabilities of a large language model with the perception and action capabilities of an intelligent agent to form a closed-loop autonomous intelligent system encompassing task understanding, planning, and execution. This system possesses semantic understanding, task planning, and self-optimization capabilities. The system can dynamically generate multi-layered remote sensing tasks based on user natural language commands and adaptively adjust and select optimal strategies during execution, thereby significantly improving the intelligence level and decision-making accuracy of remote sensing target interpretation. Compared to existing state-of-the-art technologies, this application's system improves task response accuracy by approximately 15% and multimodal fusion efficiency by 30% in complex multimodal task scenarios, demonstrating higher intelligence and practical value.
[0239] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps; or steps from different embodiments may be combined into a new technical solution.
[0240] It should be noted that the module division in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0241] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause the interpretation system to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0242] This application provides a decoding system. Figure 13 This is a schematic diagram of the structure of the decoding system provided in the embodiments of this application, such as... Figure 13As shown, the decoding system 2000 includes a memory 1001 and a processor 1002. The memory 1001 stores a computer program that can run on the processor 1002. When the processor 1002 executes the program, it implements the steps in the method provided in the above embodiments.
[0243] It should be noted that the memory 1001 is configured to store instructions and applications executable by the processor 1002, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 1002 and the various modules in the decoding system 1000. It can be implemented by flash memory or random access memory (RAM).
[0244] This application also provides a computer-readable storage medium for storing computer programs.
[0245] Optionally, the computer-readable storage medium can be applied to the decoding system in the embodiments of this application, and the computer program causes the processor to execute the various methods in the embodiments of this application, which will not be described in detail here for the sake of brevity.
[0246] This application also provides a computer program product, including computer program instructions.
[0247] Optionally, the computer program product can be applied to the decoding system in the embodiments of this application, and the computer program instructions cause the processor to execute the various methods in the embodiments of this application, which will not be described in detail here for the sake of brevity.
[0248] This application also provides a computer program.
[0249] Optionally, the computer program can be applied to the decoding system in the embodiments of this application. When the computer program runs on the processor, it causes the processor to execute the various methods in the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0250] It should be noted that the descriptions of the decoding system, storage medium, computer program product, and computer program embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the decoding system, storage medium, computer program product, and computer program embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0251] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0252] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0253] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0254] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0255] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0256] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0257] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0258] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the decoding system to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0259] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0260] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0261] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0262] The above are merely embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0263] It should be understood that if this disclosure references any user data and personal information (including but not limited to device information, behavioral data, location information, etc.) and before applying the technical solutions described in the embodiments of this disclosure, the relevant products or services should comply with the laws and regulations concerning the protection of user data and personal information, strictly process users' personal information and data in accordance with the provisions of applicable laws and regulations throughout the entire data processing lifecycle, follow the principles of legality, legitimacy, necessity, good faith, openness, and transparency, and adopt reasonable privacy design schemes and technical measures to ensure the security of user data and personal information, protect users' legitimate rights and interests, and prevent the risks of leakage, theft, or tampering of user data and personal information.
[0264] Specifically, the company must publish and display its privacy policy in a prominent position on the user interface, clearly informing users of the types, purposes, uses, and methods of processing personal information, as well as other matters that should be disclosed as required by laws and regulations; obtain users' prior informed consent or explicit authorization regarding data processing through user-initiated interaction (such as confirmation pop-ups); process or store user data securely within the legally required timeframe; adopt a series of security technologies and management measures, including but not limited to data encryption and access control; share and transfer user data within the scope permitted by law and in a legally required manner; and process user rights, including the rights to query, access, correct, delete, withdraw authorization and consent, cancel registration, and obtain copies of personal information, within the legally required timeframe.
Claims
1. A dynamic and scalable remote sensing target interpretation system based on large-scale model-agent collaboration, characterized in that, The large-scale model-agent collaborative dynamic scalable remote sensing target interpretation system includes a cognitive module, a planning module, and an execution module; among which... The cognitive module is used to decompose the text task to be processed using a large language model, obtain the sub-tasks corresponding to the text task to be processed, and obtain the direct dependency relationship between every two sub-tasks. The planning module is used to plan the task order of the subtasks based on the direct dependencies between the subtasks and construct a task execution sequence; wherein the task execution sequence includes a first subtask for serial execution and a second subtask for parallel execution. The execution module is used to perform task processing on the multimodal remote sensing image to be processed corresponding to the sub-task based on the task execution sequence, and output the task execution result.
2. The large-scale model-agent collaborative dynamic scalable remote sensing target interpretation system according to claim 1, characterized in that, The planning module includes a relationship building module, a task acquisition module, and a sequence generation module; The relationship building module is used to construct a directed acyclic graph based on the task execution steps of the subtask; The task acquisition module is used to acquire the first subtask and the second subtask based on the directed acyclic graph using a graph algorithm; wherein the first subtask is constructed based on the subtasks that have the direct dependency relationship; and the second subtask is constructed based on the subtasks that do not have the direct dependency relationship. The sequence generation module is used to generate the task execution sequence based on the first subtask and the second subtask.
3. The large-model-agent collaborative dynamic scalable remote sensing target interpretation system according to claim 2, characterized in that, The planning module is further configured to: obtain a first set of image processing algorithms corresponding to the task type of the subtask based on a first mapping relationship between task type and image processing algorithm; obtain a second set of image processing algorithms corresponding to the modality of the multimodal remote sensing image based on a second mapping relationship between image modality and image processing algorithm; and perform algorithm evaluation based on the first set of image processing algorithms and the second set of image processing algorithms to obtain the optimal image processing algorithm. The execution module is further configured to perform task processing on the multimodal remote sensing image to be processed corresponding to the sub-task based on the task execution sequence and the optimal image processing algorithm, and output the task execution result.
4. The large-scale model-agent collaborative dynamic scalable remote sensing target interpretation system according to claim 3, characterized in that, The planning module is also used to determine the same candidate image processing algorithms in the first image processing algorithm set and the second image processing algorithm set; The performance evaluation results of the candidate image processing algorithms are determined based on the utility function, and the candidate image processing algorithm with the best evaluation result is selected as the optimal image processing algorithm.
5. The large-model-agent collaborative dynamic scalable remote sensing target interpretation system according to claim 3 or 4, characterized in that, The execution module is further configured to, after outputting the task execution result, determine the error information of the third subtask based on the task execution result, and send the error information of the third subtask to the cognition module; wherein, the third subtask is a subtask in which the task processing failed; The cognitive module is also used to analyze the error information of the third subtask, obtain the error analysis result, and send the error analysis result to the planning module; The planning module is further configured to obtain the optimal candidate execution strategy corresponding to the error analysis result based on the third mapping relationship between the analysis result and the execution strategy, and send the optimal candidate execution strategy to the execution module; wherein the optimal candidate execution strategy includes at least one or more of the following: extending the execution time of the third sub-task, reducing the sampling number of the multimodal remote sensing images corresponding to the third sub-task, and updating the optimal image processing algorithm of the third sub-task; The execution module is also used to re-execute the third sub-task based on the optimal candidate execution strategy.
6. A dynamic and scalable remote sensing target interpretation method based on large model-agent collaboration, characterized in that, A dynamic and scalable remote sensing target interpretation system applied to large model-agent collaboration, the large model-agent collaborative dynamic and scalable remote sensing target interpretation system including a cognition module, a planning module, and an execution module; the method includes: The cognitive module decomposes the text task to be processed using a large language model, obtains the sub-tasks corresponding to the text task to be processed, and obtains the direct dependency relationship between every two sub-tasks. The planning module performs task order planning on the subtasks based on the direct dependencies between them, and constructs a task execution sequence; wherein, the task execution sequence includes a first subtask for serial execution and a second subtask for parallel execution; The execution module performs task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence, and outputs the task execution results.
7. The method according to claim 6, characterized in that, The method further includes: The planning module obtains a first set of image processing algorithms corresponding to the task type of the subtask based on a first mapping relationship between task type and image processing algorithm. The planning module obtains a second set of image processing algorithms corresponding to the modalities of the multimodal remote sensing image based on the second mapping relationship between image modalities and image processing algorithms. The planning module performs algorithm evaluation based on the first image processing algorithm set and the second image processing algorithm set to obtain the optimal image processing algorithm. The execution module performs task processing on the multimodal remote sensing images corresponding to the sub-tasks based on the task execution sequence, and outputs the task execution results, including: The execution module performs task processing on the multimodal remote sensing image corresponding to the subtask based on the task execution sequence and the optimal image processing algorithm, and outputs the task execution result.
8. The method according to claim 6 or 7, characterized in that, The method further includes: After outputting the task execution result, the execution module determines the error information of the third subtask based on the task execution result, and sends the error information of the third subtask to the cognition module; wherein, the third subtask is a subtask that failed to process the task. The cognitive module analyzes the error information of the third subtask, obtains the error analysis results, and sends the error analysis results to the planning module; The planning module obtains the optimal candidate execution strategy corresponding to the error analysis result based on the third mapping relationship between the analysis result and the execution strategy, and sends the optimal candidate execution strategy to the execution module; wherein, the optimal candidate execution strategy includes at least one or more of the following: extending the execution time threshold of the third sub-task, reducing the sampling number of the multimodal remote sensing images corresponding to the third sub-task, and updating the optimal image processing algorithm of the third sub-task; The execution module re-executes the third sub-task based on the optimal candidate execution strategy.
9. A dynamic and scalable remote sensing target interpretation system based on large model-agent collaboration, comprising a memory and a processor, wherein the memory stores a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 6-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 6-8.