A language model driven task resolution and dynamic interaction method
By using a language model-driven task parsing and dynamic interaction method, the task parsing capability and multi-device collaboration efficiency of the natural resource survey system in complex environments have been improved. This has solved the shortcomings of traditional systems in complex and ever-changing environments and enabled efficient and safe execution of reconnaissance tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-28
AI Technical Summary
Traditional natural resource survey systems suffer from limited task analysis capabilities, insufficient dynamic decision-making, low efficiency of multi-device collaboration, and poor environmental adaptability in complex and ever-changing environments, making it difficult to meet the needs of high efficiency, accuracy, and autonomy.
A language model-driven task parsing and dynamic interaction approach is adopted. A large model-driven intelligent engine parses task instructions, designs a conflict detection mechanism, collects multimodal data in real time for dynamic decision-making, optimizes task paths, and combines multidimensional obstacle avoidance algorithms to achieve vehicle-machine collaboration. Multimodal large language models and digital twin technology are used to improve environmental adaptability.
It improves the system's environmental adaptability and task response speed, enhances the efficiency of multi-device collaboration, reduces manual operation costs and risks, and enables efficient and safe reconnaissance missions in complex scenarios.
Smart Images

Figure CN122469683A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language models, and more specifically to a language model-driven method for task parsing and dynamic interaction. Background Technology
[0002] In the field of natural resource surveys and monitoring in complex and challenging areas, traditional manual methods face numerous challenges, including high costs, significant safety risks, and limited data acquisition. To address these challenges, with advancements in technology, particularly the rapid development of drones and unmanned vehicles, vehicle-mounted natural resource survey systems integrating multimodal sensors have emerged and are gradually becoming an effective means of solving this problem. For example, the Landexplorer-1 system developed by Professor Wang Lizhe's team at China University of Geosciences (Wuhan) integrates advanced geological survey equipment such as retractable drones and quadruped robots through a three-dimensional collaborative architecture of "vehicle-air-ground," enabling flexible, mobile, and efficient natural resource field surveys, significantly reducing manual operation costs and the safety risks for natural resource survey personnel.
[0003] However, while these systems have improved the efficiency and safety of natural resource surveys to some extent, they still have many shortcomings when facing complex and ever-changing natural environments and diverse mission requirements. In particular, traditional systems often struggle to meet the demands for high efficiency, accuracy, and autonomy in areas such as mission analysis, dynamic decision-making, multi-device collaboration, and environmental adaptability.
[0004] Specifically, the shortcomings of existing technologies are mainly reflected in the following aspects:
[0005] Limited task parsing capabilities: Traditional systems often rely on manual parsing and instruction conversion when receiving task instructions, making it difficult to accurately understand all the intentions and constraints in complex multimodal task instructions. This leads to low task execution efficiency and may even cause errors due to misunderstanding of instructions.
[0006] Insufficient dynamic decision-making capability: In complex and ever-changing natural environments, traditional systems lack real-time dynamic decision-making capabilities. They cannot adjust task allocation and execution strategies in a timely manner according to environmental changes and task requirements, resulting in poor task execution performance, or even failure due to their inability to adapt to environmental changes.
[0007] Low efficiency in multi-device collaboration: In multi-device collaborative operation scenarios, traditional systems often struggle to achieve efficient collaboration. Problems such as communication delays between devices and unreasonable task allocation frequently occur, seriously affecting overall operational efficiency and safety.
[0008] Poor environmental adaptability: Faced with complex and ever-changing natural environments, traditional systems often lack effective environmental perception and adaptability. They struggle to autonomously complete tasks in unknown or dynamically changing environments, limiting their application in complex scenarios.
[0009] To address the aforementioned issues, the applicant proposes a language model-driven method for task parsing and dynamic interaction. Summary of the Invention
[0010] The purpose of this invention is to provide a language model-driven task parsing and dynamic interaction method to solve the problems in the prior art.
[0011] To achieve the above objectives, the present invention provides the following technical solution: a language model-driven task parsing and dynamic interaction method, comprising the following steps:
[0012] Instruction fusion and conflict resolution steps:
[0013] A large model-driven intelligent engine is used to parse and extract key information from task instructions given by users through various input methods (such as natural language, voice, map annotation, etc.);
[0014] Design a conflict detection mechanism based on semantic similarity matching to identify and resolve potential conflicts in instructions, including conflicting task priorities or resource competition.
[0015] When command conflicts are detected, they are corrected in real time through dialogue interaction to ensure the reasonable allocation and execution of tasks;
[0016] Dynamic decision-making and task reallocation steps:
[0017] Collect and analyze multimodal data (such as traffic flow, weather changes, and terrain information) in real time to extract key data for risk prediction;
[0018] Generate task replanning suggestions. When the environment changes or unexpected situations occur during task execution, the feasibility of the task is reassessed by combining the task's historical records, current resource status, and unexpected events.
[0019] Based on task priority, resource availability, and task progress, the task allocation scheme is dynamically adjusted to ensure the smooth execution and efficient completion of tasks.
[0020] Optionally, it also includes vehicle-to-vehicle (V2V) collaborative arc routing optimization steps:
[0021] The design task coding method is to represent the combination of flight paths of the UAV during the execution of the task by defining an appropriate coding structure;
[0022] A stochastic variable neighborhood descent search algorithm (RVND-SA) combined with simulated annealing mechanism is used for path optimization, and the task path is continuously adjusted and optimized through multiple neighborhood search operators;
[0023] By leveraging the global optimization capabilities of simulated annealing, local optima are avoided, ensuring that the optimal task execution path is found in large-scale patrol missions.
[0024] Optionally, it also includes dynamic obstacle avoidance and path correction steps involving air and ground coordination:
[0025] Automatically select the appropriate path planning method based on the availability of the data source;
[0026] When a 3D data source is available, the 3D tangent method (3D STG) is used for path planning, taking into account 3D spatial factors such as the height difference of obstacles, terrain undulations and buildings;
[0027] When the 3D data source is insufficient or unavailable, 2D obstacle avoidance algorithms, such as the Static Elliptical Tangent Method (SETG-TG) and the Dynamic Elliptical Tangent Method (DETG-TG), are used to plan the path based on the 2D location data.
[0028] Optionally, it also includes a PPPRTK-assisted scene reconstruction step:
[0029] Build a collaborative measurement platform integrating multiple sensors, including lidar, IMU, optical camera, and GNSS module;
[0030] A fast, tightly coupled laser inertial odometry method is adopted to improve the robustness and accuracy of real-time positioning and mapping in high dynamic environments;
[0031] The Gaussian real-time rendering algorithm is used to achieve efficient visualization in complex dynamic environments.
[0032] Optional steps include scene understanding and intelligence generation:
[0033] Receive fused multimodal data, including visual features after image encoding and linguistic features after text encoding, and perform preliminary analysis through the basic inference stage;
[0034] The data enters the core processing stage, including spatial perception, language understanding, scene reasoning and data reasoning, to optimize the rich features of multimodal data;
[0035] During the task decoding phase, detailed task instructions or intelligence outputs are generated through task cleaning, online target path planning, dynamic adjustment, and visualization.
[0036] Based on the method described above, the large model-driven intelligent engine further includes:
[0037] The intelligent task self-generation module enables task decomposition, task modeling, task planning, task calibration, and task combination.
[0038] A general model library building module that enables model building, model description, model matching, and model testing;
[0039] Based on LangChain's intelligent task processing module, it enables task scheme management, task prompt construction, and intelligent task execution.
[0040] Optionally, it also includes an unmanned equipment interaction module, which further includes:
[0041] Dynamic behavior planning functionality includes control command parsing, action space, and path planning;
[0042] Real-time data transmission capabilities, including data acquisition, data storage, and management;
[0043] The group collaborative scheduling function enables autonomous collaboration and global optimization of the drone swarm.
[0044] Beneficial effects: Improved environmental adaptability: Through multimodal large language models and digital twin technology, the system can analyze complex environments in real time and dynamically adjust task strategies to adapt to various complex scenarios.
[0045] Accelerate task response speed: Real-time data processing and intelligent decision-making mechanisms significantly shorten task response time and improve operational efficiency.
[0046] Enhance multi-device collaboration efficiency: Through vehicle-machine collaborative planning and dynamic obstacle avoidance technology, achieve efficient collaborative operation between drones and ground vehicles, and improve overall collaboration efficiency.
[0047] Reduced manual operation costs and risks: Autonomous intelligent reconnaissance systems reduce reliance on manual labor, lower field safety risks, and reduce operating costs. Attached Figure Description
[0048] Figure 1 This is a diagram illustrating the overall architecture of an embodiment of the present invention.
[0049] Figure 2 This is a system interface diagram of an embodiment of the present invention;
[0050] Figure 3 This is a data flow diagram of an embodiment of the present invention;
[0051] Figure 4 This is a diagram illustrating the composition of a large-model-driven intelligent engine system according to an embodiment of the present invention.
[0052] Figure 5 This is a system architecture diagram of a large model-driven intelligent engine according to an embodiment of the present invention. Detailed Implementation
[0053] The preferred embodiments of the present invention are described below with reference to the accompanying drawings to make the technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0054] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0055] This invention relates to a language model-driven task parsing and dynamic interaction method, which has significant application value in the development of autonomous intelligent locomotive-assisted unmanned reconnaissance equipment for digital twin maps. Its core lies in achieving real-time dynamic collaboration between hardware clusters and software decision-making in complex scenarios through a multimodal large language model (MLLM), thereby comprehensively improving the environmental adaptability, task response speed, and multi-device collaboration efficiency of the reconnaissance system.
[0056] In the technical field, this invention focuses on the collaborative operation of autonomous intelligent vehicles and unmanned reconnaissance equipment, aiming to solve the high-cost and high-risk problems faced in tasks such as natural resource surveys and monitoring in complex and challenging areas through advanced technological means. Traditional reconnaissance systems are often limited by fixed task procedures and limited environmental perception capabilities, making it difficult to adapt to complex and ever-changing task scenarios. This invention, however, constructs a fully autonomous collaborative reconnaissance system with a multimodal large language model as its core engine, achieving deep integration of hardware clustering and software decision-making, providing strong support for the efficient execution of reconnaissance missions.
[0057] To achieve the aforementioned objectives, this invention proposes a novel technical solution. This solution, centered on a multimodal large language model, constructs a collaborative service system that deeply integrates artificial intelligence and spatial intelligence. Through the coupling of the large language model and multimodal intelligent technology, a closed-loop capability system encompassing spatial perception, autonomous decision-making, precise execution, and dynamic optimization is formed. Within this system, the service layer provides intelligent service support, while the data layer integrates multi-source heterogeneous data, providing a data foundation for mission analysis, situational awareness, and environmental understanding. The airborne system is equipped with a cluster of intelligent payloads, including lidar and multispectral imagers, constructing an integrated air-space-ground sensing matrix that provides comprehensive environmental perception capabilities for reconnaissance missions.
[0058] Regarding key technology modules, this invention proposes several innovations. First, a language model-based intelligent engine enables task decomposition, modeling, planning, calibration, and combination, and facilitates flexible task assignment, communication, and error correction through natural language interaction. This module not only improves task execution efficiency but also enhances the system's intelligence level. Second, a mobile navigation and positioning enhancement system provides precise navigation support for the UAV, ensuring accurate flight and task execution in complex environments. Furthermore, a language model-driven vehicle-to-machine (V2X) collaborative planning system is another core innovation of this invention. This system achieves efficient collaboration between vehicles and UAVs through two stages: task planning and dynamic collaboration, significantly improving overall operational efficiency.
[0059] In terms of specific functional implementation, this invention supports multimodal input fusion and parsing, capable of handling various input methods such as natural language, speech, and map annotations, and accurately parses user commands through NLP and image processing technologies. Simultaneously, the application of a spatiotemporal collaborative arc routing optimization algorithm makes the collaborative path planning between vehicles and drones more rational and efficient. Regarding air-ground collaborative control and obstacle avoidance, this invention combines environmental modeling and dynamic obstacle recognition technologies to achieve safe obstacle avoidance for both drones and ground vehicles, ensuring the safety of the operation process. Furthermore, the application of a three-dimensional obstacle avoidance algorithm enables the system to perform path planning and obstacle avoidance in three-dimensional space, overcoming the limitations of traditional two-dimensional obstacle avoidance algorithms.
[0060] The present invention offers significant advantages. First, by applying multimodal large language models and digital twin technology, the system can analyze complex environments in real time and dynamically adjust task strategies, thereby significantly improving environmental adaptability. Second, the application of real-time data processing and intelligent decision-making mechanisms significantly improves task response speed and operational efficiency. Furthermore, the application of vehicle-machine collaborative planning and dynamic obstacle avoidance technology significantly enhances the collaboration efficiency between the UAV and ground vehicles, improving overall operational effectiveness. Finally, the application of the autonomous intelligent reconnaissance system reduces reliance on manual labor, lowers field safety risks, and reduces operating costs, bringing significant economic benefits to enterprises.
[0061] This invention proposes a detailed system deployment and mission execution process. First, an onboard edge computing module and automated hive are deployed on the unmanned vehicle (UAV), which also carries multiple UAVs, forming an air-ground collaborative reconnaissance system. The UAVs are equipped with monocular cameras and onboard computing units to acquire attitude streams and image data in real time, providing accurate information support for the reconnaissance mission. Regarding the mission execution process, after receiving the mission instruction, the backend command center sends the instruction to the UAV to initiate the reconnaissance mission. After self-checking, the UAV analyzes the mission according to the instruction, generating a mission plan and route planning. Subsequently, the UAV navigates to the designated area and deploys the UAVs for reconnaissance. The UAVs acquire battlefield situational information and track key targets. The ground-based system rapidly reconstructs two-dimensional and three-dimensional models of the battlefield environment, identifying and analyzing key targets. Finally, the data is compressed and transmitted back to the backend command center via satellite communication for analysis and decision-making by command personnel.
[0062] Regarding the technical implementation details, this invention elaborates on each key technical module. The language model-based intelligent engine is implemented through an intelligent task self-generation module, a general model library construction module, a LangChain-based intelligent task processing module, and an unmanned equipment interaction module. These modules work together to ensure accurate task parsing and execution. The mobile navigation and positioning enhancement system utilizes PPPRTK technology to provide high-precision positioning support, ensuring accurate navigation for UAVs and ground vehicles in complex environments. The vehicle-machine collaborative planning system employs a large model-driven task parsing and dynamic interaction method to achieve task planning and dynamic collaboration, improving the efficiency of cooperation between vehicles and UAVs. Scene reconstruction and intelligence generation combine SLAM algorithms and Gaussian real-time rendering algorithms to achieve high-precision 3D reconstruction and real-time intelligence generation, providing accurate environmental information and intelligence support for command and decision-making.
[0063] To ensure system stability and reliability, this invention also proposes a system testing and optimization scheme. System testing is conducted under various complex scenarios to verify the system's environmental adaptability, task response speed, and multi-device collaboration efficiency. Problems are identified through testing, and targeted improvements and optimizations are implemented. Based on the test results, the system is optimized to improve performance and stability, ensuring stable operation in various complex scenarios and meeting user needs.
[0064] Example 1
[0065] Example Background
[0066] In natural resource surveys and monitoring in complex and challenging areas, traditional manual methods face numerous challenges, such as high operating costs, significant safety risks, and limited data acquisition. To address these issues and improve the environmental adaptability, mission response speed, and multi-device collaboration efficiency of reconnaissance systems, this invention proposes an autonomous intelligent locomotive-assisted unmanned reconnaissance device based on a language model-driven task parsing and dynamic interaction method. This device enables real-time dynamic collaboration between hardware clusters and software decision-making in complex scenarios, providing an efficient and safe solution for tasks such as natural resource surveys and monitoring.
[0067] Implementation Example Description
[0068] The autonomous intelligent vehicle-assisted unmanned reconnaissance equipment in this embodiment mainly consists of two parts: an unmanned vehicle and unmanned aerial vehicles (UAVs). The unmanned vehicle is equipped with an onboard edge computing module and an automated drone pod, capable of carrying multiple UAVs simultaneously. The UAVs are equipped with a monocular camera and an onboard computing unit, enabling them to acquire attitude stream and image data in real time, providing accurate information support for reconnaissance missions.
[0069] In terms of system architecture, this embodiment constructs a collaborative service system that deeply integrates artificial intelligence and spatial intelligence. This system includes a service layer, a data layer, a network layer, and an airborne system. The service layer is the core of the entire system, integrating a language model-driven intelligent engine, a mobile navigation enhancement system, a vehicle-machine collaborative planning system, a PPPRTK-assisted scene reconstruction system, and a scene understanding and intelligence generation system. These systems work together to achieve a complete process from task analysis to intelligence generation. The data layer integrates multi-source heterogeneous data, including sample library data, AI data, and remote sensing data, providing a solid data foundation for task analysis, situational awareness, and environmental understanding. The network layer integrates advanced technologies such as 4G / 5G networks and satellite communication, providing high-bandwidth, low-latency data transmission capabilities to ensure real-time communication between various parts of the system. The airborne system uses a multi-rotor / fixed-wing UAV collaborative platform as its carrier, configured with an intelligent payload cluster including lidar and multispectral imagers, providing comprehensive environmental awareness capabilities for reconnaissance missions.
[0070] In terms of implementation process, this embodiment achieves a complete automated process from user input to intelligence generation. Users can issue task commands through various methods such as natural language, voice, and map annotation. After receiving the command, the system uses a large language model to perform semantic parsing and extract key information such as task objectives, priorities, and resource requirements. Subsequently, the system generates specific sub-tasks and performs task allocation and scheduling. During task execution, the system considers constraints such as the vehicle and drone's endurance, payload limitations, and sensor types to ensure smooth task execution. Simultaneously, the system also possesses dynamic collaborative capabilities, enabling efficient cooperation between vehicles and drones through technologies such as air-to-ground linkage, dynamic obstacle avoidance and path correction, and task reallocation. Finally, based on the scene understanding and intelligence generation system, the system generates a structured intelligence map, providing strong support for decision-making.
[0071] Example Effects
[0072] The autonomous intelligent vehicle-to-machine (V2M) unmanned reconnaissance equipment in this embodiment has significant advantages over existing technologies. First, by applying multimodal large language models and digital twin technology, the system can analyze complex environments in real time and dynamically adjust task strategies, thus significantly improving environmental adaptability. This enables the system to operate stably in various complex scenarios, meeting user needs. Second, the application of real-time data processing and intelligent decision-making mechanisms significantly improves task response speed. The system can quickly parse user commands, generate task planning schemes, and rapidly execute reconnaissance tasks, improving operational efficiency. Furthermore, through the application of vehicle-to-machine collaborative planning and dynamic obstacle avoidance technology, the collaboration efficiency between the UAV and ground vehicles is significantly enhanced. The system can achieve dynamic obstacle avoidance and path correction in air-to-ground coordination, ensuring the safe and efficient execution of reconnaissance missions. Finally, the application of the autonomous intelligent reconnaissance system reduces reliance on manual labor, lowers field safety risks, and reduces operating costs, bringing significant economic benefits to enterprises.
[0073] Example 2
[0074] Example Background
[0075] In complex scenarios such as military reconnaissance and disaster relief, there is an urgent need for efficient and safe reconnaissance equipment. Traditional reconnaissance systems are often limited by fixed mission procedures and limited environmental awareness capabilities, making it difficult to adapt to complex and ever-changing mission scenarios. To improve the intelligence level and combat effectiveness of reconnaissance systems, this invention proposes an autonomous intelligent vehicle-assisted unmanned reconnaissance device based on a language model-driven mission parsing and dynamic interaction method. This device can achieve real-time dynamic coordination between hardware clusters and software decision-making in complex scenarios, providing strong support for military reconnaissance, disaster relief, and other missions.
[0076] Implementation Example Description
[0077] The autonomous intelligent vehicle-to-machine cooperative unmanned reconnaissance device in this embodiment also includes two main components: an unmanned vehicle and an unmanned aerial vehicle, and adopts a system architecture similar to that of Embodiment 1. However, in terms of technical implementation details, this embodiment places greater emphasis on language model-driven task parsing and dynamic interaction capabilities, as well as the optimization of vehicle-machine cooperative planning and dynamic obstacle avoidance technologies.
[0078] In terms of language model-driven task parsing, this embodiment supports multimodal input fusion and parsing, capable of handling various input methods such as natural language, speech, and map annotation. Through advanced semantic parsing technology, the system can accurately extract key information from user instructions, such as task objectives, priorities, and resource requirements. Subsequently, the system generates specific sub-tasks and performs task allocation and scheduling. To ensure the accuracy of task execution, this embodiment also designs a conflict detection mechanism based on semantic similarity matching, which can correct instruction conflicts in real time and avoid errors during task execution.
[0079] In terms of dynamic interaction, this embodiment employs an advanced vehicle-to-machine (V2X) cooperative arc routing optimization algorithm. This algorithm combines a simulated annealing mechanism with a stochastic variable neighborhood descent search algorithm (RVND-SA) to optimize task paths and improve task execution efficiency. Simultaneously, this embodiment also integrates a three-dimensional tangent method and a two-dimensional obstacle avoidance algorithm, utilizing reinforcement learning techniques to enhance the adaptability and optimization capabilities of the obstacle avoidance strategy. The application of these technologies enables the system to achieve dynamic obstacle avoidance and path correction in complex scenarios, ensuring the safe and efficient execution of reconnaissance missions.
[0080] Example Effects
[0081] The autonomous intelligent vehicle-to-machine (V2M) unmanned reconnaissance equipment in this embodiment has demonstrated significant advantages in military reconnaissance, disaster relief, and other missions. First, through the enhancement of language model-driven task parsing and dynamic interaction capabilities, the system can more accurately understand user commands and generate more reasonable mission planning schemes. This enables the system to better adapt to complex and ever-changing mission scenarios, improving combat effectiveness. Second, the application of the vehicle-to-machine collaborative arc routing optimization algorithm optimizes mission paths, reduces unnecessary movement and waiting time, and improves mission execution efficiency. Furthermore, the application of air-to-ground coordinated dynamic obstacle avoidance and path correction technology allows the system to flexibly respond to various emergencies in complex scenarios, ensuring the safe and successful completion of reconnaissance missions. Finally, through the application of PPPRTK-assisted scene reconstruction and scene understanding and intelligence generation systems, the system can provide more accurate environmental information and intelligence support, providing strong support for command and decision-making, and improving the accuracy and timeliness of combat command.
[0082] In summary, this invention, by constructing a fully autonomous collaborative reconnaissance system with a multimodal large language model as its core engine, achieves real-time dynamic collaboration between hardware clusters and software decision-making in complex scenarios. This system not only improves the environmental adaptability, task response speed, and multi-device collaboration efficiency of the reconnaissance system, but also reduces manual operation costs and risks, providing strong support for enterprise technology transformation and product structure adjustment. Through a detailed description of system deployment, task execution processes, technical implementation details, and system testing and optimization schemes, this invention provides a comprehensive technical solution for the collaborative operation of autonomous intelligent vehicles and unmanned reconnaissance equipment.
[0083] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the scope of the invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0084] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A language model-driven task parsing and dynamic interaction method, characterized in that, Includes the following steps: Instruction fusion and conflict resolution steps: A large-model-driven intelligent engine is used to parse and extract key information from task instructions given by users through various input methods; Design a conflict detection mechanism based on semantic similarity matching to identify and resolve potential conflicts in instructions, including conflicting task priorities or resource competition. When command conflicts are detected, they are corrected in real time through dialogue interaction to ensure the reasonable allocation and execution of tasks; Dynamic decision-making and task reallocation steps: Collect and analyze multimodal data in real time, and extract key data for risk prediction; Generate task replanning suggestions. When the environment changes or unexpected situations occur during task execution, the feasibility of the task is reassessed by combining the task's historical records, current resource status, and unexpected events. Based on task priority, resource availability, and task progress, the task allocation scheme is dynamically adjusted to ensure the smooth execution and efficient completion of tasks.
2. The method according to claim 1, characterized in that, It also includes vehicle-machine collaborative arc routing optimization steps: The design task coding method is to represent the combination of flight paths of the UAV during the execution of the task by defining an appropriate coding structure; A stochastic variable neighborhood descent search algorithm combined with simulated annealing mechanism is adopted for path optimization, and the task path is continuously adjusted and optimized through multiple neighborhood search operators; By leveraging the global optimization capabilities of simulated annealing, local optima are avoided, ensuring that the optimal task execution path is found in large-scale patrol missions.
3. The method according to claim 1, characterized in that, It also includes dynamic obstacle avoidance and path correction steps involving air and ground coordination: Automatically select the appropriate path planning method based on the availability of the data source; When a 3D data source is available, the 3D tangent method is used for path planning, taking into account 3D spatial factors such as the height difference of obstacles, terrain undulations and buildings. When the 3D data source is insufficient or unavailable, a 2D obstacle avoidance algorithm, such as the static elliptical tangent method and the dynamic elliptical tangent method, is used to plan the path based on the 2D position data.
4. The method according to claim 1, characterized in that, It also includes PPPRTK-assisted scene reconstruction steps: Build a collaborative measurement platform integrating multiple sensors, including lidar, IMU, optical camera, and GNSS module; A fast, tightly coupled laser inertial odometry method is adopted to improve the robustness and accuracy of real-time positioning and mapping in high dynamic environments; The Gaussian real-time rendering algorithm is used to achieve efficient visualization in complex dynamic environments.
5. The method according to claim 1, characterized in that, It also includes scene understanding and intelligence generation steps: Receive fused multimodal data, including visual features after image encoding and linguistic features after text encoding, and perform preliminary analysis through the basic inference stage; The data enters the core processing stage, including spatial perception, language understanding, scene reasoning and data reasoning, to optimize the rich features of multimodal data; During the task decoding phase, detailed task instructions or intelligence outputs are generated through task cleaning, online target path planning, dynamic adjustment, and visualization.
6. The method according to any one of claims 1 to 5, characterized in that, The large model-driven intelligent engine further includes: The intelligent task self-generation module enables task decomposition, task modeling, task planning, task calibration, and task combination. A general model library building module that enables model building, model description, model matching, and model testing; Based on LangChain's intelligent task processing module, it enables task scheme management, task prompt construction, and intelligent task execution.
7. The method according to claim 6, characterized in that, It also includes an unmanned equipment interaction module, which further includes: Dynamic behavior planning functionality includes control command parsing, action space, and path planning; Real-time data transmission capabilities, including data acquisition, data storage, and management; The group collaborative scheduling function enables autonomous collaboration and global optimization of the drone swarm.