Safe flight agent system for executing complex language instructions in high dynamic environment

By designing high and low-level structure perception and planning modules in the embodied intelligent aircraft system, combining the large language model and edge computing platform in the cloud, the problem of insufficient ability to execute complex commands and lack of security guarantees in high-dynamic environments is solved, and efficient and safe task execution is achieved.

CN119992387AActive Publication Date: 2025-05-13SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510098835.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing embodied intelligent aircraft lack the ability to execute complex language instructions in highly dynamic environments and lack considerations for security in interactive environments.

Method used

A safe flight agent system that executes complex language instructions in a highly dynamic environment is designed, using perception function modules and planning function modules with high and low-level structures, combined with a large language model and an edge computing platform in the cloud to realize real-time task management and security assessment.

Benefits of technology

Improve the mission execution efficiency and safety of flying agents in a highly dynamic environment, ensure the safety of people and property in the environment, and avoid potential privacy and collision risks through real-time security assessment and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992387A_ABST
    Figure CN119992387A_ABST
Patent Text Reader

Abstract

The invention provides a safe flight agent system for executing complex language instructions in a high dynamic environment. The system comprises a perception module and a security module based on a visual language model, a planning module based on a language large model and an unmanned aerial vehicle control interface. In a system operation stage, an intelligent agent receives a command from a user, a planning module splits a task, calls a corresponding task memory and a code generator to generate a corresponding executable code, calls a perceived interface to obtain perceptual information related to the task, and finally, after the perceptual information is checked by a security check module, the perceptual information is sent to a server. And executing the corresponding code to complete the instruction of the user. The system achieves safe and efficient execution of tasks in a highly dynamic environment by combining a high-performance low-frame-rate large model and a high-frame-rate traditional method in each module and using a security detection module involving privacy, potential risk and collision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of embodied intelligent agents and aircraft systems, and in particular to a safe flying intelligent agent system that executes complex language instructions in a highly dynamic environment. Background Art

[0002] Embodied intelligence refers to the intelligent behavior exhibited by intelligent entities through their physical forms (such as robots, virtual images, etc.) interacting with the environment. This intelligence not only comes from computing and algorithms, but also covers multi-dimensional capabilities such as perception, movement, and environmental interaction. Embodied intelligence places special emphasis on the following core elements: interaction between the body and the environment. Intelligent entities interact and feedback with the surrounding environment in real time through their perception and movement systems. For example, robots use cameras, sensors and other devices to perceive the environment and perform actions through motors, joints, etc. The combination of perception and action enables embodied intelligent entities to convert sensory information (including vision, hearing, touch, etc.) into specific actions, thereby completing complex tasks. This fusion of perception and action is the key foundation for realizing intelligent behavior. Adaptability and learning ability. Embodied intelligent entities have the ability to adapt to environmental changes and can continuously optimize their behavior and decision-making through learning and experience. For example, robots are able to improve their path planning and task execution through repeated trials.

[0003] Embodied intelligent entities operate in the physical world, so they must consider physical constraints and resource limitations, such as energy consumption, computing power, mechanical structure, etc. The research goal of embodied intelligence is to develop intelligent entities that can perform tasks autonomously, flexibly, and efficiently in a physical environment, which has broad application prospects in the fields of robotics, autonomous driving, virtual reality, and augmented reality.

[0004] At present, some methods and systems have been proposed for embodied intelligence based on aircraft. However, due to the real-time defects of large models, the ability of intelligent agents to complete complex user instructions in highly dynamic environments is still very lacking. In addition, current embodied intelligence systems rarely consider safety assurance in interactive environments.

[0005] At present, the long inference time of large models and the limitation of communication rate have led to the fact that the ability of embodied intelligence in high-dynamic scenarios is greatly limited. Due to the huge number of parameters of high-performance models, the inference time of the model is difficult to be effectively controlled under the existing hardware platform. For aircraft, it is difficult to deploy models with huge parameters on lightweight edge computing platforms, and they are generally deployed in the cloud remotely and communicate with aircraft. However, in perception tasks, high-definition images need to be transmitted to the model for calculation. The inference time of the model itself plus the communication delay will further deepen the system delay. High latency will cause the ability of the intelligent agent to perform tasks to drop sharply. How to ensure the performance of aircraft in high-dynamic environments is a key and poorly solved problem. At the same time, as a highly mobile intelligent agent, how to ensure the privacy and physical safety of objects in the environment when performing tasks is an important and less deeply considered issue. Summary of the invention

[0006] In view of the defects in the prior art, an object of the present invention is to provide a safe flight intelligent agent system that executes complex language instructions in a high dynamic environment.

[0007] According to one aspect of the present invention, a safe flight intelligent agent system for executing complex language instructions in a high dynamic environment is provided, comprising:

[0008] A perception function module, wherein the perception function module adopts a high-low structure, wherein the high-level structure is a detection visual language model installed on the cloud platform, and the low-level structure is a perception module installed on the flying agent; the detection visual language model combines information of the perception module and observation signals collected by the flying agent to perform open set detection; the perception module performs closed set detection based on observation signals and information of the detection visual language model;

[0009] The task manager is installed on the cloud platform and includes a task decomposer, a state manager and a memory module. The task decomposer receives user instructions, performs semantic understanding and task splitting, the state manager manages task status, and the memory module stores process variables, task completion status, and interaction information with users;

[0010] A planning function module adopts a high-low structure, the high-level structure is a large language model planner installed on the cloud platform, and the low-level structure is a planning module installed on the flying intelligent body; the large language model planner generates an execution code based on the task decomposition result of the task manager, the information of the detected visual language model, and the information of the planning module; the planning module obtains a planning path based on the task decomposition result of the task manager, the information of the perception module and the execution code of the large language model planner.

[0011] Preferably, the task decomposer receives a command from a user and calls the task memory in the memory module to split the task, specifically:

[0012] T t =f decomposition (I t ,M)={i 1,t ,i 2,t ,.,i n,t}

[0013] Among them I t represents the user command received by the agent at time t and serves as the command input of the task decomposer. M is the memory generated by the agent during operation, including historical perception results and historical task completion status. decomposition Represents the task decomposer model function of the agent; i n,t Indicates the executable subtasks decomposed from the user's instructions, including: search, follow, navigate, patrol, and direct movement of the drone; T t Represents the set of splits of the agent's instructions to the user at time t.

[0014] Preferably, the large language model planner generates an execution code based on the task decomposition result of the task manager, the information of the detected visual language model, and the information of the planning module, specifically:

[0015] C n,t =f planner (i n,t ,P,L,M)

[0016] S n,t =f safe_execute (C n,t )

[0017] Among them, i n,t It is expressed as the agent's nth executable subtask at time t, f planner It is a code generator for executable subtasks based on the large language model; P is the prompt corresponding to the large language model code generator for executable tasks, L is a set of control and perception interfaces that can be called in the code generator, including the visual language model interface and the related interfaces of the UAV control. The input of the visual language model interface is the information of detecting the visual language model and the target of the planning module, and the output is the target position provided to the large language model planner; the input of the related interface of the UAV control is the target point of different tasks of the planning module or the ID of the object to be tracked, and the output is the execution code of the intelligent aircraft; C n,t The executable Python code generated for the nth subtask at time t;

[0018] Among them, fsafe_execute It is a safe execution module for Python code. This module executes the subtask code and returns a Boolean value S n,t When it is True, it means that the code has no syntax problems and can be executed. Otherwise, it means that the generated code has syntax problems and needs to be rewritten.

[0019] Preferably, the planning module includes:

[0020] The obstacle avoidance planner has the mobile target point and the information of the perception module as input and the obstacle avoidance path planned in real time as output;

[0021] The tracking planner takes as input the ID of the object to be tracked and the information of the perception module, and outputs a path that avoids obstacles in real time and can follow the target object.

[0022] Preferably, it also includes a safety function module, which adopts a high-low structure, the high-level structure is a safety visual language model installed on the cloud platform, and the low-level structure is a collision detector installed on the flying intelligent body; the safety visual language model performs privacy safety detection and potential safety detection based on the observation signal collected by the flying intelligent body; the collision detector performs collision detection based on the observation signal collected by the flying intelligent body and the planned path of the planning function module.

[0023] Preferably, the privacy security detection is specifically:

[0024] pl t ,pr t =f privacy_check (O t )

[0025] Among them, f privacy_check It is a privacy detection module based on a fine-tuned visual language model.

[0026] O t is the observation signal collected by the flying agent at time t and the state of the drone, pr t and rl t ,is the privacy level and the reason for the privacy score;

[0027] The potential safety detection is specifically:

[0028] rl t ,rr t= f risk_check (O t )

[0029] Among them, f risk_check It is a potential risk detection module, rl t ,rr tThey are the potential risk score and the reason for the potential risk evaluation.

[0030] Preferably, the collision detection is specifically:

[0031] Co t =f collision_check (O t , p t )

[0032] f collision_check It is a collision detection module, whose input is the observation signal O at time t t And the path planning information p at time t t The collision detection function predicts the position of dynamic objects in the environment through the Kalman filtering method and performs collision detection with the planned trajectory.

[0033] Preferably, the safety function module performs safety assessment and safety processing according to the privacy detection results, potential safety detection results and collision detection results.

[0034] Preferably, the safety assessment is specifically:

[0035] pl t and rl t is an integer from 1 to 4, pr t With rr t The reason for the score, in the form of a string;

[0036] 1 means no security threat, 2 means slight security threat but under control, 3 means moderate security threat and the user needs to be informed to terminate the task, and 4 means severe security threat.

[0037] Co t It is the Boolean value of the collision detection result at time t; when this variable is True, it means that the drone is based on p at this moment. t and the observed signal O t A collision risk is detected; otherwise, there is no collision risk.

[0038] Preferably, the security processing includes:

[0039] When the privacy and potential risk evaluation is 3, the security function module directly sends a warning message to the task manager, which sends it to the user and terminates the task;

[0040] When the risk assessment is 4, the safety processing module controls the flight agent to terminate the mission and return to the last safe point;

[0041] For collision detection, when the safety function module detects a collision, the collision detector directly notifies the flight agent to immediately perform emergency braking to avoid the collision, and at the same time communicates with the task manager to inform the user.

[0042] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0043] The safe flight intelligent agent system that executes complex language instructions in a high-dynamic environment in an embodiment of the present invention, combined with the high-performance, low-frame-rate large prediction model in each module and the high-frame-rate perception, planning and other modules, can ensure efficient execution of tasks in a high-dynamic environment.

[0044] In the safe flight intelligent agent system that executes complex language instructions in a high dynamic environment in an embodiment of the present invention, each functional module is designed with a high-low layer structure to address the problem of high latency. The high layer ensures the performance of the intelligent agent, and the low layer ensures the rapid response of the intelligent agent to the environment.

[0045] The safe flight intelligent agent system that executes complex language instructions in a high dynamic environment in an embodiment of the present invention integrates a large model and a memory module in the task manager and planning function module respectively to achieve better environmental understanding and task planning performance.

[0046] The safe flight intelligent agent system for executing complex language instructions in a high-dynamic environment in an embodiment of the present invention is designed to address the safety issues of aircraft. A safety function module that integrates a large visual language model is designed to evaluate the current safety of the drone, thereby ensuring the safety of people and property in the environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0048] Figure 1 An intelligent agent framework based on a language model and a visual language model according to an embodiment of the present invention;

[0049] Figure 2 A planning module framework based on a large language model according to an embodiment of the present invention;

[0050] Figure 3 This is a security function module framework based on a visual language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0052] In one embodiment of the present invention, a safe flight intelligent agent system for executing complex language instructions in a high-dynamic environment is provided. Figure 1 As shown, it mainly includes:

[0053] The perception function module adopts a high-low structure. The high-level structure is the detection visual language model installed on the cloud platform, and the low-level structure is the perception module installed on the flying intelligent body. The detection visual language model combines the information of the perception module and the observation signal collected by the flying intelligent body to perform open set detection. The perception module performs closed set detection based on the observation signal and the information of the detection visual language model.

[0054] The task manager is installed on the cloud platform and includes a task decomposer, a state manager, and modules. The task decomposer receives user instructions, performs semantic understanding and task splitting, and the state manager manages task status. The memory module stores process variables, task completion status, and interaction information with users.

[0055] Planning function module, the planning function module adopts a high-low structure. The high-level structure is the large language model planner installed on the cloud platform, and the low-level structure is the planning module installed on the flying intelligent body; the large language model planner generates execution code based on the task splitting results of the task manager, the information of the detection visual language model, and the information of the planning module; the planning module obtains the planning path based on the task decomposition results of the task manager, the information of the perception module and the execution code of the large language model planner.

[0056] The above embodiments, combined with the high-performance, low-frame-rate large prediction model in each module and the high-frame-rate perception, planning and other modules, can efficiently execute in highly dynamic tasks.

[0057] The observation signal is acquired by a camera device installed in the flying intelligent body. In order to further perceive the information content of the observation signal, in a preferred embodiment of the present invention, a better perception function module is provided, and its detection visual language model and perception module are in a high-low hierarchical relationship. The detection visual language model provides high-level open set detection, and the perception module provides low-level closed set detection of the 3D position information of the object. They interact to obtain real-time open set detection results. The input of the perception module is the image (observation signal) and the position information of the drone, and the output is the closed set 3D position information and category of pedestrians, bicycles, motorcycles, and vehicles.

[0058] The above embodiment integrates a large model and an end-to-end 3D position detection model, which can better understand environmental information and obtain accurate detection target locations and categories.

[0059] Through the perception function module in the above implementation, the current environment information of the flying intelligent body and the detected environment information can be obtained. At this time, the user issues a language-based fuzzy command to the intelligent body. In order to obtain a correct understanding and analysis of the fuzzy command, in a preferred implementation of the present invention, the task instruction is decomposed, managed and stored by a task manager. Specifically, the task is decomposed using a task decomposer based on a large language model, and the decomposed subtasks include patrolling, searching, navigation, following and language control of the drone movement.

[0060] In some specific embodiments, the flight agent task decomposition process can be expressed as:

[0061] T t =f decomposition (I t ,M)={i 1,t ,i 2,t ,.,i n,t}

[0062] Among them I t represents the user instruction received by the agent at time t and serves as the instruction input of the task decomposer. M is the task memory generated when the agent performs the task. decomposition The model function representing the task decomposition in the agent's task decomposer; i n,t represents the nth decomposition task of the user instruction at time t, and these subtasks belong to one of the executable subtasks mentioned above; T t Represents the set of splits of the agent's instructions to the user at time t.

[0063] At the same time, the state manager can switch the current operating state of the intelligent agent based on the data information fed back by each module, thereby maintaining the stable operation of the intelligent agent.

[0064] Note that at each stage, process variables, task completion status, and interaction information with the user will be stored in the memory module. The memory information will be provided to the large model in different modules, so that the agent can understand its own state and use information from past experience to plan, thereby improving the overall performance of the agent.

[0065] After the task decomposition is completed through the task manager, it is necessary to obtain the executable code of each subtask. Therefore, in a preferred embodiment, the executable code is generated by the large language model planner. In this embodiment, the task planner uses the large language model to coordinate the perception of the perception module and the perception information interface based on the VLM visual language model and the special code generator of each subtask to perform task planning by generating executable python code. The process can be expressed as:

[0066] c n,t =f task_planner (i n,t ,P,L,M)

[0067] Among them, i n,t It is expressed as the nth executable subtask command of the agent at time t. task_planner It is a task planner based on the large language model. p is the prompt corresponding to the large language model executable code generator. L is a set of control and perception interfaces that can be called in the code generator, which includes the visual language model interface and the related interfaces of the UAV control. The input of the visual language model interface is the information of the visual language model detection and the target of the planning module, and the output is the target position provided to the large language model planner; the input of the related interface of the UAV control is the target point of different tasks of the planning module or the ID of the object to be tracked, and the output is the execution code of the intelligent aircraft; C n,t The executable Python code generated for the nth subtask at time t.

[0068] During the code generation process, the code generator for each subtask based on the large language model will improve the code generation related to the executable subtask in the task planner and finally output the complete code. This process can be expressed as:

[0069] C n,t =f code_generator_j (c n,t )

[0070] f code_generator_j is a code generator based on a large language model for the jth subtask, j∈{search, follow, navigate, patrol, direct movement of drones}. The code generator generates executable complete code C according to commands in different categories. n,t .

[0071] In order to ensure that the code is correct, in a preferred embodiment of the present invention, the syntax of the generated complete code is checked, specifically:

[0072] S n,t =f safe_execute (C n,t )

[0073] It is a safe execution module for Python code. This module will execute the subtask code and return a Boolean value S n,t When it is True, it means that the code has no syntax problems and can be executed. Otherwise, it means that the generated code has syntax problems and needs to be rewritten. When the code is checked and there is no syntax error and can be executed, the code is executed.

[0074] Furthermore, in a preferred embodiment of the present invention, after the code is executed, a task completion detection module is used to detect whether the task is completed. The process is shown as follows.

[0075] F n,t ,E n,t =f exec_check (C n,t ,O t )

[0076] Among them, f exec_check It is the subtask status confirmation module. t is the real-time state of the drone and the observation signal of the sensor, F n,t A Boolean value indicating whether the task execution failed. When this variable is True, it proves that the task execution failed, resulting in the failure of the entire task. Otherwise, it proves that the subtask execution was successful, and the next task will be executed. n,t A Boolean variable that determines whether a subtask needs to be interacted with by the user again to determine the user's intention. When this variable is True, the system will interact with the user again and provide the current task execution status. After the user confirms the requirement, the subtask will be re-planned according to the user's needs. Figure 2 As shown, the process will maintain interaction with the user until the task is completed and the next command is issued.

[0077] In another preferred embodiment, the planning module is composed of an obstacle avoidance planner and a tracking planner. The input of the obstacle avoidance planner is the information of the moving target point and the perception module, and the output is the obstacle avoidance path planned in real time. The input of the tracking planner is the ID of the object to be tracked and the information of the perception module, and the output is a path that avoids obstacles in real time and can follow the target object.

[0078] In the above embodiment, the cloud-based large language model planner calls the obstacle avoidance planner and the tracking planner according to the task. The cloud-based large language model planner sends the obtained code to the agent, and the agent calls each module for execution.

[0079] In order to ensure the safety performance of the entire system, in a preferred embodiment of the present invention, a safety function module is also designed in the entire system. The safety function module is a relatively independent module that monitors the privacy security, potential risk security and planned collision security of the intelligent agent in the environment in real time. Figure 3 As shown, specifically:

[0080] pl t ,pr t =f privacy_check (O t )

[0081] rl t ,rr t =f risk_check (O t )

[0082] Co t =f collision_check (O t , p t )

[0083] where f privacy_check and f risk_check They are the privacy detection module and the potential risk detection module in the detection module. Both modules are based on the fine-tuned visual language model. t is the sensor observation signal of the drone at time t, as well as the state of the drone. t ,pr t and rl t ,rr t They are privacy level, privacy score reason, potential risk score, and potential risk evaluation reason. t and rl t An integer from 1 to 4. t With rr t The reason for the score, in the form of a string. collision_check It is based on the traditional collision detection method, the input of which is the observation signal O at time t. t And the path planning information p at time t t , the Kalman filter method is used to predict the position of dynamic objects in the environment and perform collision detection with the planned trajectory. If a collision is detected, an early warning is given.

[0084] In this step, the drone's perception information is input into the privacy security and potential security detection module in real time. The detector can analyze the scene for privacy and potential security issues based on different rules based on the perception information, and give reasons and scores. At the same time, if the drone is flying and performing tasks at this time point, the collision detector will perform real-time collision detection on the flight path and dynamic objects in the environment based on the current perception information to ensure that the drone does not collide with objects in the environment.

[0085] Furthermore, in a preferred implementation, the safety evaluation module will evaluate the current safety situation of the drone and summarize the safety situation and send it to the safety processing module. For the privacy and potential risk modules, they will output a hazard level from 1 to 4. 1 means no safety threat, 2 means slight safety threat but within the controllable range, 3 means medium safety threat and the user needs to be informed to terminate the mission, and 4 means serious safety threat and the drone will forcibly terminate the mission. If the current situation is safe, it will not affect the operation of the drone. For collision detection, if detected, the drone's mission will be stopped directly.

[0086] Furthermore, in another preferred implementation, the safety processing module controls the drone to take actions to maintain a safe state according to the results of the safety evaluation for different safety issues. When the privacy and potential risk evaluation is 3, a warning message will be sent to the user, and the user will be asked to terminate the task. When the risk evaluation is 4, the drone will terminate the task and return to the last safe point. As for collision detection, when the system detects a collision, the agent immediately performs emergency braking to avoid the collision and informs the user.

[0087] In the above embodiment, during the system operation phase, the agent receives commands from the user, the planning module splits the task and calls the corresponding task memory and code generator to generate the corresponding executable code, calls the perception interface to obtain the task-related perception information, and finally executes the corresponding code to complete the user's instructions after checking by the security check module. The system achieves safe and efficient execution of tasks in a highly dynamic environment by combining a large model with high performance and low frame rate in each module, a traditional method with high frame rate, and a security detection module involving privacy, potential risks and collisions.

[0088] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A safe flight agent system that executes complex language instructions in a highly dynamic environment, characterized in that: include: A perception function module, wherein the perception function module adopts a high-low structure, wherein the high-level structure is a detection visual language model installed on the cloud platform, and the low-level structure is a perception module installed on the flying agent; the detection visual language model combines information of the perception module and observation signals collected by the flying agent to perform open set detection; the perception module performs closed set detection based on the observation signals and information of the detection visual language model; The task manager is installed on the cloud platform and includes a task decomposer, a state manager and a memory module. The task decomposer receives user instructions, performs semantic understanding and task splitting; the state manager manages task status; and the memory module stores process variables, task completion status and interaction information with users. A planning function module, wherein the planning function module adopts a high-low structure, wherein the high-level structure is a large language model planner installed on the cloud platform, and the low-level structure is a planning module installed on the flying intelligent body; The large language model planner generates an execution code based on the task splitting result of the task manager, the information of the detected visual language model, and the information of the planning module; The planning module obtains a planning path based on the task splitting result of the task manager, the information of the perception module and the execution code of the large language model planner.

2. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 1, characterized in that: The task decomposer receives a command from a user and calls the task memory in the memory module to split the task, specifically: T t =f decomposition (I t ,M)={i 1,t ,i 2,t ,.,i n,t } Among them I t represents the user instruction received by the agent at time t and serves as the instruction input of the task decomposer. M is the task memory generated by the agent during operation, which includes historical perception results and historical task completion status. decomposition Represents the task decomposer model function of the agent; i n,t Indicates the executable subtasks decomposed from the user's instructions, including: search, follow, navigate, patrol, and direct movement of the drone; T t Represents the set of splits of the agent's instructions to the user at time t.

3. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 2, characterized in that: The large language model planner generates an execution code based on the task decomposition result of the task manager, the information of the detected visual language model, and the information of the planning module, specifically: C n,t =f planner (i n,t ,P,L,M) S n,t =f safe_execute (C n,t ) Among them, i n,t It is expressed as the agent's nth executable subtask at time t, f planner It is a code generator for executable subtasks based on the large language model; P is the prompt corresponding to the large language model code generator for executable tasks, L is a set of control and perception interfaces that can be called in the code generator, including the visual language model interface and the related interfaces of the UAV control. The input of the visual language model interface is the information of detecting the visual language model and the target of the planning module, and the output is the target position provided to the large language model planner; the input of the related interface of the UAV control is the target point of different tasks of the planning module or the ID of the object to be tracked, and the output is the execution code of the intelligent aircraft; C n,t The executable Python code generated for the nth subtask at time t; Among them, f safe_execute It is a safe execution module for Python code. This module executes the subtask code and returns a Boolean value S n,t When it is True, it means that the code has no syntax problems and can be executed. Otherwise, it means that the generated code has syntax problems and needs to be rewritten.

4. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 3, characterized in that: The planning module comprises: The obstacle avoidance planner has the mobile target point and the information of the perception module as input and the obstacle avoidance path planned in real time as output; The tracking planner takes as input the ID of the object to be tracked and the information of the perception module, and outputs a path that avoids obstacles in real time and can follow the target object.

5. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 1, characterized in that: It also includes a safety function module, which adopts a high-low layer structure. The high-level structure is a safety visual language model installed on the cloud platform, and the low-level structure is a collision detector installed on the flying intelligent body; the safety visual language model performs privacy safety detection and potential safety detection based on the observation signal collected by the flying intelligent body; the collision detector performs collision detection based on the observation signal collected by the flying intelligent body and the planned path of the planning function module.

6. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 5, characterized in that: The privacy security detection is specifically: pl t ,pr t =f privacy_check (O t , Among them, f privacy_check It is a privacy detection module; O t is the observation signal collected by the flying agent at time t and the state of the drone, pr t and rl t , is the privacy level, privacy score reason; The potential safety detection is specifically: rl t ,rr t =f risk_check (O t ) Among them, f risk_check It is a potential risk detection module, rl t ,rr t They are the potential risk score and the reason for the potential risk evaluation.

7. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 5, characterized in that: The collision detection is specifically: Co t =f collision_check (O t ,p t ) f collision_check It is a collision detection module, whose input is the observation signal O at time t t And the path planning information pt at time t, the collision detection function predicts the position of dynamic objects in the environment through the Kalman filtering method and performs collision detection with the planned trajectory.

8. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 5, characterized in that: The safety function module performs safety assessment and safety processing according to the privacy detection results, potential safety detection results and collision detection results.

9. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 8, characterized in that: The safety assessment includes: pl t and rl t is an integer from 1 to 4, pr t With rr t The reason for the score, in the form of a string; 1 means no security threat, 2 means slight security threat but under control, 3 means moderate security threat and the user needs to be informed to terminate the task, and 4 means severe security threat. Co t It is the Boolean value of the collision detection result at time t; when this variable is True, it means that the drone is based on p at this moment. t and the observed signal O t A collision risk is detected; otherwise, there is no collision risk.

10. The safe flight intelligent agent system for executing complex language instructions in a high dynamic environment according to claim 9, characterized in that: The security processing includes: When the privacy and potential risk evaluation is 3, the security function module directly sends a warning message to the task manager, which sends it to the user and terminates the task; When the risk assessment is 4, the safety processing module controls the flight agent to terminate the mission and return to the last safe point; For collision detection, when the safety function module detects a collision, the collision detector directly notifies the flight agent to immediately perform emergency braking to avoid the collision, and at the same time communicates with the task manager to inform the user.

Citation Information

Patent Citations

  • Universal system of intelligent robot with body, construction method and use method

    CN117549310A

  • Universal system of intelligent robot with body

    CN118906074A

  • Intelligent automatic wheelchair driving system and method based on large language model

    CN118939228A

  • Neural task planner for autonomous vehicles

    US20210223774A1

  • Using scene understanding to generate context guidance in robotic task execution planning

    WO2024182721A1

Cited By

  • Multi-level industrial unmanned aerial vehicle control system and method

    CN120972747A

  • Multi-agent-based unmanned aerial vehicle autonomous flight control method and server

    CN120993955A

  • Geospatial analysis execution method and system oriented to GeoJSON (Geographic JavaScript Object Notation)

    CN121388068A