Multi-layer intelligent development framework of multi-modal LLM Agent for unmanned aerial vehicle system
By adopting a multi-level intelligent control framework in the UAV system, combining multi-modal large language model, vision, speech model and classic control algorithms, the problems of drones' difficulties in dealing with complex environments and insufficient adaptability are solved, achieving higher adaptability and task success rate.
Patent Information
- Application Number
- CN202411805941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-13
AI Technical Summary
When facing complex and real-time changing environments, existing drone systems have difficulty in dealing with, unable to adapt, and lack the ability to adapt to their flexibility, resulting in mission failure or inefficiency.
A multi-level intelligent control framework based on hierarchical design is adopted, including the brain (multimodal large language model) responsible for high-level task planning, the cerebellum (visual and speech models) responsible for scene perception and feedback, and control execution (classic control algorithm) responsible for underlying real-time control.
Through multimodal perception and layered decision-making, the adaptability and adaptability of the drone are improved, real-time path planning and obstacle avoidance are achieved, and the stability and success rate of task execution are enhanced.
Smart Images

Figure CN119987724A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and unmanned aerial vehicle control, and relates to an intelligent control method for an unmanned aerial vehicle system based on a multimodal large language model (LLM), in particular to a multi-level offline intelligent development framework suitable for unmanned aerial vehicle systems, which integrates multimodal perception, voice interaction, visual detection and other functions to improve the intelligence level and adaptability of unmanned aerial vehicles. Background Art
[0002] Drones are widely used in military reconnaissance, logistics and transportation, environmental monitoring, disaster relief and other fields. Most drone systems use a single perception mode and preset control algorithm, which makes it difficult to handle and unable to adapt to complex real-time changing environments. In recent years, large language models (LLMs) and multimodal models have emerged. They have powerful fuzzy intent understanding and task planning capabilities, but due to their high computing power requirements and response delays, they are not suitable for short-cycle real-time control tasks. Although small models have fast reasoning speed, they lack the ability to autonomously understand complex tasks and intentions.
[0003] Most current drone systems have certain limitations in design and operation mechanisms. They often use a single perception mode, which makes it difficult to fully and accurately grasp the overall picture of complex environments. At the same time, the preset control algorithm makes the drone appear rigid and dull when facing the ever-changing real-time environment, and lacks the ability to adapt flexibly. Once the environment exceeds the processing range of the preset algorithm, the drone will get into trouble and cannot make effective adjustments and decisions in time, resulting in mission failure or inefficiency. This difficulty in handling complex environments and the lack of adaptive capabilities have seriously restricted the application and development of drones in a wider range of fields and more complex tasks. Summary of the invention
[0004] To solve the above problems, the present invention adopts a multi-level intelligent control framework based on hierarchical design, which is as follows:
[0005] 1. Brain (large model)
[0006] 1.1. Functions and tasks
[0007] The high-level center is composed of a multimodal large language model (LLM), which has powerful fuzzy intent understanding and complex task planning capabilities. It is responsible for processing high-level tasks such as natural language interaction and complex task planning, such as parsing user voice commands and planning the overall cruise mission.
[0008] 1.2. Deployment method
[0009] After using tools to quantize and optimize the large model, offline inference is performed on the drone’s onboard hardware (such as the NVIDIA Jetson XaaverNX).
[0010] 1.3. Advantages
[0011] The reasoning speed is relatively slow, but it can deeply understand complex instructions and is suitable for non-urgent tasks that require in-depth analysis.
[0012] 2. Cerebellum (small model of vision and speech)
[0013] 2.1. Functions and tasks
[0014] The intermediate center includes a visual detection model, Beidou GPS joint inertial navigation, a lidar processing module and a voice processing model, and is responsible for scene positioning, perception, obstacle avoidance and voice recognition. Using Beidou GPS joint inertial navigation positioning, based on the YOLO series target detection model and OCR algorithm, it realizes environmental visual perception, identifies target objects and obstacles, performs offline voice recognition and synthesis through the voice module, parses user commands and provides status feedback.
[0015] 2.2. Deployment method
[0016] It does not require high computing power and is suitable for processing scenarios with moderate real-time requirements. It can maintain a relatively fast response speed and has certain voice and visual processing capabilities.
[0017] 2.3. Advantages
[0018] Relatively efficient, it can effectively perceive and analyze rapidly changing environmental information, such as dynamic obstacle avoidance and scene recognition.
[0019] 3. Control execution (classical control algorithm)
[0020] 3.1. Functions and tasks
[0021] The low-level center adopts classic control algorithms such as PID to achieve real-time control of the drone's underlying hardware. It is responsible for hardware-level operations such as sensor data reception, real-time motion control (such as flight attitude control, navigation adjustment), and motor drive to ensure flight stability and basic operations.
[0022] 3.2. Deployment method
[0023] It runs directly in the drone's underlying control hardware and provides real-time closed-loop control capabilities.
[0024] 3.3. Advantages
[0025] It has extremely high timeliness and can respond quickly to sensor feedback. It is suitable for short-cycle real-time control tasks such as flight attitude adjustment and obstacle avoidance.
[0026] When the entire system is running, after the user issues a task, the brain (large model) first parses the task and plans an action plan. Then, during the execution process, the cerebellum (small models of vision and speech) perceives environmental changes and generates feedback. Finally, the control execution (classical control algorithm) adjusts the drone's flight path or posture in real time based on the feedback.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1. Multimodal perception enhances adaptability
[0029] 1.1. Find the best path in real time
[0030] In logistics and transportation, using multimodal perception, drones can use vision and voice to identify and avoid obstacles in real time during delivery, and complete delivery tasks based on the optimal transportation route planned by the large model. In disaster rescue scenarios, with the help of multimodal perception such as vision and voice, trapped people or obstacles can be quickly identified, and the flight path can be quickly adjusted through real-time control systems to effectively respond to complex and changing disaster scene environments.
[0031] 1.2. Enhanced adaptability
[0032] Integrating multimodal input information enables the drone to autonomously perceive complex environmental changes in real time and make adjustments, greatly enhancing its adaptability in different scenarios.
[0033] 2. Efficient hierarchical decision-making and control
[0034] 2.1. Exploiting the advantages of stratification
[0035] The three-layer system combines the advantages of large and small models. The brain (large model) is responsible for handling complex task planning. The cerebellum (small model of vision and speech) focuses on tasks such as scene perception, positioning, obstacle avoidance and speech recognition. Control execution (classical control algorithm) is responsible for real-time control of the underlying hardware to ensure flight stability and basic operations.
[0036] 2.2. Meeting task requirements
[0037] It takes into account both efficient reasoning and real-time response, and can work collaboratively at different levels according to different mission requirements, so that drones can make efficient decisions and control when performing various tasks.
[0038] 3. Offline deployment and edge computing
[0039] 3.1.Technical implementation and application
[0040] By quantitatively optimizing the model, the large model can be successfully run offline on resource-limited embedded devices (such as drone onboard hardware). In tasks such as logistics and transportation, disaster relief, smart agriculture, and urban monitoring, drones do not need to rely on external computing resources in real time and can operate independently.
[0041] 3.2. Improve operational capabilities
[0042] It has greatly improved the independent operation capability of drones, enabling them to operate stably in different environments and expanding their application scope.
[0043] 4. High security and stability
[0044] 4.1. Security and stability guarantee
[0045] The classic control algorithm used in the low-level center plays a key role in various application scenarios. In logistics and transportation, it ensures stable flight attitude and accurate navigation to avoid risks such as cargo falling; in disaster relief, it ensures safe flight and accurate rescue of drones in complex environments; in smart agriculture, it ensures stable spraying operations; and in urban monitoring, it maintains stable patrols and monitoring.
[0046] 4.2. Effect
[0047] Provide reliable real-time control for drones when performing missions, so that they have higher safety and operational stability, effectively reduce accident risks and improve mission success rate.
[0048] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings.
[0050] Figure 1 Shows the overall framework diagram of the present invention;
[0051] Figure 2 The task analysis and planning flow chart of the present invention is shown, wherein the perception module is responsible for perceiving and processing multimodal information from the external environment; the core module is responsible for basic tasks such as memory, thinking and decision-making; and the motion module uses tools to perform tasks and affect the surrounding environment. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] The principle and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0054] 1. Overall framework
[0055] refer to Figure 1 ,This invention builds a multi-level intelligent control framework based on the ,layered design concept, which mainly covers three levels:
[0056] 1.1. Brain (big model): As a high-level center, its core is the multimodal large language model (LLM), which is mainly responsible for processing high-level tasks such as natural language interaction and complex task planning. It is a key part of the entire system to understand user intentions and plan tasks globally.
[0057] 1.2. Cerebellum (small model of vision and speech): It constitutes the intermediate center, integrating the visual detection model, Beidou GPS joint inertial navigation, lidar processing module and speech processing model, focusing on tasks such as scene positioning, perception, obstacle avoidance and speech recognition, providing UAVs with environmental perception and interaction capabilities during mission execution.
[0058] 1.3. Control execution (classic control algorithm): As a low-level center, it uses classic control algorithms such as PID to achieve direct real-time control of the drone's underlying hardware and ensure the basic operation and stability of the drone's flight. These three levels work together and closely cooperate to achieve the drone's intelligent control goals.
[0059] 2. Overall process
[0060] 2.1. Task analysis and planning (brain) can be referenced Figure 2 The core module.
[0061] The user issues mission instructions to the drone through natural language (such as voice) or other means. After receiving these fuzzy instructions, the multimodal large language model (LLM) in the brain parses them with its powerful fuzzy intent understanding ability. Subsequently, the LLM breaks down the task into a series of specific action plans, laying the foundation for the execution of subsequent tasks. For example, when faced with a user's instruction to "conduct a detailed survey of a specific area and record key information", the brain will comprehensively consider various factors, such as regional geographical features, survey priorities, etc., to plan a detailed survey route, determine data collection points, and set task priorities.
[0062] 2.2. Scene perception and feedback (cerebellum) can be referenced Figure 2 The perception module.
[0063] The cerebellum plays an important role during the execution of tasks. The visual detection model uses Beidou GPS combined with inertial navigation for positioning, and visually perceives the environment based on the YOLO series target detection model and OCR algorithm, and can accurately identify surrounding target objects (such as cargo labels in logistics transportation, traffic signs in urban monitoring, etc.) and obstacles. At the same time, the voice module realizes offline voice recognition and speech synthesis functions. On the one hand, it analyzes the user's instructions during the task execution process (such as farmers' voice instructions for crop spraying operations in smart agriculture), and on the other hand, it generates status feedback or prompt voice according to the task status (such as reporting the on-site situation to the command center during disaster relief).
[0064] 2.3. Real-time control and adjustment (control execution) can refer to Figure 2 The motion module.
[0065] The control execution module makes real-time adjustments to the flight path or attitude of the drone based on the feedback information provided by the cerebellum. It is responsible for receiving data from various sensors (such as accelerometers, gyroscopes, magnetometers, barometers, etc.) and performing real-time motion control based on these data. During the flight, the drone's flight attitude is accurately controlled (such as adjusting the roll angle to maintain a stable heading when encountering crosswinds), navigation adjustments are made (correcting the flight direction based on GPS signals and preset paths), and the motor is driven to achieve basic operations such as takeoff, landing, hovering, acceleration, and deceleration of the drone, ensuring the stability and accuracy of the drone's flight and effectively achieving instant avoidance of obstacles.
[0066] 3. Specific implementation of each module
[0067] 3.1. Brain (large model)
[0068] Functional construction: Based on the multimodal large language model (LLM), it has excellent fuzzy intent understanding and complex task planning capabilities, providing key support for the system to understand complex instructions and plan tasks.
[0069] Task execution: Processing various high-level tasks such as natural language interaction and complex task planning. For example, for the user's survey task instructions, it can comprehensively plan the survey route, reasonably plan the data collection points, and scientifically set the task priority, etc., demonstrating strong task planning capabilities.
[0070] Deployment method: With the help of specific tools, large models are quantitatively optimized so that they can be used for offline reasoning on the drone's onboard hardware (such as the NVIDIA Jetson Xaaver NX). Although the reasoning speed is relatively slow, it has obvious advantages when dealing with tasks that are not urgent but require in-depth analysis (such as complex regional survey planning).
[0071] 3.2. Cerebellum (small model of vision and speech)
[0072] Functional construction: It is composed of visual detection model, Beidou GPS joint inertial navigation, lidar processing module and voice processing model to form a comprehensive environmental perception and interaction system.
[0073] Mission execution: Accurate positioning is achieved through Beidou GPS combined with inertial navigation to obtain the precise location information of the drone. Based on the YOLO series target detection model and OCR algorithm, it performs detailed visual perception of the environment and accurately identifies various target objects and obstacles in different application scenarios (such as logistics and transportation, urban monitoring, etc.). The voice module realizes offline speech recognition and speech synthesis functions, and plays an important role in command parsing and information feedback in scenarios such as smart agriculture and disaster relief.
[0074] Deployment method: Since it does not require high computing power, it is suitable for scene processing with moderate real-time requirements. During operation, it can maintain a relatively fast response speed, process visual and voice information in a timely manner, and has certain voice and visual processing capabilities. It can effectively respond to rapidly changing environmental information such as the rapid handling of goods in logistics warehouses, and realize key functions such as dynamic obstacle avoidance and scene recognition.
[0075] 3.3. Control Execution (Classic Control Algorithm)
[0076] Function construction: Use classic control algorithms such as PID to achieve real-time and precise control of the underlying hardware of the drone.
[0077] Mission execution: Responsible for receiving various sensor data and performing real-time motion control based on these data. During flight, accurately adjust the flight attitude of the drone according to different flight conditions (such as crosswind interference), accurately adjust navigation based on GPS signals and preset paths, and drive the motor to realize various basic operations of the drone to ensure flight stability and accuracy.
[0078] Deployment method: It runs directly in the underlying control hardware of the drone to form a real-time closed-loop control function with extremely high timeliness. It can respond quickly to sensor feedback and perform excellently in short-cycle real-time control tasks (such as flight attitude adjustment, obstacle avoidance, etc.).
[0079] Although the spirit and principle of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements contained in the spirit and scope of the attached claims.
[0080] Regarding the limitation of the protection scope of the present invention, those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the protection scope of the present invention.
Claims
1. The multi-modal LLM Agent is used for the multi-level intelligent development framework of UAV systems, which is characterized by: The framework includes: Based on hierarchical design, a multi-level intelligent control framework was constructed, which includes three levels: brain (large model), cerebellum (small models of vision and speech) and control execution (classical control algorithm); the brain, as a high-level center, is responsible for processing high-level tasks such as natural language interaction and complex task planning; the cerebellum, as an intermediate center, focuses on tasks such as scene positioning, perception, obstacle avoidance and speech recognition; control execution, as a low-level center, realizes real-time control of the underlying hardware of the drone; the three levels collaborate with each other to jointly realize the intelligent control of the drone.
2. The multi-modal LLM Agent according to claim 1 is used for a multi-level intelligent development framework of an unmanned aerial vehicle system, characterized in that: The framework process is as follows: S21: Task analysis and planning (brain) The user issues a task to the drone through natural language (such as voice) or other means. The multimodal large language model (LLM) in the brain receives and parses the fuzzy instructions and breaks down the task into a series of specific action plans. S22: Scene perception and feedback (cerebellum) In the process of executing tasks, Xiaoce uses visual detection algorithms to identify environmental changes and generates status feedback or prompts based on the voice module. The visual detection model uses Beidou GPS combined with inertial navigation for positioning, and realizes environmental visual perception based on the YOLO series target detection model and OCR algorithm to identify surrounding target objects, obstacles, etc. The voice module realizes offline voice recognition and voice synthesis to analyze user instructions and status feedback. S23: Real-time control and adjustment (control execution) According to the feedback from the cerebellum, the control execution makes real-time adjustments to the UAV's flight path or attitude to achieve instant obstacle avoidance and flight stability maintenance; the control execution is responsible for hardware-level operations such as sensor data reception, real-time motion control (such as flight attitude control, navigation adjustment), and motor drive.
3. The process of using the multi-modal LLM Agent in the multi-level intelligent development framework of the unmanned aerial vehicle system according to claim 2 is characterized in that: The specific implementation steps of the process are: S31: The task analysis and planning function is constructed by a multimodal large language model (LLM), which has strong fuzzy intent understanding and complex task planning capabilities; it can handle high-level tasks such as natural language interaction and complex task planning. For example, it can process user voice commands such as "conduct a detailed survey of a specific area and record key information" and comprehensively plan the survey task, including determining the survey route, planning data collection points, setting task priorities, etc. The deployment method is to use tools to quantize and optimize large models and perform offline inference on the drone's onboard hardware (such as NVIDIA Jetson Xavier NX); S32: The scene perception and feedback function is constructed by a visual detection model, Beidou GPS joint inertial navigation, a lidar processing module, and a voice processing model. It can use Beidou GPS joint inertial navigation for positioning, obtain the precise location information of the drone, and perform visual perception of the environment based on the YOLO series target detection model and OCR algorithm, such as accurately identifying cargo labels and transportation destination signs in logistics transportation, identifying traffic signs and violations in urban monitoring, etc. It realizes offline speech recognition and speech synthesis through the voice module, such as parsing farmers' voice instructions for crop spraying operations in smart agriculture, and reporting the on-site situation to the command center in voice during disaster relief. Its deployment does not require high computing power and is suitable for scene processing with moderate real-time requirements. It can maintain a relatively fast response speed and process visual and voice information in a timely manner. It also has certain voice and visual processing capabilities, and can effectively respond to rapidly changing environmental information such as rapid handling of goods in logistics warehouses, and realize dynamic obstacle avoidance, scene recognition and other functions. S33: The real-time control and adjustment function is constructed using classic control algorithms such as PID to achieve real-time control of the drone's underlying hardware; it is responsible for receiving data from various sensors (such as accelerometers, gyroscopes, magnetometers, barometers, etc.) and performing real-time motion control based on the data; during flight, it accurately controls the drone's flight attitude, such as adjusting the roll angle in time to maintain a stable heading when encountering crosswinds; makes navigation adjustments, such as correcting the flight direction based on GPS signals and preset paths; drives the motor to achieve basic operations such as drone takeoff, landing, hovering, acceleration, and deceleration to ensure flight stability and accuracy; this functional module can run directly in the drone's underlying control hardware to provide real-time closed-loop control functions; it has extremely high timeliness and can quickly respond to sensor feedback.
Citation Information
Cited By
Hypersonic aircraft intelligent control system based on large language model
CN120386268A
Control device integrating cerebellum model and cross-modal attention
CN120839814A
Multi-level industrial unmanned aerial vehicle control system and method
CN120972747A