Unmanned aerial vehicle navigation reasoning method and system based on visual perception and large language model
By combining lightweight visual perception and large language model drone navigation technology, the autonomous perception and decision-making of drones in complex environments is achieved, the problem of insufficient navigation accuracy and autonomy in the existing technology is solved, and the navigation performance and task execution efficiency of drones in unknown environments is improved.
Patent Information
- Application Number
- CN202511062805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing drone navigation technology is difficult to achieve high-precision navigation in complex environments, especially in unknown environments or emergencies, and the limited computing resources make it difficult to deploy high-performance models.
The lightweight multi-level visual perception model is combined with the large language model, and visual perception is converted into structured data through visual perception, and logical reasoning is used to use knowledge base and prompt engineering to achieve autonomous navigation decision-making, and the knowledge base is updated in real time at the edge.
UAVs achieve accurate understanding and intelligent decision-making in complex environments, improve navigation performance and task execution efficiency, reduce dependence on cloud computing, and enhance the interpretability and credibility of decisions.
Smart Images

Figure CN120558243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of unmanned aerial vehicle (UAV) navigation and artificial intelligence, and in particular to a UAV navigation reasoning method and system based on visual perception and a large language model. Background Art
[0002] With the vigorous development of the low-altitude economy, drones are being used more and more widely in logistics distribution, power inspection, emergency rescue and other fields. In complex and changeable low-altitude environments, such as urban areas with high-rise buildings and rugged outdoor scenes, drones need to quickly perceive environmental information and make accurate decisions to ensure flight safety and efficient mission completion. Although deep learning technology has promoted the widespread application of machine vision in drone platforms, achieving highly autonomous and intelligent navigation decision-making capabilities is still the core trend of the industry's in-depth development towards complex scenarios.
[0003] Current mainstream drone navigation technologies face multiple bottlenecks. Traditional satellite navigation and inertial navigation technologies have inherent defects. Satellite navigation is prone to positioning inaccuracies in signal-blocked areas, while inertial navigation has the problem of error accumulation over time. Both are difficult to meet the needs of high-precision navigation in complex dynamic environments. Although existing visual navigation methods can achieve a certain degree of environmental perception, they lack the ability to understand semantic information such as obstacles and passable areas in complex scenarios. They lack the ability to deeply analyze and logically reason about environmental semantics, making it difficult to cope with diverse mission requirements.
[0004] Multimodal large models are large-scale AI models that can simultaneously process two or more different types of data (such as images, text, voice, and video) and achieve information interaction and semantic understanding through cross-modal fusion techniques. Their application in the drone sector is primarily reflected in visual language navigation (VLN) technology. However, existing VLN technology is still limited to pre-set scenarios and explicit instructions, lacking the ability to make autonomous judgments and decisions for beyond-visual-range missions, unknown environments, or unexpected situations (such as inclement weather and unexpected obstacles). Furthermore, due to the inherent shortcomings of multimodal large models, VLN technology faces challenges such as excessively high computing performance requirements and high costs. Its training and inference processes require processing massive amounts of pixel data from high-resolution images and complex language sequences, placing extremely high demands on computing chips and data transmission bandwidth. However, the limited embedded hardware resources onboard drones make it difficult to support efficient model operation. Furthermore, the massive data annotation, high-performance computing equipment, and power consumption required for model training make the technology's application costs far beyond the reach of commercial drone scenarios.
[0005] Therefore, there is an urgent need to break through the limitations of traditional visual language navigation technology, give drones the ability to autonomously perceive and make decisions in unknown environments, and improve their navigation performance and task execution efficiency in complex scenarios. Summary of the Invention
[0006] The purpose of this invention is to provide a UAV navigation reasoning method and system based on visual perception and large language model, which deeply integrates visual semantic information with the logical reasoning ability of large language model (LLM), adopts lightweight deployment of multi-level visual perception model and large language model, and realizes the UAV's accurate understanding and intelligent decision-making of complex environments.
[0007] The technical solutions of the present invention are as follows: The UAV navigation reasoning method based on visual perception and large language model includes the following steps: S1, collect images of the target area through the camera carried by the drone; S2. Performing pixel-level perception on the image through a pre-trained visual perception model and converting the visual information into structured data; S3. Using retrieval-enhanced generation combined with a predefined knowledge base to convert the structured data into key sentences that can be processed by a large language model; S4. Using prompt engineering and thought chain technology, the large language model outputs navigation instructions based on the key sentences and the current task; S5. When encountering an unknown type of target, the active learning process is initiated to achieve real-time update of the edge knowledge base through human intervention.
[0008] As a further preferred embodiment of the method of the present invention, the structured data in step S2 includes: entity-relationship-attribute triple knowledge obtained by extracting and associating target categories, spatial positions between targets, and target physical parameters in visual perception results.
[0009] As a further preferred embodiment of the method of the present invention, the retrieval enhancement generation described in step S3 includes: retrieving relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, then weighting the retrieval results, and extracting the knowledge fragments that best match the current scenario.
[0010] As a further preferred embodiment of the method of the present invention, the knowledge base includes a rule base, a case base and a constraint base. The rule base is used to clarify safety distances and traffic rules, the case base is used to store typical scenario solutions, and the constraint base is used to record drone performance limitations.
[0011] As a further preferred embodiment of the method of the present invention, the prompt engineering technology described in step S4 includes: constructing a hierarchical prompt template including a base layer and a task layer, wherein the base layer embeds target categories, instruction styles, and planning constraints, and the task layer combines visual perception results to generate a composite prompt of scene semantics + target state + task requirements.
[0012] As a further preferred embodiment of the method of the present invention, the thought chain technology described in step S4 includes: breaking down a complex navigation task into multiple sub-steps, and guiding the model to reason step by step until the navigation instruction generation is completed.
[0013] As a further preferred embodiment of the method of the present invention, the process of outputting navigation instructions by the large language model based on the key sentences and the current task in step S4 includes: judging the distribution of targets in the current scenario and factors that may affect the flight of the drone based on the rule library, and referring to solutions to similar scenarios in the case library, and then combining the constraint condition library to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
[0014] On the other hand, the present invention also discloses a UAV navigation reasoning system based on visual perception and a large language model, comprising: An image acquisition module is used to acquire images of the target area through a camera carried by the UAV; An environmental perception module, configured to perform pixel-level perception of the image using a pre-trained visual perception model and convert visual information into structured data; A semantic understanding module, configured to convert the structured data into key sentences that can be processed by a large language model by using retrieval-enhanced generation combined with a predefined knowledge base; A navigation reasoning module, which uses prompt engineering and thought chain technology to output navigation instructions based on the key sentences and the current task by a large language model; The knowledge update module is used to start the active learning process when encountering unknown types of targets, and to achieve real-time updates of the edge knowledge base through human intervention.
[0015] As a further preference of the system of the present invention, the structured data in the environmental perception module is entity-relationship-attribute triple knowledge obtained by extracting and associating target categories, spatial positions between targets, and target physical parameters in visual perception results.
[0016] As a further preferred embodiment of the system of the present invention, the semantic understanding module includes: The retrieval enhancement generation unit is used to retrieve relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, and then sort the retrieval results by weight to extract the knowledge fragments that best match the current scenario.
[0017] As a further preference of the system of the present invention, the knowledge base includes: Rule base unit, used to clarify safety distance and traffic rules; Case library unit, used to store typical scenario solutions; Constraint library unit, used to record the performance limitations of the drone.
[0018] As a further preferred embodiment of the system of the present invention, the navigation reasoning module includes: The prompt engineering unit is used to build a hierarchical prompt template consisting of a base layer and a task layer. The base layer embeds target categories, instruction styles, and planning constraints. The task layer combines visual perception results to generate composite prompts of scene semantics, target status, and task requirements. The thinking chain unit is used to break down complex navigation tasks into multiple sub-steps, guiding the model to reason step by step until the navigation instruction generation is completed.
[0019] As a further preference of the system of the present invention, the process of outputting navigation instructions by the large language model according to the key sentences and the current task in the navigation reasoning module includes: judging the distribution of targets in the current scenario and the factors that may affect the flight of the drone based on the rule library unit, and referring to the solutions to similar scenarios in the case library unit, and then combining the constraint library unit to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
[0020] The present invention is beneficial in that: This invention builds a localized knowledge system and combines it with model lightweight processing technology to free drones from their strong dependence on cloud computing. Even in emergency scenarios with network interruptions (such as earthquake-stricken areas and remote mountainous areas), it can quickly complete environmental perception and decision-making, thereby improving the availability of drones under complex communication conditions.
[0021] The present invention utilizes triple knowledge structures to structure visual information, decomposes complex decisions through thought chain technology, and verifies decision compliance with the help of a rule engine, making the entire decision-making process clear and explainable. This not only facilitates technical personnel to review and optimize decision logic, but also solves the problem of poor explainability of traditional decisions based on black box models, and improves the credibility of decisions in scenarios with extremely high requirements for decision accuracy (such as medical rescue and material delivery).
[0022] This invention integrates technologies from multiple fields such as computer vision, natural language processing, and knowledge engineering to achieve a full-link connection from visual perception to semantic understanding to intelligent decision-making; multi-level semantic mapping technology solves the ambiguity problem of visual-language cross-modality, and the task-adapted prompt engineering framework combined with thought chain reasoning improves the interpretability of complex decisions in drone navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a flowchart of a UAV navigation reasoning method based on visual perception and a large language model according to an embodiment of the present invention; Figure 2 This is a framework diagram of a UAV navigation reasoning system based on visual perception and a large language model according to an embodiment of the present invention; Figure 3 This is a framework diagram of a vision-language cross-modal deep fusion collaborative decision-making according to an embodiment of the present invention; Figure 4 This is a flowchart of the construction and evolution of a knowledge system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described clearly and completely below with reference to the accompanying drawings and specific embodiments.
[0025] Existing visual language navigation (VLN) technology is limited to preset scenarios and clear instructions, and lacks autonomous judgment and decision-making capabilities for beyond-visual-range tasks, unknown environments, or emergencies (such as bad weather and unexpected obstacles). In order to overcome the limitations of traditional visual language navigation technology, give drones the ability to autonomously perceive and make decisions in unknown environments, and improve the navigation performance and mission execution efficiency of drones in complex scenarios, this embodiment discloses a drone navigation reasoning method based on visual perception and a large language model. Figure 1 As shown, the following steps are included: S101: Collecting images of the target area through the camera carried by the drone; S102: Performing pixel-level perception on the image using a pre-trained visual perception model, and converting the visual information into structured data; S103: Using retrieval enhancement generation combined with a predefined knowledge base to convert the structured data into key sentences that can be processed by a large language model; S104: Using prompt engineering and thought chain technology, the large language model outputs navigation instructions based on the key sentences and the current task; S105: When encountering an unknown type of target, the active learning process is started, and the edge knowledge base is updated in real time through human intervention.
[0026] The structured data in step S102 includes: entity-relationship-attribute triples obtained by extracting and associating target categories (entities), spatial positions (relationships) between targets, and target physical parameters (attributes, such as size and distance) in the visual perception results; The retrieval enhancement generation in step S103 includes: retrieving relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, then sorting the retrieval results by weight, and extracting the knowledge fragments that best match the current scenario; the knowledge base includes a rule base, a case base, and a constraint base. The rule base is used to clarify safety distances and traffic rules, the case base is used to store typical scenario solutions, and the constraint base is used to record drone performance limitations.
[0027] The prompt engineering technology in step S104 includes: constructing a hierarchical prompt template including a basic layer and a task layer, wherein the basic layer embeds target categories, instruction styles, and planning constraints, and the task layer combines visual perception results to generate a composite prompt of scene semantics + target status + task requirements; the thinking chain technology includes: breaking down complex navigation tasks into multiple sub-steps, guiding the model to reason step by step until the navigation instruction generation is completed; and the process of the large language model outputting navigation instructions based on key sentences and the current task includes: judging the distribution of targets such as vehicles, pedestrians, buildings and vegetation in the current scene and factors that may affect the flight of the drone based on the rule library, and referring to solutions for similar scenarios in the case library, and then combining the constraint condition library to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
[0028] Corresponding to the above method, this embodiment also discloses a UAV navigation reasoning system based on visual perception and large language model, such as Figure 2 Shown, including: The image acquisition module 101 is used to acquire images of the target area through the camera carried by the drone; An environmental perception module 102 is configured to perform pixel-level perception of the image using a pre-trained visual perception model and convert visual information into structured data; Semantic understanding module 103, configured to convert the structured data into key sentences that can be processed by a large language model by using search-enhanced generation combined with a predefined knowledge base; The navigation reasoning module 104 is configured to use prompt engineering and thought chain technology to output navigation instructions based on the key sentences and the current task using a large language model; The knowledge updating module 105 is used to start the active learning process when encountering an unknown type of target, and to achieve real-time update of the edge knowledge base through human intervention.
[0029] In one embodiment, the structured data is entity-relationship-attribute triples obtained by extracting and associating target categories (entities), spatial positions (relationships) between targets, and physical parameters (attributes, such as size and distance) of the targets in visual perception results; In one embodiment, the semantic understanding module includes: The retrieval enhancement generation unit is used to retrieve relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, and then sort the retrieval results by weight to extract the knowledge fragments that best match the current scenario.
[0030] In one embodiment, the knowledge base includes: Rule base unit, used to clarify safety distance and traffic rules; Case library unit, used to store typical scenario solutions; Constraint library unit, used to record the performance limitations of the drone.
[0031] In one embodiment, the navigation reasoning module includes: The prompt engineering unit is used to build a hierarchical prompt template consisting of a base layer and a task layer. The base layer embeds target categories, instruction styles, and planning constraints. The task layer combines visual perception results to generate composite prompts of scene semantics, target status, and task requirements. The thinking chain unit is used to break down complex navigation tasks into multiple sub-steps, guiding the model to reason step by step until the navigation instruction generation is completed.
[0032] In one embodiment, the process of outputting navigation instructions by the large language model based on key sentences and the current task in the navigation reasoning module includes: judging the distribution of targets such as vehicles, pedestrians, buildings and vegetation in the current scene and factors that may affect the flight of the drone based on the rule library unit, and referring to the solutions to similar scenarios in the case library unit, and then combining the constraint library unit to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
[0033] In order to make the embodiments of the present invention clearer, Figure 3 A specific embodiment is shown, which fully and clearly describes the technical solution in the embodiment of the present invention.
[0034] First, the camera on the drone is used to collect images of the target area. In this embodiment, an industrial-grade RGB camera is used to cover the entire scene at a resolution of 1920×1080 and a frame rate of 30fps. Image noise (such as dust and light reflection) is preprocessed through Gaussian filtering to ensure that the visual data meets the requirements of semantic analysis accuracy. Figure 3 The "original image input" step in the process.
[0035] Then, the image is perceived at the pixel level through a multi-level visual perception model. The specific steps are as follows: The UNet semantic segmentation network model is used to achieve pixel-level semantic segmentation of images. Figure 3 Output of the "Semantic Segmentation" module; By identifying entities such as buildings, vegetation, vehicles, pedestrians, etc. in the scene, a semantic map containing scene semantics and target distribution is generated. Figure 3 Output of the "Environment Modeling" module; The YOLOv8 model is used to identify small-sized targets in the image and output the target coordinates and shape. Figure 3 Output of the "Entity Detection" module; Map the target categories in the results into entities, map the spatial position relationships into relationships, and map parameters such as size / speed / priority / neighborhood attributes into attributes; The semantic information is converted into triple knowledge of "entity-relationship-attribute", such as "<vehicle, location, 50 meters ahead>", realizing the conversion from visual information to structured data. Figure 3 The "semantic segmentation / environment modeling / entity detection → scene information" data link is a triplet construction method that can adapt to the knowledge expression of multi-task scenarios (such as rescue and inspection).
[0036] The retrieval-enhanced generation technology is combined with a predefined knowledge base to achieve semantic understanding. The specific steps are as follows: The knowledge base includes a rule base (such as safety distance thresholds), a case base (typical scenario solutions), and a constraint base (drone performance limitations). For example, when a drone enters an urban area to perform an emergency rescue mission, information such as "grid-like distribution of roads" in the knowledge base can help the model quickly identify the direction of the road. Features such as tree height and shape in the knowledge base can also be used to help distinguish trees from other obstacles. Retrieval enhancement generation technology retrieves relevant knowledge based on the entities and relationships in triple knowledge, sorts the results, calculates the semantic similarity between triples and knowledge base entries using the TF-IDF algorithm, and extracts knowledge fragments with a similarity ≥ 0.8; The knowledge fragments are reorganized into key sentences that conform to the input format of the large language model (LLM). For example, when a drone detects an obstacle in front, the system can automatically generate a standardized statement such as "Obstacle: tree, location: left front, height: 2m" based on this system.
[0037] Through hierarchical prompt engineering and thought chain technology, the large language model is driven to output accurate navigation instructions. The specific steps are as follows: The basic layer preferably embeds common constraints (such as "pedestrian safety distance ≥ 5 meters") to adapt to the safety requirements of multiple scenarios; The task layer integrates environmental perception results (e.g., "Semantic map: collapsed building area 200 meters ahead, including three suspected human targets") to construct a composite prompt of "scene semantics + target status + task requirements." This prompt construction method can improve the targeted generation of model instructions. The "search and rescue path planning" task is preferably decomposed into the sub-steps of "identifying dangerous areas → marking personnel positions → avoiding obstacles → generating optimal paths", guiding the large language model to output executable instructions (such as "rotate 45 degrees to the left, fly 5 meters forward, and descend 5 meters in altitude"), corresponding to the attached Figure 3 The "large language model → output instruction" module realizes a closed loop from knowledge understanding to action decision-making. This task decomposition logic can cover multiple types of navigation tasks (such as inspection, rescue, and surveying and mapping).
[0038] When encountering unknown types of targets (such as new building structures, special obstacles), it is preferred to trigger the active learning process, such as Figure 4 The specific steps are as follows: Ternary knowledge groups (such as <new building-structure-attribute>) are supplemented through manual annotation and updated to the decision-making knowledge base after similarity retrieval and context fusion verification; the knowledge base is updated using the "incremental learning + knowledge distillation" method to ensure the compatibility of new knowledge with historical knowledge, and continuously improve the system's adaptability to complex scenarios.
[0039] The present invention uses a drone to collect visible light images of the target area, and uses a semantic segmentation algorithm to perform pixel-level classification of targets such as buildings, vegetation, vehicles, pedestrians, etc. in the image, generating structured data containing the spatial distribution of scene elements; the segmentation results are further converted into language fragments that can be understood by a large language model (LLM), and the scene semantics are deeply analyzed through a pre-trained large language model, and an analysis report containing traffic conditions and flight environment assessment is output, and navigation suggestions are generated based on real-time scene information; a technical chain of "image semantic segmentation-cross-modal semantic conversion-intelligent decision generation" is constructed, and by combining computer vision and natural language processing technologies, cross-modal information fusion from visual perception to semantic understanding is realized; it can be applied to scenarios such as drone urban inspections, traffic monitoring, and low-altitude navigation, effectively solving the problems of drones' understanding and decision-making of dynamic scenes in complex environments, and providing data-driven intelligent support for drone autonomous flight, with the advantages of strong real-time performance, high environmental adaptability, and good decision-making reliability.
[0040] The above embodiments are only used to illustrate the technical solutions of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.
Claims
1. UAV navigation reasoning method based on visual perception and large language model, characterized by: The following steps are involved: S1, collect images of the target area through the camera carried by the drone; S2. Performing pixel-level perception on the image through a pre-trained visual perception model and converting the visual information into structured data; S3. Using retrieval-enhanced generation combined with a predefined knowledge base to convert the structured data into key sentences that can be processed by a large language model; S4. Using prompt engineering and thought chain technology, the large language model outputs navigation instructions based on the key sentences and the current task; S5. When encountering an unknown type of target, the active learning process is initiated to achieve real-time update of the edge knowledge base through human intervention.
2. The UAV navigation reasoning method based on visual perception and large language model according to claim 1 is characterized in that: The structured data in step S2 includes: entity-relationship-attribute triple knowledge obtained by extracting and associating target categories, spatial positions between targets, and target physical parameters in visual perception results.
3. The UAV navigation reasoning method based on visual perception and large language model according to claim 2 is characterized in that: The retrieval enhancement generation described in step S3 includes: retrieving relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, then weighting the retrieval results, and extracting the knowledge fragments that best match the current scenario.
4. The UAV navigation reasoning method based on visual perception and large language model according to claim 1 or 3 is characterized in that: The knowledge base includes a rule base, a case base and a constraint condition base. The rule base is used to clarify safety distances and traffic rules, the case base is used to store typical scenario solutions, and the constraint condition base is used to record drone performance limitations.
5. The UAV navigation reasoning method based on visual perception and large language model according to claim 1 is characterized in that: The prompt engineering technology described in step S4 includes: constructing a hierarchical prompt template including a base layer and a task layer, wherein the base layer embeds target categories, instruction styles, and planning constraints, and the task layer combines visual perception results to generate a composite prompt of scene semantics + target state + task requirements.
6. The UAV navigation reasoning method based on visual perception and large language model according to claim 1 is characterized in that: The thought chain technology described in step S4 includes: breaking down the complex navigation task into multiple sub-steps, guiding the model to reason step by step until the navigation instruction generation is completed.
7. The UAV navigation reasoning method based on visual perception and large language model according to claim 4 is characterized in that: The process of outputting navigation instructions by the large language model according to the key sentences and the current task in step S4 includes: judging the distribution of targets in the current scenario and factors that may affect the flight of the drone based on the rule library, and referring to solutions for similar scenarios in the case library, and then combining the constraint library to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
8. UAV navigation reasoning system based on visual perception and large language model, characterized by: include: An image acquisition module is used to acquire images of the target area through a camera carried by the UAV; An environmental perception module, configured to perform pixel-level perception of the image using a pre-trained visual perception model and convert visual information into structured data; A semantic understanding module, configured to convert the structured data into key sentences that can be processed by a large language model by using retrieval-enhanced generation combined with a predefined knowledge base; A navigation reasoning module, which uses prompt engineering and thought chain technology to output navigation instructions based on the key sentences and the current task by a large language model; The knowledge update module is used to start the active learning process when encountering unknown types of targets, and to achieve real-time updates of the edge knowledge base through human intervention.
9. The UAV navigation reasoning system based on visual perception and large language model according to claim 8 is characterized in that: The structured data in the environmental perception module is entity-relationship-attribute triple knowledge obtained by extracting and associating target categories, spatial positions between targets, and target physical parameters in visual perception results.
10. The UAV navigation reasoning system based on visual perception and large language model according to claim 9 is characterized in that: The semantic understanding module includes: The retrieval enhancement generation unit is used to retrieve relevant rules and cases from the knowledge base based on the entities and relationships in the triple knowledge, and then sort the retrieval results by weight to extract the knowledge fragments that best match the current scenario.
11. The UAV navigation reasoning system based on visual perception and large language model according to claim 8 or 10, characterized in that: The knowledge base includes: Rule base unit, used to clarify safety distance and traffic rules; Case library unit, used to store typical scenario solutions; Constraint library unit, used to record the performance limitations of the drone.
12. The UAV navigation reasoning system based on visual perception and large language model according to claim 8 is characterized in that: The navigation reasoning module includes: The prompt engineering unit is used to build a hierarchical prompt template consisting of a base layer and a task layer. The base layer embeds target categories, instruction styles, and planning constraints. The task layer combines visual perception results to generate composite prompts of scene semantics, target status, and task requirements. The thinking chain unit is used to break down complex navigation tasks into multiple sub-steps, guiding the model to reason step by step until the navigation instruction generation is completed.
13. The UAV navigation reasoning system based on visual perception and large language model according to claim 11 is characterized in that: The process of outputting navigation instructions by the large language model based on the key sentences and the current task in the navigation reasoning module includes: judging the distribution of targets in the current scenario and factors that may affect the flight of the drone based on the rule library unit, and referring to solutions to similar scenarios in the case library unit, and then combining the constraint library unit to generate reasonable navigation reasoning, and finally giving recommendations on safe paths and flight directions.
Citation Information
Patent Citations
Visual language navigation method and system based on large language model
CN118031964A
Large model driven body intelligent agent zero sample target navigation method
CN118258396A
Intelligent agent, indoor navigation method and equipment thereof, medium and product
CN119443287A
General security risk monitoring method and system based on large model capability
CN119964083A
Automated label generation using a machine-learned language model
US20250200356A1
Cited By
Unmanned aerial vehicle visual language navigation zero fine tuning method based on fine-grained cognitive function module integration
CN121230728A
Unmanned aerial vehicle scene understanding method, system and device and storage medium
CN121353962A
An unmanned airport scene understanding method, system, device and storage medium
CN121353962B
Unmanned aerial vehicle natural language multi-modal navigation method and system based on multi-dimensional thinking chain
CN121363964A
Unmanned aerial vehicle air navigation visual language motion control method with active dialogue capability
CN121521133A