Unmanned aerial vehicle navigation method and device based on vision and language, terminal equipment and storage medium
Through the navigation method combining vision and language, the pre-trained model is used for cross-modal matching and dynamic control, the problem of drone navigation in niche scenarios and unknown environments is solved, and the stable autonomous flight of drones in complex environments is achieved.
Patent Information
- Application Number
- CN202510585615.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
AI Technical Summary
Existing UAV navigation technology is difficult to achieve effective navigation in niche scenarios or unknown environments, especially when facing complex electromagnetic conditions and unforeseen interference in urban environments, the direct control method lacks flexibility and autonomy, resulting in flight instability.
Using vision and language-based navigation methods, by acquiring multi-view images and natural language navigation instructions of the drone environment, pre-trained visual language models and large language models are used to match across modalities, identify landmark features and generate feasible paths, and combine six-degree of freedom dynamic model and nonlinear predictive control to realize drone navigation.
Implement autonomous navigation of drones in unknown environments, reduce dependence on high-quality labeled data, improve navigation flexibility and autonomy, and ensure stable flight of drones in complex environments.
Smart Images

Figure CN120333456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) navigation, and particularly to a UAV navigation method, device, terminal device, and storage medium based on vision and language. Background Art
[0002] Traditional UAV navigation usually adopts direct control methods, that is, controlling the movement of the UAV by sending preset instructions to the UAV. These methods are usually based on predefined paths or real-time sensor data and are crucial for tasks that require precise maneuvering. The direct control methods have high efficiency in tasks with clear goals and stable environments. However, natural language instructions may be ambiguous, leading to misunderstandings by the UAV, which may cause incorrect operations of the UAV or task failures. In urban environments, in the face of collapsed building structures, complex electromagnetic conditions, or unexpected interferences, the direct control methods lack flexibility, may lead to unstable flights, and thus limit their dynamic interaction capabilities. In addition, the direct control methods lack autonomy and cannot independently process tasks, so they are not very suitable for long-term or remote operations.
[0003] End-to-end UAV navigation systems utilize machine learning techniques to directly map sensor inputs to control actions, thus eliminating the need for manual feature extraction and decision-making processes. These systems show potential in processing complex data structures and providing better generalization capabilities in different environments. Instruction-level control involves the interpretation of high-level commands or instructions to guide the behavior of the UAV. This method makes the interaction with the UAV closer to human behavior, enabling it to adapt to changing environments according to natural language instructions. Large language models (LLMs) adopting the Transformer architecture have played an important role in enhancing instruction-level control. By providing the ability to understand and generate human-like text, they simplify the control of the UAV and enable it to handle complex and real-time task adjustments. However, end-to-end strategies require a large amount of high-quality labeled data to achieve good performance, which may be difficult to obtain in specific or niche applications. In addition, these models may not generalize well to new or unseen environments, resulting in performance degradation. Therefore, it is difficult for the prior art to achieve UAV navigation in niche scenarios or unknown environments. Summary of the Invention
[0004] Embodiments of the present invention provide a UAV navigation method, device, terminal device, and storage medium based on vision and language. The present invention can solve the problem that it is difficult for existing UAV navigation technologies to achieve UAV navigation in niche scenarios or unknown environments.
[0005] An embodiment of the present invention provides a UAV navigation method based on vision and language, including:
[0006] Obtaining images of each perspective of the environment where the UAV is located and natural language navigation instructions;
[0007] Divide the images of each perspective of the environment where the drone is located into a preset number of specific regions, monitor the landmark information of each specific region in the images of each perspective through a pre-trained vision-language model, and obtain the visual features of each specific region; through a pre-trained large language model, extract landmark phrases from natural language navigation instructions, and extract landmark word features for each landmark phrase;
[0008] Based on the pre-trained vision-language model, through the cross-attention mechanism, perform cross-modal matching between the landmark word features corresponding to each landmark phrase and the visual features to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; based on each landmark phrase and the potential landmark candidates corresponding to each landmark phrase, determine all target landmarks through a pre-trained large language model; obtain a feasible path according to all target landmarks;
[0009] Based on the feasible path, realize the navigation of the drone.
[0010] Further, the obtaining of the images of each perspective of the environment where the drone is located includes:
[0011] Based on the shooting device equipped on the drone, control the drone to rotate itself to obtain the images of each perspective of the environment where the drone is located.
[0012] Further, the performing, based on the pre-trained vision-language model, of cross-modal matching between the landmark word features corresponding to each landmark phrase and the visual features through the cross-attention mechanism to obtain a preset number of potential landmark candidates corresponding to each landmark phrase includes:
[0013] Based on the pre-trained vision-language model, through the cross-attention mechanism, perform cross-modal matching between the word features corresponding to each landmark phrase and the visual features of each specific region to obtain a similarity matrix of the word features and the visual features of each specific region, and perform normalization processing on the similarity matrix to obtain a weight matrix;
[0014] Based on the weight matrix, perform aggregation processing on the weights of all word features of each specific region to obtain the comprehensive matching scores of each specific region, sort the comprehensive matching scores in descending order, and obtain a preset number of potential landmark candidates corresponding to each landmark phrase according to the sorting result.
[0015] Further, the determining of all target landmarks through a pre-trained large language model based on each landmark phrase and the potential landmark candidates corresponding to each landmark phrase includes:
[0016] Generate descriptive texts for each landmark phrase and its corresponding potential landmark candidates through a pre-trained large language model, and perform natural language matching evaluation on the descriptive texts to obtain the natural language metric similarity scores of each potential landmark candidate;
[0017] Take the potential landmark candidate with the highest natural language metric similarity score as the target landmark.
[0018] Further, control the drone to achieve drone navigation through the following steps:
[0019] Build a six-degree-of-freedom dynamics model, obtain the state data of the drone through the six-degree-of-freedom dynamics model, and based on the state data of the drone and the state data of the landmark obstacles on the feasible path, obtain the control instructions of the drone through a non-linear predictive control model;
[0020] Convert the control instructions of the drone into motor instructions through a low-level attitude controller, and complete the navigation of the drone based on the motor instructions.
[0021] Further, the six-degree-of-freedom dynamics model includes:
[0022]
[0023] In the formula, is the derivative of the position of the drone, v(t) is the linear velocity of the drone, is the derivative of the velocity, the rotation matrix R(φ,θ)∈SO(3) describes the attitude of the drone in Euler form, T is the thrust, g is the acceleration due to gravity, A is the linear damping coefficient matrix, A x is the linear damping coefficient corresponding to the x-axis direction, A y is the linear damping coefficient corresponding to the y-axis direction, A z is the linear damping coefficient corresponding to the z-axis direction, is the derivative of the roll angle, τ φ is the time constant of the roll channel, K φ is the gain of the roll channel, φ ref (t) is the system reference input roll angle, φ(t) is the roll angle, is the derivative of the pitch angle, τ θ is the time constant of the pitch channel, K θ is the gain of the pitch channel, θ ref (t) is the system reference input pitch angle, θ(t) is the pitch angle.
[0024] Further, during the navigation of the drone, it also includes:
[0025] Store the historical flight trajectory of the drone and the corresponding natural language navigation instructions in the form of a graph;
[0026] When there are ambiguous navigation instructions in the natural language navigation instructions, based on the stored historical flight trajectories and the corresponding natural language navigation instructions, a feasible path is derived.
[0027] Another embodiment of the present invention provides a vision- and language-based UAV navigation device, including:
[0028] A data acquisition module, a feature extraction module, a feasible path determination module, and a navigation determination module:
[0029] The data acquisition module is used to acquire images of each perspective of the environment where the UAV is located and natural language navigation instructions;
[0030] The feature extraction module is used to divide the image of each perspective of the environment where the UAV is located into a preset number of specific regions, monitor the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model to obtain the visual features of each specific region; through a pre-trained large language model, extract landmark phrases from the natural language navigation instructions, and extract landmark word features for each landmark phrase;
[0031] The feasible path determination module is used to, based on the pre-trained vision-language model, perform cross-modal matching between the landmark word features corresponding to each landmark phrase and the visual features through a cross-attention mechanism to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, determine all target landmarks through a pre-trained large language model; obtain a feasible path according to all target landmarks;
[0032] The navigation determination module is used to implement the navigation of the UAV based on the feasible path.
[0033] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a vision- and language-based UAV navigation method as described in any one of the above embodiments.
[0034] Another embodiment of the present invention provides a storage medium, where the storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute a vision- and language-based UAV navigation method as described in any one of the above embodiments.
[0035] By implementing the present invention, the following beneficial effects are achieved:
[0036] The present invention discloses a method, apparatus, terminal device and storage medium for drone navigation based on vision and language. The method includes obtaining images of each perspective of the environment where the drone is located and natural language navigation instructions; dividing the image of each perspective of the drone's environment into a preset number of specific regions, and monitoring the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model to obtain the visual features of each specific region; extracting landmark phrases from the natural language navigation instructions through a pre-trained large language model, and extracting landmark word features for each landmark phrase; based on the pre-trained vision-language model, performing cross-modal matching between the landmark word features corresponding to each landmark phrase and the visual features through a cross-attention mechanism to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; determining all target landmarks based on each landmark phrase and the potential landmark candidates corresponding to each landmark phrase through a pre-trained large language model; obtaining a feasible path based on all target landmarks; and realizing the navigation of the drone based on the feasible path. By integrating the vision-language model and the large language model, and leveraging the features and knowledge learned by the pre-trained model from massive data, the present invention significantly reduces the dependence on a large amount of high-quality labeled data. The cross-modal matching of landmark word features and visual features achieved through the cross-attention mechanism also avoids the need for additional labeled data. Additionally, by dividing the drone perspective images into specific regions and extracting visual features, the model can capture local invariant features in the environment. Combining the understanding of landmark phrases by the language model, it can effectively identify potential landmarks and generate feasible paths even in the face of a completely new or unseen environment, thus realizing the navigation of the drone. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 FIG. is a schematic flowchart of a method for drone navigation based on vision and language according to an embodiment of the present invention.
[0038] Figure 2 FIG. is a schematic structural diagram of a method for drone navigation based on vision and language according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.
[0041] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality" means more than two unless otherwise specifically defined.
[0042] As Figure 1 shown, to solve the problem that it is difficult for existing UAV navigation technologies to achieve UAV navigation in niche scenarios or unknown environments, an embodiment of the present invention provides a UAV navigation method based on vision and language, including the following steps:
[0043] Step S1, obtain images of each perspective of the environment where the UAV is located and natural language navigation instructions;
[0044] In a preferred embodiment, the obtaining of images of each perspective of the environment where the UAV is located includes:
[0045] Based on the shooting device equipped on the UAV, obtain images of each perspective of the environment where the UAV is located by controlling the UAV to rotate itself.
[0046] Specifically, the UAV is equipped with a wide-angle camera at the front end of its fuselage. The camera can rotate 90 degrees up and down along the pitch axis but cannot rotate left and right. In order to comprehensively perceive the surrounding environment, the UAV needs to rotate itself to capture images of each perspective of the environment where the UAV is located; at the same time, obtain natural language navigation instructions. The natural language navigation instructions include azimuth descriptions and landmark pointers, and can be obtained through voice recognition or text input interfaces, providing a semantic basis for path planning.
[0047] Step S2, divide the image of each perspective of the environment where the UAV is located into a preset number of specific regions, monitor the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model, and obtain the visual features of each specific region; through a pre-trained large language model, extract landmark phrases from the natural language navigation instructions, and extract landmark word features for each landmark phrase;
[0048] When the drone identifies landmarks, it converts the spatial data of these landmarks into text form for subsequent motion analysis. However, a simple spatial description relying solely on the drone's perspective is insufficient for a large language model to accurately judge the relative positions between landmarks. For example, a road detected in the front view of the drone may be in the center or off to one side, but only knowing it is in the center would prompt the drone to drive along the road.
[0049] Therefore, in the present invention, by introducing a high-resolution spatial descriptor, the spatial descriptor divides the image of each perspective of the environment where the drone is located into a preset number of specific regions. Schematically, the spatial descriptor divides the image of each perspective of the environment where the drone is located into nine specific regions, and each region provides a fine-grained spatial descriptor, using "#0" to "#8" to represent each specific region. For example, "#0" represents the region in the upper left quadrant. The pre-trained vision-language model monitors the landmark information in each specific region of the image of each perspective and extracts the visual features of each specific region.
[0050] Schematically, the pre-trained vision-language model can be GroundingDINO;
[0051] In addition, this application uses a pre-trained large language model as a text analysis tool to extract landmark phrases from natural language navigation instructions and extract landmark word features for each landmark phrase;
[0052] Specifically, the pre-trained large language model extracts landmark phrases from natural language navigation instructions and compiles them into a set L, which contains elements from l1 to l n The pre-trained large language model decomposes natural language navigation instructions into multiple landmark phrases to facilitate step-by-step reasoning and landmark identification. At the same time, it parses each landmark phrase to generate landmark word features; among them, the landmark phrase extraction process is expressed as:
[0053] L = LLM(T, prompt);
[0054] In the formula, LLM represents the large language model performing a specific task, T is the natural language navigation instruction, and prompt is the prompt information;
[0055] Schematically, the pre-trained large language model can be BERT, GPT 4, and GPT 4 Turbo.
[0056] Step S3: Based on the pre-trained vision-language model, through the cross-attention mechanism, cross-modal matching is performed between the landmark word features corresponding to each landmark phrase and each visual feature to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, all target landmarks are determined through the pre-trained large language model; a feasible path is obtained according to all target landmarks.
[0057] In a preferred embodiment, the cross-modal matching between the landmark word features corresponding to each landmark phrase and each visual feature through the cross-attention mechanism based on the pre-trained vision-language model to obtain a preset number of potential landmark candidates corresponding to each landmark phrase includes:
[0058] Based on the pre-trained vision-language model, through the cross-attention mechanism, cross-modal matching is performed between the word features corresponding to each landmark phrase and the visual features of each specific region to obtain a similarity matrix of the word features and the visual features of each specific region, and the similarity matrix is normalized to obtain a weight matrix.
[0059] Based on the weight matrix, the weights of all word features in each specific region are aggregated to obtain the comprehensive matching score of each specific region, the comprehensive matching scores are sorted in descending order, and a preset number of potential landmark candidates corresponding to each landmark phrase are obtained according to the sorting result.
[0060] Specifically, based on the pre-trained vision-language model, through the cross-attention mechanism, the similarity between the landmark word features of each landmark phrase and the visual features of each specific region of the image is calculated to generate an original similarity matrix. Schematically, the dot product similarity or cosine similarity can be used to calculate the similarity between the landmark word features of each landmark phrase and the visual features of each specific region of the image; the similarity matrix is normalized to obtain the weight distribution of each word feature in different visual regions. For each visual region, the weights of all word features associated with it are weighted and summed to obtain the comprehensive matching score of this region. After sorting from high to low according to the scores, the preset number of regions with the highest rankings are selected as potential landmark candidates. For example, the top 3 regions with the highest scores are selected as potential landmark candidates. This process enables the key landmark phrases in the semantic instruction to establish associations with multiple potential regions in the scene, avoiding navigation deviations caused by the failure of a single region match. Among them, the landmark candidates are obtained through visual features. Applying the SoftMax function to the similarity matrix, the similarity of each landmark feature word to all specific regions is converted into a probability distribution to obtain a weight matrix, and the weight matrix satisfies the following formula:
[0061]
[0062] In the formula, W ijis the normalized attention weight between the i-th landmark word feature and the j-th specific region, S ij represents the original similarity score between the i-th landmark word feature and the j-th specific region, represents the exponential operation on the original similarity score, N regions is the total number of specific regions into which the image is divided.
[0063] For each specific region j, aggregate the weights of all landmark word features through the following formula:
[0064]
[0065] In the formula, RegionScore j is the comprehensive matching score of the specific region j, N words is the total number of landmark word features in the landmark phrase, w i is the adjustable importance weight of the i-th landmark word feature.
[0066] In a preferred embodiment, based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, determine all target landmarks through a pre-trained large language model; including:
[0067] Generate descriptive texts for each landmark phrase and its corresponding potential landmark candidates through a pre-trained large language model, and perform natural language matching evaluation on the descriptive texts to obtain the natural language metric similarity scores of each potential landmark candidate;
[0068] Take the potential landmark candidate with the highest natural language metric similarity score as the target landmark.
[0069] Specifically, generate descriptive texts for each landmark phrase and its corresponding potential landmark candidates through a pre-trained large language model. Among them, the descriptive texts include (1) a detailed description of the features, (2) the basis for matching: the former is a detailed explanation of the visual features or geographical features to help understand why a certain potential match is considered relevant to the query; the latter explains how the system relates a certain potential match to the query according to feature fusion and similarity evaluation. Then, perform natural language matching evaluation on these descriptive texts, calculate the similarity between each landmark phrase and each potential match, obtain the natural language metric similarity scores of each potential landmark candidate, and select the potential landmark candidate with the highest similarity score as the target landmark.
[0070] In the present invention, through a pre-trained vision-language model, the cross-attention mechanism is used to perform cross-modal matching between the landmark word features in the landmark phrase and the visual features of the image, and a preset number of potential landmark candidates are found for each landmark phrase; then, based on these landmark phrases and their candidates, relying on the powerful semantic understanding and analysis capabilities of the pre-trained large language model, all accurate target landmarks that meet the requirements are determined from the candidates; finally, based on all the determined target landmarks, a feasible path connecting these landmarks is obtained through a suitable path planning method, thereby realizing the function from text description to actual path planning.
[0071] Step S4: Based on the feasible path, realize the navigation of the drone.
[0072] In a preferred embodiment, the drone is controlled to realize the navigation of the drone through the following steps:
[0073] Construct a six-degree-of-freedom dynamics model, obtain the state data of the drone through the six-degree-of-freedom dynamics model, and based on the state data of the drone and the state data of the landmark obstacles on the feasible path, obtain the control instructions of the drone through a non-linear predictive control model;
[0074] Through a low-level attitude controller, convert the control instructions of the drone into motor instructions, and based on the motor instructions, complete the navigation of the drone.
[0075] Specifically, the six-degree-of-freedom dynamics model refers to a physical model that describes the position and attitude changes of the drone in three-dimensional space. This model provides precise physical constraints for non-linear predictive control by calculating the position derivative, velocity derivative, and attitude angle derivative in real time; the non-linear predictive control model refers to a control method based on rolling horizon optimization, which can be specifically implemented by using a model predictive control algorithm. For example, in each control cycle, an optimization problem including dynamic constraints and obstacle avoidance is solved. This model realizes path tracking in a dynamic environment by fusing the path planning results and real-time obstacle information.
[0076] During the flight of the drone, it continuously obtains its own position, speed, and attitude information, and at the same time monitors the state of dynamic obstacles on the path. The six-degree-of-freedom dynamics model calculates the motion state of the drone in real time through a system of differential equations, providing a physical constraint boundary for predictive control. The non-linear predictive control model solves the optimal control sequence including obstacle avoidance constraints based on the current state and future path planning, generating control instructions that meet the dynamic characteristics. The drone state data and obstacle trajectories are input into the non-linear model predictive control module as parameters of the solver, and these parameters also include the reference control action u ref , the obstacle radius r obs , the reference position x ref and the safety radius r safety. The low-level attitude controller calculates the thrust component and moment component of the control instruction, converts them into specific rotational speed instructions for the four motors, and drives the UAV to fly along the planned path. These three links form a closed-loop control system, ensuring motion feasibility through physical model constraints, achieving dynamic obstacle avoidance through predictive control, and guaranteeing control accuracy at the underlying execution level.
[0077] In a preferred embodiment, the six-degree-of-freedom dynamics model includes:
[0078]
[0079] where is the derivative of the UAV's position, v(t) is the linear velocity of the UAV, is the derivative of the UAV's velocity, the rotation matrix R(φ,θ) ∈ SO(3) describes the UAV's attitude in Euler form, T is the thrust, g is the acceleration due to gravity, A is the linear damping coefficient matrix, A x is the linear damping coefficient corresponding to the x-axis direction, A y is the linear damping coefficient corresponding to the y-axis direction, A z is the linear damping coefficient corresponding to the z-axis direction, is the derivative of the roll angle, τ φ is the time constant of the roll channel, K φ is the gain of the roll channel, φ ref (t) is the system reference input roll angle, φ(t) is the roll angle, is the derivative of the pitch angle, τ θ is the time constant of the pitch channel, K θ is the gain of the pitch channel, θ ref (t) is the system reference input pitch angle, θ(t) is the pitch angle.
[0080] In a preferred embodiment, during the navigation of the UAV, it further includes:
[0081] Storing the UAV's historical flight trajectory and the corresponding natural language navigation instructions in the form of a graph;
[0082] When there are fuzzy navigation instructions in the natural language navigation instructions, deriving a feasible path based on the stored historical flight trajectory and the corresponding natural language navigation instructions.
[0083] Specifically, in some cases, natural language navigation instructions may contain consecutive ambiguous descriptions and have no landmarks, such as "turn left, then turn right, then go straight". There are no ground objects for reference in such instructions. The drone cannot derive a navigable path from such ambiguous instructions, and this situation will directly lead to the repeated execution of steps. Therefore, the historical flight trajectory of the drone and the corresponding natural language navigation instructions are stored in the form of a graph, where the nodes represent the previously encountered landmarks, and the edges store the specific navigation instructions between the landmarks. For any two nodes, a navigable path can be retrieved through the shortest path algorithm. Therefore, when there are ambiguous navigation instructions in the natural language navigation instructions, a feasible path can be derived based on the stored historical flight trajectory and the corresponding natural language navigation instructions.
[0084] As Figure 2 shown, it is a schematic structural diagram of a vision- and language-based drone navigation device provided by an embodiment of the present invention, including:
[0085] A data acquisition module, a feature extraction module, a feasible path determination module, and a navigation determination module:
[0086] The data acquisition module is used to acquire images of each perspective of the environment where the drone is located and natural language navigation instructions;
[0087] The feature extraction module is used to divide the image of each perspective of the environment where the drone is located into a preset number of specific regions, monitor the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model to obtain the visual features of each specific region; through a pre-trained large language model, extract landmark phrases from the natural language navigation instructions, and extract landmark word features for each landmark phrase;
[0088] The feasible path determination module is used to, based on the pre-trained vision-language model, perform cross-modal matching of the landmark word features corresponding to each landmark phrase with each visual feature through a cross-attention mechanism to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, determine all target landmarks through a pre-trained large language model; and obtain a feasible path according to all target landmarks;
[0089] The navigation determination module is used to implement the navigation of the drone based on the feasible path.
[0090] It can be understood that the above device item embodiment corresponds to the method item embodiment of the present invention, and it can implement a vision- and language-based drone navigation method provided by any one of the above method item embodiments of the present invention.
[0091] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.
[0092] Those skilled in the art can clearly understand that for the convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein.
[0093] Another preferred embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a vision- and language-based UAV navigation method as described in any one of the above embodiments.
[0094] It should be noted that the terminal device mentioned here can be computing devices such as desktop computers, notebooks, palm computers, and cloud servers. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that, for example, it may also include input / output devices, network access devices, buses, etc.
[0095] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device through various interfaces and lines.
[0096] The memory can be used to store the computer program. By running or executing the computer program stored in the memory and invoking the data stored in the memory, the processor realizes various functions of the terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0097] Another preferred embodiment of the present invention provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute a method for visual and language-based UAV navigation according to any one of the present invention.
[0098] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0099] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A vision- and language-based UAV navigation method, characterized in that, Including: Obtaining images of each perspective of the environment where the drone is located and natural language navigation instructions; Dividing the image of each perspective of the environment where the drone is located into a preset number of specific regions, monitoring the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model to obtain the visual features of each specific region; extracting landmark phrases from the natural language navigation instructions through a pre-trained large language model, and extracting landmark word features for each landmark phrase; Based on the pre-trained vision-language model, cross-modal matching is performed between the landmark word features corresponding to each landmark phrase and the visual features through a cross-attention mechanism to obtain a preset number of potential landmark candidates corresponding to each landmark phrase; Based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, all target landmarks are determined through a pre-trained large language model; a feasible path is obtained based on all target landmarks; Based on the feasible path, the navigation of the drone is realized.
2. The method for visual and language-based UAV navigation according to claim 1, wherein, The obtaining of images of each perspective of the environment where the drone is located includes: Based on the shooting device equipped on the drone, images of each perspective of the environment where the drone is located are obtained by controlling the rotation of the drone itself.
3. The method for visual- and language-based UAV navigation according to claim 2, wherein, The cross-modal matching of the landmark word features corresponding to each landmark phrase and the visual features through the cross-attention mechanism based on the pre-trained vision-language model to obtain a preset number of potential landmark candidates corresponding to each landmark phrase includes: Based on the pre-trained vision-language model, cross-modal matching is performed between the word features corresponding to each landmark phrase and the visual features of each specific region through a cross-attention mechanism to obtain a similarity matrix of the word features and the visual features of each specific region, and the similarity matrix is normalized to obtain a weight matrix; Based on the weight matrix, the weight of all word features in each specific region is aggregated to obtain the comprehensive matching score of each specific region, the comprehensive matching scores are sorted in descending order, and a preset number of potential landmark candidates corresponding to each landmark phrase are obtained according to the sorting result.
4. The method for visual- and language-based UAV navigation according to claim 3, wherein, The determination of all target landmarks through the pre-trained large language model based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase includes: Generating descriptive texts for each landmark phrase and the corresponding potential landmark candidates through a pre-trained large language model, and performing natural language matching evaluation on the descriptive texts to obtain the natural language index similarity scores of each potential landmark candidate; Taking the potential landmark candidate with the highest natural language index similarity score as the target landmark.
5. The method for drone navigation based on vision and language according to claim 4, wherein The drone is controlled to realize the navigation of the drone through the following steps: A six-degree-of-freedom dynamics model is constructed, state data of the drone is obtained through the six-degree-of-freedom dynamics model, and based on the state data of the drone and the state data of the landmark obstacles on the feasible path, a control instruction of the drone is obtained through a non-linear predictive control model; Through a low-level attitude controller, the control instruction of the drone is converted into a motor instruction, and based on the motor instruction, the navigation of the drone is completed.
6. The method for visual and language-based UAV navigation according to claim 5, wherein The six-degree-of-freedom dynamics model includes: In the formula, is the position derivative of the UAV, v(t) is the linear velocity of the UAV, is the velocity derivative of the UAV. The rotation matrix R(φ,θ) ∈ SO(3) describes the attitude of the UAV in Euler form, T is the thrust, g is the acceleration due to gravity, A is the linear damping coefficient matrix, and A x is the linear damping coefficient corresponding to the x-axis direction, and A y is the linear damping coefficient corresponding to the y-axis direction, and A z is the linear damping coefficient corresponding to the z-axis direction. is the roll angle derivative, and τ φ is the time constant of the roll channel, and K φ is the gain of the roll channel, φ ref (t) is the system reference input roll angle, and φ(t) is the roll angle. is the pitch angle derivative, and τ θ is the time constant of the pitch channel, and K θ is the gain of the pitch channel, θ ref (t) is the system reference input pitch angle, and θ(t) is the pitch angle.
7. The method for visual and language-based UAV navigation according to claim 6, wherein, During the navigation of the drone, it also includes: Store the historical flight trajectory of the drone and the corresponding natural language navigation instructions in the form of a graph; When there are ambiguous navigation instructions in the natural language navigation instructions, derive a feasible path based on the stored historical flight trajectory and the corresponding natural language navigation instructions.
8. An unmanned aerial vehicle navigation device based on vision and language, characterized in that, Including: A data acquisition module, a feature extraction module, a feasible path determination module, and a navigation determination module: The data acquisition module is used to acquire images of each perspective of the environment where the drone is located and natural language navigation instructions; The feature extraction module is used to divide the image of each perspective of the environment where the drone is located into a preset number of specific regions, monitor the landmark information of each specific region in the image of each perspective through a pre-trained vision-language model, and obtain the visual features of each specific region; through a pre-trained large language model, extract landmark phrases from the natural language navigation instructions, and extract landmark word features for each landmark phrase. The feasible path determination module is used to, based on the pre-trained vision-language model, perform cross-modal matching between the landmark word features corresponding to each landmark phrase and the visual features through a cross-attention mechanism, and obtain a preset number of potential landmark candidates corresponding to each landmark phrase; Based on each landmark phrase and each potential landmark candidate corresponding to each landmark phrase, determine all target landmarks through a pre-trained large language model; obtain a feasible path according to all target landmarks; The navigation determination module is used to, based on the feasible path, implement the navigation of the drone.
9. A terminal device, characterized in that, Including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a vision- and language-based drone navigation method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute a vision- and language-based drone navigation method according to any one of claims 1 to 7.
Citation Information
Cited By
Visual language navigation method for cross-modal alignment in dynamic shielding environment
CN121067831A
Unmanned aerial vehicle visual language navigation zero fine tuning method based on fine-grained cognitive function module integration
CN121230728A