An Unmanned Aerial Vehicle Full-Autonomous Navigation Method with Fine-Grained Environment Perception Ability
By carrying multiple sensors on the drone to obtain multimodal data, and using deep reinforcement learning and large-model technology to build a fully autonomous navigation intelligent body of the drone, solving the problems of navigation accuracy and battery life of the drone in the fine-grained environment in the existing technology, and achieving high accuracy, robustness and long-term autonomous navigation of drones.
Patent Information
- Application Number
- CN202411186768.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-28
AI Technical Summary
The existing drone autonomous navigation methods cannot independently identify the target object and use the corresponding mode data according to the fine-grained environment in which the target object is located, resulting in navigation accuracy being affected by interference such as light changes, occlusions and no GPS signals in the field, and the battery life time is reduced, making it difficult to navigate independently in complex environments for a long time.
A fully autonomous navigation method for drone with fine-grained environment perception capabilities is proposed. By carrying high-definition cameras, infrared cameras and lidars, multi-modal data sets are obtained, and image panoramic segmentation large models, multi-modal models and deep reinforcement learning algorithms are used to build a fully autonomous navigation agent of drone, which can independently identify target objects and use corresponding modal data for navigation according to different fine-grained environments.
It realizes fully autonomous navigation of drones with high accuracy, robustness and long-term operation in complex environments, enhancing the battery life of drones and the reliability of navigation results.
Smart Images

Figure CN119091329B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision unmanned aerial vehicle navigation methods, and in particular relates to a fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability. Background Art
[0002] Multimodal UAV autonomous navigation technology is an advanced technology that integrates multiple sensors and algorithms, enabling UAVs to achieve autonomous navigation in complex environments.
[0003] In recent years, the most researched method in the field of autonomous navigation of UAVs is the autonomous navigation method of UAVs based on the Deep Q-learning (Deep Q) algorithm. In the autonomous navigation task of UAVs, the Deep Q algorithm can deal with the local observability of the environment and the lack of perceptual information. Specifically, the researchers proposed a strategy network based on the memory enhancement mechanism, which can integrate historical memory information and current observations, extract the temporal dependencies of observation data, thereby enhancing the state estimation ability under local observable conditions and avoiding the algorithm from falling into the local optimal solution. However, the existing autonomous navigation methods of UAVs cannot autonomously find the target object and cannot use the corresponding modal data according to the fine-grained environment where the target object is located. Instead, the same modal data is used to drive the autonomous navigation of the UAV, resulting in the accuracy of the UAV navigation being affected by fine-grained interference limitations such as lighting changes, obstructions, and no GPS signals in the wild. At the same time, the endurance of the UAV may also be reduced, making it difficult for the UAV to navigate autonomously in a complex environment for a long time and effectively. Summary of the invention
[0004] In order to solve the technical problems existing in the background technology, the present invention aims to provide a fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability, and proposes a fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability, which can autonomously identify target objects and use corresponding modes according to fine-grained environments such as different lighting changes in the target object's environment, different obstructions, and no GPS in the wild. In particular, the method is used in fine-grained environments such as lighting changes in the target object's environment, frequent appearance of obstructions, and wild environments, and low battery power of the unmanned aerial vehicle. The method realizes a highly accurate, robust, and long-term operation of a fully autonomous navigation algorithm for an unmanned aerial vehicle by autonomously using corresponding modes in a complex environment.
[0005] In order to solve the technical problem, the technical solution of the present invention is:
[0006] A fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability, the method comprising:
[0007] S1: Use a drone equipped with a high-definition camera, infrared camera, and lidar to shoot the same target in different light intensities, obstructions, and fine-grained environments in the wild, and obtain a multimodal dataset including RGB images, infrared images, lidar signals, and global positioning system coordinates;
[0008] S2: Build a fully autonomous navigation agent for UAVs with a large image panoptic segmentation model SAM, a large multimodal model CLIP, a ViT backbone network with pre-trained public network weights, and a multi-layer perceptron;
[0009] S3: constructing training labels for the shortest path using the RGB image and the global positioning system coordinates, and preliminarily training the UAV fully autonomous navigation agent using the DeepQ-learning algorithm to obtain a UAV fully autonomous navigation agent that can autonomously search for target objects;
[0010] S4: constructing training labels using the multimodal data set, and training the above-mentioned drone fully autonomous navigation agent capable of autonomously searching for target objects through a prompt learning method to obtain a full-modal drone fully autonomous navigation agent capable of using full-modal data;
[0011] S5: Use the multimodal data set to construct different UAV fully autonomous navigation fine-grained environments, and use the DeepQ-learning algorithm to train the full-modal UAV fully autonomous navigation agent to obtain a UAV fully autonomous navigation agent that has the ability to perceive fine-grained environments and use corresponding modal data for navigation.
[0012] Further, the step S1 comprises:
[0013] S101: Use the same drone to carry a high-definition camera, an infrared camera, and a laser radar to respectively shoot different target objects in urban and outdoor scenes, and obtain RGB images, infrared images, and laser radar signals of the same object;
[0014] S102: obtaining global positioning GPS coordinates of the content in the video through a map;
[0015] S103: The RGB image, the infrared image, the laser radar signal and the global positioning system coordinates are matched one by one to obtain a multimodal data set including four modes of RGB, infrared image, laser radar signal and global positioning system coordinates.
[0016] Further, the step S2 comprises:
[0017] S201: Take the user instruction as the input of the existing large model CLIP to obtain the feature F of the user's intended target u ;
[0018] S202: Using the RGB image taken by the drone in real time as the input of the existing image panorama segmentation model SAM, the current scene of the drone can be analyzed and the object features F in the current scene can be obtained. s ;
[0019] S203: Using the feature F of the user's intended target u and the object features F in the current scene s The trained ViT backbone network and a multi-layer perceptron are used as input to determine whether the target in the current scene is the user's intended target. If it is the user's intended target, the global positioning GPS of the user's intended target is returned; if it is not the user's intended target, the global positioning GPS of the user's intended target is not returned;
[0020] The trained ViT backbone network is the network structure of the public Vision Transformer, and the network weights trained by GoogleNET are loaded; the multilayer perceptron consists of a fully connected layer with an input channel of 1024 and an output channel of 2048, a Relu activation function, and a fully connected layer with an input channel of 2048 and an output channel of 2.
[0021] Further, the step S3 comprises:
[0022] S301: Using the starting position of the drone as the path origin and the GPS of the user's intended target as the path end point, calibrate the optimal path in the above multimodal data set by manual clicking, and obtain the RGB image and GPS of the object drone at the optimal path point, and use them as the optimal path label; the optimal path is the path with the shortest flight time required for the drone to travel from the starting position to the destination;
[0023] S302: Taking the RGB image and the optimal path label in the multimodal data set as input, setting the reward function, and using the Deep Q-learning algorithm to train the above-mentioned UAV autonomous full navigation intelligent agent, according to the reward given by the reward function, the intelligent agent that obtains the maximum expected reward value is the UAV autonomous navigation intelligent agent of the RGB modality;
[0024] Further, the step S4 comprises:
[0025] S401: Based on the RGB image of the optimal path, the corresponding infrared image, lidar signal and GPS are used as training labels;
[0026] S402: Using the visual cue learning method, freeze the parameters of the RGB modality UAV autonomous navigation agent, add the corresponding infrared image token Token_T, the corresponding lidar signal token Token_Radia, and the corresponding GPS signal token Token_GPS to the ViT backbone network of the RGB modality UAV navigation agent, and initialize them to zero respectively;
[0027] S403: RGB modality UAV autonomous navigation agent is based on RGB images, infrared images, lidar signals and GPS signals in multimodal datasets, and uses transformer multi-head cross attention mechanism for global nonlinear fusion to obtain fusion feature vectors T at different levels. i , i=3,4,5;
[0028] S404: Calculate T using cross entropy loss i and T i ' loss, after the training converges, a full-modal UAV fully autonomous navigation agent capable of driving in all modes can be obtained.
[0029] Furthermore, the global nonlinear fusion based on the transformer multi-head cross attention mechanism includes:
[0030] S4031: Using multiple first single-layer fully connected networks l 1i () The multimodal fusion tokens T at different levels i Linearly map them into query vectors q i , that is, q i = l 1i (T i ), i=3,4,5;
[0031] S4032: Using multiple second single-layer fully connected networks l 2i () Different levels of multimodal image tokens X i Linearly map them into key vectors k i , that is, k i = l 2i (X i ), i=3,4,5;
[0032] S4033: Using multiple third single-layer fully connected networks l 3i () Different levels of multimodal image tokens X i Linearly map them into value vectors v i , that is, v i = l 3i (X i ), i=3,4,5;
[0033] S4034: Calculate the query vector qi and the key vector k i Perform sinusoidal spatial position embedding respectively to obtain the position vector q′ i and k′ i ;
[0034] S4035: The obtained value vector v i , position vector q′ i and k′ i , using the transformer-based multi-head cross attention mechanism model MultiHC i () performs global nonlinear fusion, and obtains the first fusion feature, namely, the feature T of each level. i , that is, T i =MultiHC i (q′ i ,k′ i ,v i ), i=3,4,5.
[0035] Further, the step S5 specifically includes:
[0036] S501: Based on the multimodal data set, respectively construct the unilluminated urban autonomous navigation fine-grained environment under the premise of low battery of the UAV, the urban autonomous navigation fine-grained environment with frequent occlusions, the field target without GPS fine-grained navigation environment, the unilluminated field target without GPS fine-grained autonomous navigation environment and the unilluminated autonomous urban navigation fine-grained environment under the premise of sufficient battery of the UAV, the urban autonomous fine-grained navigation environment with frequent occlusions, the field target without GPS fine-grained autonomous navigation environment and the unilluminated field target without GPS fine-grained autonomous navigation environment;
[0037] S502: Add four new fine-grained environment tokens Tokens_envirRGB, Tokens_envirRa, Tokens_envirT, and Tokens_envirGPS to the fully autonomous navigation agent of the fully modal UAV driven by the full modality, initialize them to zero, and fuse them with other tokens by multiplication, so as to transform the fully autonomous navigation agent of the full modality UAV into a fully autonomous navigation agent of the UAV with fine-grained environmental awareness;
[0038] S503: Set the enabled modal reward function, freeze the original parameters of the fully modal driven UAV autonomous navigation agent, use the RGB images, infrared images, lidar signals and GPS signals in the multimodal dataset as input, and use the Deep Q-learning method to train the multimodal UAV autonomous navigation agent with environmental awareness in the constructed fine-grained environment. The training is terminated when the maximum expected reward can be obtained, and the fully autonomous navigation of the UAV with fine-grained environmental perception capability is obtained.
[0039] Furthermore, the low battery condition of the drone is that the drone can only use one type of sensor, that is, it can only use a high-definition camera, an infrared camera, a laser radar, or a GPS;
[0040] The high power of the drone is based on the premise that the drone uses multiple sensors, namely, a high-definition camera, an infrared camera, a laser radar, and a GPS;
[0041] The method of constructing the unilluminated autonomous urban fine-grained navigation environment under the premise of low battery of the drone is to use the laser radar signal on the optimal path point as the training label;
[0042] The method of constructing the autonomous fine-grained urban navigation environment with frequent obstructions under the premise of low battery of the drone is to use the GPS on the optimal path point as the training label;
[0043] The method of constructing the field target non-GPS autonomous fine-grained navigation environment under the premise of low battery of the UAV is to use the RGB image on the optimal path point as the training label;
[0044] The method for constructing the autonomous fine-grained navigation environment of the unmanned aerial vehicle without GPS and without light in the field under the premise of low battery is to use the thermal image on the optimal path point as the training label.
[0045] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the above-mentioned methods for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability is implemented.
[0046] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned methods for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability.
[0047] Compared with the prior art, the advantages of the present invention are:
[0048] The present invention adopts a training method based on a combination of prompt learning and deep Q-learning driven by multimodal data to train a fully autonomous navigation method for a UAV with fine-grained navigation environment perception capability. The method can autonomously search for user intended targets and autonomously decide to use different modes of autonomous navigation according to different fine-grained navigation environments, so that the autonomous navigation has stronger endurance and the navigation results are more robust and accurate.
[0049] Different from the existing UAV autonomous navigation method based on deep Q-learning, the present invention adopts multimodal data-driven prompt learning and deep Q-learning; it drives the autonomous navigation of the UAV with multimodal data to provide data support for autonomous navigation, obtains a fully autonomous navigation agent based on RGB modal data, realizes the basic functions of fully autonomous navigation, adds new tokens to it, and uses prompt learning technology to use multimodal data to drive it to use multimodal data for autonomous navigation, and obtains a full-modal UAV fully autonomous navigation agent; further, basic multimodal data is used to build different fine-grained navigation environments, and deep Q-learning technology is used to realize fully autonomous navigation that perceives fine-grained navigation environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 , the main flow chart of a fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability according to the present invention. DETAILED DESCRIPTION
[0051] The specific implementation mode of the present invention is described below in conjunction with embodiments:
[0052] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to match the contents disclosed in the specification so that people familiar with this technology can understand and read them, and are not used to limit the conditions under which the present invention can be implemented. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.
[0053] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" cited in this specification are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.
[0054] Embodiment 1:
[0055] See attached Figure 1 A specific implementation of the fully autonomous navigation method of a UAV with fine-grained environmental perception capability proposed by the present invention includes the following steps:
[0056] S1: Use a drone equipped with a high-definition camera, infrared camera, and lidar to shoot the same target in different light intensities, obstructions, and fine-grained environments in the wild, and obtain a multimodal dataset containing RGB images, infrared images, radar signals, and global positioning GPS;
[0057] In some more specific embodiments, it may include:
[0058] S11: Use the same drone to carry a high-definition camera, infrared camera, and lidar to shoot different target objects in urban and outdoor scenes, and obtain RGB images, infrared images, and lidar signals of the same object;
[0059] S12: obtaining available global positioning (GPS) coordinates of the content in the above-mentioned video through a map;
[0060] S13: The RGB image, infrared image, lidar video, and global positioning system coordinates are matched one by one to obtain a multimodal data set including four modes: RGB, infrared image, lidar signal, and global positioning system coordinates.
[0061] S2: Build a fully autonomous navigation agent for UAVs using a large image panoptic segmentation model SAM, an existing large multimodal model CLIP, a ViT backbone network with pre-trained public network weights, and a multi-layer perceptron;
[0062] In some more specific embodiments, it may include:
[0063] S21: Taking the user instruction as the input of the existing large model CLIP, the feature Fu of the user's intended target can be obtained;
[0064] S22: Using the RGB image taken by the drone in real time as the input of the existing image panoramic segmentation model SAM, the current scene of the drone can be analyzed and the object features Fs in the current scene can be obtained;
[0065] S23: Using the feature Fu of the user's intended target and the feature Fs of the object in the current scene as the input of the existing trained ViT backbone network and a multi-layer perceptron, it can be determined whether the target in the current scene is the user's intended target. If it is the user's intended target, the global positioning GPS of the user's intended target is returned; if it is not the user's intended target, the global positioning GPS of the user's intended target is not returned.
[0066] According to some preferred embodiments of the present invention, the ViT feature extractor is the publicly available network structure of VisionTransformer (ViT), and loads the network weights trained by GoogleNET.
[0067] According to some preferred embodiments of the present invention, the multilayer perceptron is composed of a fully connected layer with an input channel of 1024 and an output channel of 2048, a Relu activation function, and a fully connected layer with an input channel of 2048 and an output channel of 2.
[0068] S3: Using RGB images and global positioning system coordinates as the shortest path to construct training labels, the Deep Q-learning algorithm is used to preliminarily train the above-mentioned UAV fully autonomous navigation agent, and a UAV fully autonomous navigation agent that can autonomously find the target object can be obtained;
[0069] In some more specific embodiments, it may include:
[0070] S31: Taking the starting position of the drone as the path origin and the GPS of the user's intended target as the path end point, calibrate the optimal path in the above multimodal dataset by manual clicking by multiple people, and obtain the RGB image and GPS of the object drone at the optimal path point, which are used as the shortest path label;
[0071] S32: Taking the RGB image and the optimal path label in the multimodal dataset as input, setting the reward function, and using the Deep Q-learning algorithm to train the above-mentioned UAV autonomous full navigation agent, the agent that obtains the maximum expected reward value based on the reward given by the reward function is the UAV autonomous navigation agent of the RGB modality.
[0072] According to some preferred embodiments of the present invention, the optimal path is a path that takes the least flight time for the drone to travel from a starting position to a destination.
[0073] According to some preferred embodiments of the present invention, the reward function is as follows:
[0074] 1. If the RGB image in the multimodal dataset is consistent with the optimal path and the output of the RGB modality UAV autonomous navigation agent is 1, the reward is +100;
[0075] 2. If the RGB image in the multimodal dataset is inconsistent with the optimal path and the output of the RGB modality UAV autonomous navigation agent is 0, the reward is +100;
[0076] 3. If the RGB image in the multimodal dataset is consistent with the optimal path, but the output of the RGB modality UAV autonomous navigation agent is 0, a penalty of -50 is imposed;
[0077] 4. If the RGB image in the multimodal dataset is inconsistent with the optimal path, but the output of the RGB modality UAV autonomous navigation agent is 1, a penalty of -50 is applied;
[0078] 5. Except for the above cases, the reward value is 0.
[0079] S4: Construct training labels with RGB images, infrared images, lidar and global positioning system coordinates, and use prompt learning technology to train the above-mentioned drone fully autonomous navigation agent that can autonomously find target objects, so as to obtain a full-modal drone full-navigation agent that can use full-modal data;
[0080] In some more specific embodiments, it may include:
[0081] S41: Based on the RGB image of the optimal path, the corresponding infrared image, lidar signal, and GPS are used as training labels;
[0082] S42: Using visual cue learning technology, the parameters of the RGB modality UAV autonomous navigation agent are frozen, and the corresponding infrared image token Token_T, the corresponding lidar signal token Token_Radia, and the corresponding GPS signal token Token_GPS are added to the ViT backbone network of the RGB modality UAV navigation agent, and they are initialized to zero respectively.
[0083] S43: RGB modality UAV autonomous navigation agent integrates RGB images, infrared images, lidar signals, and GPS signals in multimodal data sets, and performs global nonlinear fusion based on the transformer multi-head cross attention mechanism to obtain fusion feature vectors T at different levels i , i=3,4,5;
[0084] S44: RGB modality UAV autonomous navigation agent uses the RGB images, infrared images, lidar signals, and GPS signals in the above training labels, and performs global nonlinear fusion based on the transformer multi-head cross attention mechanism to obtain different levels of fusion feature vectors T i ', i = 3, 4, 5;
[0085] S45: Calculate T using cross entropy loss i and T i ' loss, after the training converges, a full-modal UAV fully autonomous navigation agent capable of driving in all modes can be obtained.
[0086] According to some preferred embodiments of the present invention, the global nonlinear fusion based on the transformer multi-head cross attention mechanism includes:
[0087] S441: Using multiple first single-layer fully connected networks l 1i () The multimodal fusion tokens T at different levels i Linearly map them into query vectors q i , that is, q i = l 1i (T i), i=3,4,5;
[0088] S442: Using multiple second single-layer fully connected networks l 2i () Different levels of multimodal image tokens X i Linearly map them into key vectors k i , that is, k i = l 2i (X i ), i=3,4,5;
[0089] S443: Using multiple third single-layer fully connected networks l 3i () Different levels of multimodal image tokens X i Linearly map them into value vectors v i , that is, v i = l 3i (X i ), i=3,4,5;
[0090] S444: The obtained query vector q i and the key vector k i Perform sinusoidal spatial position embedding respectively to obtain the position vector q′ i and k′ i ;
[0091] S445: The obtained value vector v i , position vector q′ i and k′ i , using the transformer-based multi-head cross attention mechanism model MultiHC i () performs global nonlinear fusion, and obtains the first fusion feature, namely, the feature T of each level. i , that is, T i =MultiHC i (q′ i ,k′ i ,v i ), i=3,4,5.
[0092] S5: Construct different UAV fully autonomous navigation fine-grained environments with different multimodal data, and use the Deep Q-learning algorithm to train the above-mentioned full-modal UAV fully autonomous navigation agent again to obtain a UAV fully autonomous navigation agent that has the ability to perceive fine-grained environments and use corresponding modal data for navigation.
[0093] In some more specific embodiments, it includes:
[0094] S51: Based on the multimodal dataset, we construct the following fine-grained urban autonomous navigation environment without lighting under the premise of low battery of UAV, the fine-grained urban autonomous navigation environment with frequent occlusions, the fine-grained navigation environment of outdoor targets without GPS, the fine-grained autonomous navigation environment of outdoor targets without GPS, and the fine-grained autonomous urban navigation environment without lighting under the premise of sufficient battery of UAV, the fine-grained urban autonomous navigation environment with frequent occlusions, the fine-grained autonomous navigation environment of outdoor targets without GPS, and the fine-grained autonomous navigation environment of outdoor targets without GPS;
[0095] S52: Add four new fine-grained environment tokens Tokens_envirRGB, Tokens_envirRa, Tokens_envirT, and Tokens_envirGPS to the full-modal UAV fully autonomous navigation agent driven by the full modality, initialize them to zero, and fuse them with other tokens by multiplication, so as to transform the full-modal UAV fully autonomous navigation agent into a UAV fully autonomous navigation agent with fine-grained environmental awareness;
[0096] S53: Set the enabled modal reward function, freeze the original parameters of the fully modal driven UAV autonomous navigation agent, take the RGB image, infrared image, lidar signal, and GPS signal in the multimodal dataset as input, and use the deep q-learning method to train the above multimodal UAV autonomous navigation agent with environmental awareness in the above constructed fine-grained environment. End the training when the maximum expected reward can be obtained, and a fully autonomous UAV navigation agent with fine-grained environmental perception capability can be obtained.
[0097] According to some preferred embodiments of the present invention, the low battery condition of the drone is that the drone can only use one type of sensor, that is, it can only use a high-definition camera or an infrared camera or a laser radar or a GPS.
[0098] According to some preferred embodiments of the present invention, the high power of the drone is based on the premise that the drone can use a variety of sensors, that is, it can use a high-definition camera, an infrared camera, a laser radar, and a GPS.
[0099] According to some preferred embodiments of the present invention, the method for constructing the unilluminated autonomous urban fine-grained navigation environment under the premise of low battery of the drone is to use the laser radar signal on the optimal path point as the training label.
[0100] According to some preferred embodiments of the present invention, the method for constructing an urban autonomous fine-grained navigation environment with frequent obstructions when the drone is low on battery is to use the GPS on the optimal path point as a training label.
[0101] According to some preferred embodiments of the present invention, the method for constructing an autonomous fine-grained navigation environment for field targets without GPS under the premise of low battery of the drone is to use RGB images on optimal path points as training labels.
[0102] According to some preferred embodiments of the present invention, the method of constructing an autonomous fine-grained navigation environment for a lightless outdoor target without GPS under the premise of low battery of the drone is to use thermal images on the optimal path points as training labels.
[0103] According to some preferred embodiments of the present invention, the enabling modality reward function is as follows:
[0104] 1. If the drone is in a low-battery state and there is no illumination in the autonomous urban fine-grained navigation environment, the environmental perception token will fuse all tokens except the token Token_Radia corresponding to the lidar signal to zero, and the reward will be +100;
[0105] 2. If the obstructions frequently appear in the urban autonomous fine-grained navigation environment under the premise of low battery of the drone, the environmental perception token will merge all tokens except the token Token_GPS corresponding to the GPS signal to zero, and the reward will be +100;
[0106] 3. If the drone is low on battery and there is no GPS autonomous fine-grained navigation environment for the outdoor target, and the environmental perception token returns all tokens except the RGB parameters to zero, the reward is +100;
[0107] 4. If the drone is in a low-battery autonomous fine-grained navigation environment with no light and no GPS, and the environmental perception token returns all tokens except the token Token_T corresponding to the infrared image to zero, the reward is +100;
[0108] 5. If the drone is able to freely use different modal data to navigate to the target with sufficient power, the reward will be +100;
[0109] 6. Except for the above cases, the reward is 0.
[0110] Embodiment 2:
[0111] This embodiment provides a terminal device, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of a fully autonomous navigation method for a drone with fine-grained environmental perception capability, including the following steps:
[0112] S1: Use a drone equipped with a high-definition camera, infrared camera, and lidar to shoot the same target in different light intensities, obstructions, and fine-grained environments in the wild, and obtain a multimodal dataset including RGB images, infrared images, lidar signals, and global positioning system coordinates;
[0113] S2: Build a fully autonomous navigation agent for UAVs with a large image panoptic segmentation model SAM, a large multimodal model CLIP, a ViT backbone network with pre-trained public network weights, and a multi-layer perceptron;
[0114] S3: constructing training labels for the shortest path using the RGB image and the global positioning system coordinates, and preliminarily training the UAV fully autonomous navigation agent using the DeepQ-learning algorithm to obtain a UAV fully autonomous navigation agent that can autonomously search for target objects;
[0115] S4: constructing training labels using the multimodal data set, and training the above-mentioned drone fully autonomous navigation agent capable of autonomously searching for target objects through a prompt learning method to obtain a full-modal drone fully autonomous navigation agent capable of using full-modal data;
[0116] S5: Use the multimodal data set to construct different UAV fully autonomous navigation fine-grained environments, and use the DeepQ-learning algorithm to train the full-modal UAV fully autonomous navigation agent to obtain a UAV fully autonomous navigation agent that has the ability to perceive fine-grained environments and use corresponding modal data for navigation.
[0117] Embodiment 3:
[0118] This embodiment provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.
[0119] The processor may load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the above-mentioned embodiment related to a fully autonomous navigation method of a drone with fine-grained environmental perception capability; the processor may load and execute the following steps of one or more instructions in the computer-readable storage medium:
[0120] S1: Use a drone equipped with a high-definition camera, infrared camera, and lidar to shoot the same target in different light intensities, obstructions, and fine-grained environments in the wild, and obtain a multimodal dataset including RGB images, infrared images, lidar signals, and global positioning system coordinates;
[0121] S2: Build a fully autonomous navigation agent for UAVs with a large image panoptic segmentation model SAM, a large multimodal model CLIP, a ViT backbone network with pre-trained public network weights, and a multi-layer perceptron;
[0122] S3: constructing training labels for the shortest path using the RGB image and the global positioning system coordinates, and preliminarily training the UAV fully autonomous navigation agent using the DeepQ-learning algorithm to obtain a UAV fully autonomous navigation agent that can autonomously search for target objects;
[0123] S4: constructing training labels using the multimodal data set, and training the above-mentioned drone fully autonomous navigation agent capable of autonomously searching for target objects through a prompt learning method to obtain a full-modal drone fully autonomous navigation agent capable of using full-modal data;
[0124] S5: Use the multimodal data set to construct different UAV fully autonomous navigation fine-grained environments, and use the DeepQ-learning algorithm to train the full-modal UAV fully autonomous navigation agent to obtain a UAV fully autonomous navigation agent that has the ability to perceive fine-grained environments and use corresponding modal data for navigation.
[0125] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0127] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0129] The present invention discloses a fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability, which comprises: constructing a multimodal autonomous navigation data set comprising RGB images, infrared images, laser radars carried by the unmanned aerial vehicle, and global positioning, which are taken by the unmanned aerial vehicle under conditions of different light intensities and obstructions; using the infrared image, RGB image, laser radar carried by the unmanned aerial vehicle, and global positioning system of the destination as the shortest path label, and using the deep Q-learning algorithm to preliminarily train the fully autonomous navigation intelligent agent of the unmanned aerial vehicle; using the infrared image, RGB image, laser radar, and global positioning system of the destination as the shortest path label, and using prompt learning to train the above-mentioned fully autonomous navigation intelligent agent of the unmanned aerial vehicle again, so as to obtain a full-modal fully autonomous navigation intelligent agent of the unmanned aerial vehicle that can use the infrared image, RGB image, laser radar signal, and global positioning simultaneously; respectively constructing a multimodal autonomous navigation data set comprising RGB images, infrared images, laser radars carried by the unmanned aerial vehicle, and global positioning, which are taken by the unmanned aerial vehicle in a non-lighting autonomous environment, an obstruction environment, and a field environment, and using the infrared image or RGB image or laser radar or global positioning system of the destination as the shortest path label, and using the deep The Q-learning algorithm trains the above-mentioned full-modal UAV fully autonomous navigation agent again to obtain a UAV navigation agent that can perceive fine-grained environments and correctly use the correct modality in different fine-grained environments.
[0130] The preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
[0131] Many other changes and modifications may be made without departing from the concept and scope of the present invention.It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.
Claims
1. A fully autonomous navigation method for an unmanned aerial vehicle with fine-grained environmental perception capability, characterized in that: The method comprises: S1: Use a drone equipped with a high-definition camera, infrared camera, and lidar to shoot the same target in different light intensities, obstructions, and fine-grained environments in the wild, and obtain a multimodal dataset including RGB images, infrared images, lidar signals, and global positioning system coordinates; S2: Build a fully autonomous navigation agent for UAVs with a large image panoptic segmentation model SAM, a large multimodal model CLIP, a ViT backbone network with pre-trained public network weights, and a multi-layer perceptron; S3: constructing training labels for the optimal path using the RGB image and the global positioning system coordinates, and preliminarily training the UAV fully autonomous navigation agent using the Deep Q-learning algorithm to obtain a UAV fully autonomous navigation agent that can autonomously search for target objects; S4: constructing training labels using the multimodal data set, and training the above-mentioned drone fully autonomous navigation agent capable of autonomously searching for target objects through a prompt learning method to obtain a full-modal drone fully autonomous navigation agent capable of using full-modal data; S5: Using the multimodal data set to construct different UAV fully autonomous navigation fine-grained environments, using the Deep Q-learning algorithm to train the multimodal UAV fully autonomous navigation intelligent agent, and obtaining a UAV fully autonomous navigation intelligent agent capable of perceiving fine-grained environments and navigating using corresponding modal data; including: S501: Based on the multimodal data set, respectively construct the unilluminated urban autonomous navigation fine-grained environment under the premise of low battery of the UAV, the urban autonomous navigation fine-grained environment with frequent occlusions, the field target without GPS fine-grained navigation environment, the unilluminated field target without GPS fine-grained autonomous navigation environment and the unilluminated autonomous urban navigation fine-grained environment under the premise of sufficient battery of the UAV, the urban autonomous fine-grained navigation environment with frequent occlusions, the field target without GPS fine-grained autonomous navigation environment and the unilluminated field target without GPS fine-grained autonomous navigation environment; S502: Add four new fine-grained environmental tokens Tokens_envirRGB, Tokens_envirRa, Tokens_envirT, and Tokens_envirGPS to the fully-modal UAV fully-autonomous navigation agent driven by the full modality, initialize them to zero, and fuse them with other Tokens by multiplication to transform the full-modal UAV fully-autonomous navigation agent into a UAV fully-autonomous navigation agent with fine-grained environmental awareness. The other Tokens include: infrared image token Token_T, lidar signal token Token_Radia, and GPS signal token Token_GPS.
2. The method for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability according to claim 1 is characterized in that: The step S1 comprises: S101: Use the same drone to carry a high-definition camera, an infrared camera, and a laser radar to respectively shoot different target objects in urban and outdoor scenes, and obtain RGB images, infrared images, and laser radar signals of the same object; S102: obtaining global positioning GPS coordinates of the content in the above-mentioned video through a map; S103: The global positioning GPS coordinates are matched one by one with the RGB image, the infrared image, and the laser radar signal to obtain a multimodal data set including four modes of RGB, infrared image, laser radar signal and global positioning system coordinates.
3. The method for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability according to claim 1 is characterized in that: The step S2 comprises: S201: Take the user instruction as the input of the existing large model CLIP to obtain the feature F of the user's intended target u ; S202: Using the RGB image taken by the drone in real time as the input of the existing image panorama segmentation model SAM, the current scene of the drone can be analyzed and the object features F in the current scene can be obtained. s ; S203: Using the feature F of the user's intended target u and the object features F in the current scene s The trained ViT backbone network and a multi-layer perceptron are used as input to determine whether the target in the current scene is the user's intended target. If it is the user's intended target, the global positioning GPS of the user's intended target is returned; if it is not the user's intended target, the global positioning GPS of the user's intended target is not returned; The trained ViT backbone network is the network structure of the public Vision Transformer, and the network weights trained by GoogleNET are loaded; the multilayer perceptron consists of a fully connected layer with an input channel of 1024 and an output channel of 2048, a Relu activation function, and a fully connected layer with an input channel of 2048 and an output channel of 2.
4. The method for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability according to claim 1 is characterized in that: The step S3 comprises: S301: Using the starting position of the drone as the path origin and the GPS of the user's intended target as the path end point, calibrate the optimal path in the above multimodal data set by manual clicking, and obtain the RGB image and GPS of the object drone at the optimal path point, and use them as the optimal path label; the optimal path is the path with the shortest flight time required for the drone to travel from the starting position to the destination; S302: Taking the RGB image and the optimal path label in the multimodal dataset as input, setting the reward function, and using the DeepQ-learning algorithm to train the above-mentioned UAV autonomous full navigation agent, the agent that obtains the maximum expected reward value based on the reward given by the reward function is the UAV autonomous navigation agent in the RGB modality.
5. The method for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability according to claim 1 is characterized in that: The step S4 comprises: S401: Based on the RGB image of the optimal path, the corresponding infrared image, lidar signal and GPS are used as training labels; S402: Using the visual cue learning method, freeze the parameters of the RGB modality UAV autonomous navigation agent, add the corresponding infrared image token Token_T, the corresponding lidar signal token Token_Radia, and the corresponding GPS signal token Token_GPS to the ViT backbone network of the RGB modality UAV navigation agent, and initialize them to zero respectively; S403: RGB modality UAV autonomous navigation agent collects RGB images, infrared images, lidar signals, and GPS signals in multimodal data sets, and performs global nonlinear fusion based on the transformer multi-head cross attention mechanism to obtain fusion feature vectors T at different levels. i , i=3,4,5; S404: The RGB modality UAV autonomous navigation agent uses the RGB images, infrared images, lidar signals, and GPS signals in the above training labels and performs global nonlinear fusion based on the transformer multi-head cross attention mechanism to obtain different levels of fusion feature vectors T i ', i = 3, 4, 5; S405: Calculate T using cross entropy loss i and T i ' loss, after the training converges, a full-modal UAV fully autonomous navigation agent capable of driving in all modes can be obtained.
6. The method for fully autonomous navigation of a UAV with fine-grained environmental perception capability according to claim 5 is characterized in that: The global nonlinear fusion based on the transformer multi-head cross attention mechanism includes: S4031: Using multiple first single-layer fully connected networks l 1i () The multimodal fusion tokens T at different levels i Linearly map them into query vectors q i , that is, q i = l 1i (T i ), i=3,4,5; S4032: Using multiple second single-layer fully connected networks l 2i () Different levels of multimodal image tokens X i Linearly map them into key vectors k i , that is, k i = l 2i (X i ), i=3,4,5; S4033: Using multiple third single-layer fully connected networks l 3i () Different levels of multimodal image tokens X i Linearly map them into value vectors v i , that is, v i = l 3i (X i ), i=3,4,5; S4034: Calculate the query vector q i and the key vector k i Perform sinusoidal spatial position embedding respectively to obtain the position vector q′ i and k′ i ; S4035: The obtained value vector v i , position vector q′ i and k′ i , using the transformer-based multi-head cross attention mechanism model Mult i HC i () performs global nonlinear fusion, and obtains the first fusion feature, namely, the feature T of each level. i , that is, T i =MultiHC i (q′ i ,k′ i ,v i ), i=3,4,5.
7. The method for fully autonomous navigation of an unmanned aerial vehicle with fine-grained environmental perception capability according to claim 1, characterized in that: The step S5 further includes: S503: Set the enabled modal reward function, freeze the original parameters of the fully modal driven UAV autonomous navigation agent, use the RGB images, infrared images, lidar signals and GPS signals in the multimodal dataset as input, and use the Deep Q-learning method to train the multimodal UAV autonomous navigation agent with environmental awareness in the constructed fine-grained environment. The training is terminated when the maximum expected reward can be obtained, and the fully autonomous navigation of the UAV with fine-grained environmental perception capability is obtained.
8. The method for fully autonomous navigation of a UAV with fine-grained environmental perception capability according to claim 7 is characterized in that: The low battery condition of the drone is that the drone can only use one sensor, that is, only a high-definition camera, an infrared camera, a laser radar or a GPS; The high power of the drone is based on the premise that the drone uses multiple sensors, namely, a high-definition camera, an infrared camera, a laser radar, and a GPS; The method of constructing the unilluminated autonomous urban fine-grained navigation environment under the premise of low battery of the drone is to use the laser radar signal on the optimal path point as the training label; The method of constructing the autonomous fine-grained urban navigation environment with frequent obstructions under the premise of low battery of the drone is to use the GPS on the optimal path point as the training label; The method of constructing the field target non-GPS autonomous fine-grained navigation environment under the premise of low battery of the UAV is to use the RGB image on the optimal path point as the training label; The method for constructing the autonomous fine-grained navigation environment of the unmanned aerial vehicle without GPS and without light in the field under the premise of low battery is to use the thermal image on the optimal path point as the training label.
9. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a fully autonomous navigation method for a drone with fine-grained environmental perception capability as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements a fully autonomous navigation method for a drone with fine-grained environmental perception capability as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Indoor navigation method based on vision and radar information fusion and reinforcement learning
CN116263335A
Multi-modal fusion method for heterogeneous data of intelligent networked vehicle multi-source sensor
CN118445748A