Visual autonomous navigation method for unmanned aerial vehicle
Through cross-modal comparison learning and near-end strategy optimization algorithms, combined with RGB and deep image information, the problem of poor navigation of drones in complex environments is solved, and efficient and robust autonomous navigation capabilities are achieved.
Patent Information
- Application Number
- CN202510282213.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-10
AI Technical Summary
Existing autonomous navigation methods for drone are poorly performed in complex or perceptually restricted environments, making it difficult to maintain robustness and efficiency, especially under conditions such as light changes and texture loss.
By aligning the information of the RGB image with the information of the depth image, a cross-modal comparison learning model is constructed, a modal-independent feature representation is generated using the InfoNCE loss function, and a visual motion strategy is trained in combination with a proximity strategy optimization algorithm.
It significantly improves the adaptability and robustness of the drone in complex environments, achieves efficient and stable autonomous navigation, and has high migration capabilities.
Smart Images

Figure CN120121054A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of artificial intelligence and robotics technology, and in particular relates to a visual autonomous navigation method for an unmanned aerial vehicle. Background Art
[0002] With the rapid development of UAV technology, its application prospects in environmental monitoring, logistics transportation, disaster relief and other fields are becoming increasingly broad. In UAV systems, autonomous navigation is one of the core issues of autonomous flight of UAVs. The core challenge of realizing autonomous navigation of UAVs is how to achieve efficient and stable obstacle avoidance and target navigation in complex and unknown environments. Traditional UAV navigation methods mostly rely on simultaneous localization and mapping and motion structure recovery technologies. These methods achieve three-dimensional reconstruction and pose estimation by extracting and matching feature points. Although these methods have shown certain feasibility in the field of autonomous flight of UAVs, they require high-intensity feature matching and optimization calculations, rely heavily on accurate data input from sensors, and are difficult to maintain robustness under interference conditions such as illumination changes and texture loss, resulting in poor performance in complex or perception-restricted environments. In contrast, human pilots rely only on images captured by UAV cameras to achieve precise control of UAVs through intuitive perception and decision-making. This capability has inspired the study of vision-based navigation strategies, which attempt to replace complex state estimation with end-to-end mapping of perception and control.
[0003] In recent years, deep reinforcement learning has gradually become a research hotspot in the field of autonomous navigation of UAVs. It demonstrates high adaptability by learning control strategies through interaction with the environment. However, end-to-end autonomous navigation methods based on deep reinforcement learning still face problems such as large data demand, high training difficulty, and poor cross-scenario migration. Among them, since the strategy optimization process requires a large number of state-action pair samples, the sampling efficiency is low, and the coupling of perceptual feature extraction and action decision-making further increases the complexity of training. At the same time, the difference in visual features in unseen environments makes the migration and robustness of traditional models insufficient.
[0004] In order to meet the above challenges, some studies in recent years have begun to try to improve the stability of visual features through contrastive learning. For example, by contrastive learning of RGB images with different data enhancements, task-related features are extracted to enhance visual robustness under changes in lighting and texture. However, these methods are usually limited to a single modality and lack in-depth use of multimodal data, resulting in their adaptability in complex environments still needs to be improved. Summary of the invention
[0005] In view of this, the present invention aims to overcome the deficiencies of the above problems in the prior art, and proposes a method for autonomous visual navigation of drones. By aligning the information of RGB images with the information of depth images, the present invention can not only retain task-related features, but also significantly improve the robustness and generalization ability of feature representation, providing an efficient and stable solution for the autonomous navigation of drones in complex environments.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows:
[0007] The first aspect of the present invention provides a method for autonomous visual navigation of drones, including the following steps:
[0008] Step 1: Model the interaction between the drone and the environment as a partially observable Markov decision process. The goal of this decision process is to construct an optimal visual motion strategy that enables the drone to select actions to maximize the expected value of the cumulative discounted reward.
[0009] Step 2: Construct a cross-modal contrastive learning model, and the objective function is defined as:
[0010]
[0011] where I(Z; X RGB ) and I(Z; X Depth ) respectively represent the mutual information between the learned feature representation Z and the RGB and depth inputs.
[0012] Step 3: Use the proximal policy optimization algorithm based on the actor-critic structure to train the visual motion strategy.
[0013] Furthermore, in the above Step 1, the decision process is represented by a six-tuple (S, A, p, R, O, γ), where S represents the state space of the environment, A represents the action space of the agent, P is the state transition model, defined as p: S × A × S → [0, 1], indicating the probability that the current state s t transfers to the next state s t after taking the action a t+1 , R: is the reward function, responsible for assigning real-valued rewards to the actions performed by the agent, O represents the observation space, and γ ∈ (0, 1) is the discount factor, used to balance the weights of immediate rewards and future rewards;
[0014] The goal of the decision process is to construct an optimal visual motion strategy π * : O → A, enabling the drone to select actions to maximize the expected value of the cumulative discounted reward, and its objective expression is:
[0015] Furthermore, in step 2, the InfoNCE loss is used as a lower bound estimate of the mutual information, and its definition is as follows:
[0016]
[0017] where τ is the temperature parameter, and f sim (z RGB , Z Depth ) represents the similarity function between the feature representations extracted by the query q RGB and the corresponding key k Depth .
[0018] Furthermore, in step 2, the training process of the cross-modal contrast learning model is as follows:
[0019] In each sampling iteration, a batch of original RGB images and their corresponding depth maps are randomly selected to form positive and negative sample pairs for contrast learning;
[0020] Data augmentation operations are performed;
[0021] The feature extraction network adopts the ResNet-50 structure;
[0022] After obtaining the image embeddings, the gradients are calculated according to the loss function, and the model parameters are updated by the stochastic gradient descent method;
[0023] After the training is completed, the parameters of the feature extraction network are frozen.
[0024] Furthermore, in step 3, the actor network maximizes the expected reward by generating actions, while constraining the policy update amplitude to ensure the stability of training. Its loss function is defined as follows:
[0025]
[0026] where r t (θ) represents the probability ratio of the policy before and after the update, is the advantage estimate value, and ∈ is the clipping parameter used to limit large updates;
[0027] The critic network reduces the variance of the policy update by estimating the cumulative reward. Its loss function is defined as:
[0028]
[0029] where V(s t ) represents the predicted state value, and V target is the true state value, which is calculated as the sum of the current reward and the discounted future reward.
[0030] Further, in step 3, the policy network directly maps the input of the visual motion policy to the original control instruction value, and the control instruction value is scaled by a coefficient α and truncated for the part exceeding the threshold to generate the finally executed speed.
[0031] The second aspect of the present invention provides a drone visual autonomous navigation device, which specifically includes:
[0032] A modeling unit for modeling the interaction between the drone and the environment as a partially observable Markov decision process, and the goal of this decision process is to construct an optimal visual motion policy so that the drone can select actions to maximize the expected value of the cumulative discounted reward;
[0033] A learning unit for constructing a cross-modal contrastive learning model, and the objective function is defined as:
[0034]
[0035] where I(Z; X RGB ) and I(Z; X Depth ) respectively represent the mutual information between the learned feature representation Z and the RGB and depth inputs;
[0036] A training unit for training the visual motion policy by using the proximal policy optimization algorithm based on the actor-critic structure.
[0037] The third aspect of the present invention provides an electronic device, including a processor and a memory communicatively connected to the processor and used for storing executable instructions of the processor, and the processor is used to execute the above-mentioned drone visual autonomous navigation method.
[0038] The fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the above-mentioned drone visual autonomous navigation method.
[0039] Compared with the prior art, the present invention has the following advantages:
[0040] Through cross-modal contrastive learning technology, the present invention uses the joint feature learning of RGB images and depth images to generate a visual representation with robustness and task relevance, greatly improving the adaptability of the drone in complex environments.
[0041] The present invention uses the proximal policy optimization (PPO) algorithm for visual motion policy training, combines multi-frame visual features and drone state information, realizes the end-to-end mapping from perception to control, and shows excellent learning efficiency and stability in the high-dimensional state space.
[0042] In simulation experiments and physical experiments, the present invention demonstrates strong robustness against illumination changes, color interference, and texture differences, and can maintain a high success rate of tasks in complex and unknown environments.
[0043] By fine-tuning the feature extraction network, the present invention effectively reduces the gap between simulation and reality, achieves high navigation performance without additional fine-tuning steps, and verifies the transferability of the strategy.
[0044] The method of the present invention is not only applicable to the autonomous navigation task of drones, but also can be extended to other perception and control fields that require high robustness and transfer ability, and has high theoretical and application value. Brief Description of the Drawings
[0045] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0046] Figure 1 is a schematic diagram of the cross-modal visual motion strategy framework of the present invention;
[0047] Figure 2 is a schematic diagram of the training and testing environment of the present invention;
[0048] Figure 3 is a schematic diagram of the trajectories of various benchmark methods in the testing environment of the present invention;
[0049] Figure 4 is a schematic diagram of the attention visualization of the model of the present invention in the simulation environment;
[0050] Figure 5 is a schematic diagram of the attention visualization of the model of the present invention in the actual environment;
[0051] Figure 6 is a schematic diagram of the drone trajectory in the indoor physical experiment of the present invention. Detailed Description of the Invention
[0052] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0053] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0054] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0055] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0056] Embodiment 1:
[0057] The present invention provides a method for visual autonomous navigation of an unmanned aerial vehicle (UAV). The complete structural framework is as shown in the attached Figure 1 drawings. In the training stage, RGB images and depth images are collected simultaneously for training a contrastive learning network. This network generates modality-independent feature representations through a similarity loss function, enabling the features extracted from the RGB input to have the same expressive ability as the depth images. After training is completed, the network parameters are frozen, and consecutive RGB frames are input into the network to extract features. These features are concatenated with the state observations and then input into the policy network, finally generating consecutive velocity commands for the UAV. Specifically, the present invention includes the following steps:
[0058] Step 1: Model the interaction between the UAV and the environment as a partially observable Markov decision process. The goal of this decision process is to construct an optimal visual motion policy that enables the UAV to select actions to maximize the expected value of the cumulative discounted reward.
[0059] Step 2: Construct a cross-modal contrastive learning model, and the objective function is defined as:
[0060]
[0061] Among them, I(Z; X RGB ) and I(Z; X Depth ) respectively represent the mutual information between the learned feature representation Z and the RGB and depth inputs;
[0062] Step 3: Use the proximal policy optimization algorithm based on the actor-critic structure to train the visual motion policy.
[0063] The problem studied in this invention is the autonomous navigation problem of an unmanned aerial vehicle (UAV) relying only on an on-board monocular camera. Considering the problem of the limited perception field of the on-board monocular camera, this invention models the interaction between the UAV and the environment as a partially observable Markov decision process, which is represented by a six-tuple (S, A, P, R, O, γ). Among them, S represents the state space of the environment, A represents the action space of the agent, P is the state transition model, defined as P: S × A × S →
[0064] [0, 1], representing the probability that the current state s t transitions to the next state s t after taking the action a t+1 . R: is the reward function, responsible for assigning a real-valued reward to the action executed by the agent. O represents the observation space, and γ ∈ (0, 1) is the discount factor, used to balance the weights of immediate rewards and future rewards. Specifically, at each time step t, the agent receives an observation value o t ∈ O from the current state s t ∈ S, selects an action a t ∈ A based on this observation value, and transfers to a new state s t+1 ∣ s t , a t ) according to the state transition model P(s t+1 , while obtaining a reward r t ∈ R. The goal of this decision process is to construct an optimal visual motion policy π * : O → A, enabling the UAV to select actions to maximize the expected value of the cumulative discounted reward, and its objective expression is: During the actual execution process, the UAV uses the optimized policy
[0065]
[0066] to iteratively select actions until the task goal is completed or a predetermined termination condition is met.
[0067] Among them, represents the mathematical expectation, 's subscript indicates starting from Sampling of s and a in the distribution.
[0068] In the present invention, the interaction between the unmanned aerial vehicle and the environment is modeled as a partially observable Markov decision process, which includes the system state, and the image information is a subset of the state. The present invention processes the image information through cross-modal contrast learning to enhance the learning effect, and the processed image information is used for visual motion training.
[0069] Traditional contrast learning methods form positive sample pairs by performing different data augmentations on the same RGB input and mapping it into a shared feature space; while dissimilar samples form negative sample pairs. Under this framework, contrast learning is regarded as a retrieval task, where each query needs to find its correct match from a set of candidate samples. Based on the traditional contrast learning principle, the present invention proposes a cross-modal contrast learning method that includes RGB images and depth images to enable RGB images to generate visual feature representations as robust as depth images. In this method, the query q represents an RGB image of size M×M, and the goal is to find its corresponding positive sample k 0 , k 1 , …, k N} from the candidate set K = {k + , that is, the corresponding M×M depth map k Depth . The query q and its positive sample k + have the same dimension to ensure a consistent feature space representation. The positive sample pair is represented by (q, k + ), while the negative sample pair is composed of (q, K\{k +}).
[0070] To further optimize the learned feature space, the present invention adopts a joint information bottleneck framework, retaining only the task-related shared information in the RGB and depth modalities. Its objective function is defined as:
[0071]
[0072] where I(Z; X RGB ) and I(Z; X Depth ) respectively represent the mutual information between the learned feature representation Z and the RGB and depth inputs. This objective encourages the model to eliminate redundant or modality-specific noise and focus on task-related information.
[0073] Directly calculating I(Z; X RGB ) and I(Z; X Depth ) is challenging because it requires approximating the high-dimensional joint distribution p(z,x) and the marginal distribution p(z). To address this complexity issue, the present invention adopts the InfoNCE loss as a lower bound estimate of the mutual information, which is defined as: L
[0074]
[0075] Among them, τ is the temperature parameter, D represents the distribution of a mini batch of data, denotes the expectation of the following expression over this distribution, and f sim (z RGB , z Depth ) represents the similarity function between the feature representations extracted from the query q RGB and the corresponding key k Depth . The present invention adopts dot product similarity as the similarity metric, which is defined as:
[0076]
[0077] This objective function encourages the feature representation z RGB extracted from the RGB image q RGB to align closely with the feature representation z Depth extracted from its corresponding depth map k Depth , while being far from the feature representations of other samples in the batch. In addition, this loss function can be interpreted as a multi-class cross-entropy objective, where the model acts as a classifier to identify the correct positive samples from a set of unrelated negative samples, thus separating the true matches from the interfering samples.
[0078] The training process of this cross-modal contrastive learning method is as follows: In each sampling iteration, a batch of original RGB images and their corresponding depth maps are randomly selected to form positive and negative sample pairs for contrastive learning. Subsequently, data augmentation operations are applied, including random cropping, resizing, random flipping, and Gaussian blur. The feature extraction network adopts the ResNet-50 structure. After obtaining the image embeddings, the gradients are calculated according to the above loss function, and the model parameters are updated by the stochastic gradient descent method. After training is completed, the parameters of the feature extraction network are frozen so that the learned representations can be effectively applied to downstream tasks.
[0079] As the core of the downstream visual navigation task, the visual motion strategy of the present invention generates real-time control instructions by interpreting perceptual information. The monocular RGB image is processed by the previously frozen feature extraction network to generate a feature representation and input it into the policy network to guide the actions of the drone. The present invention adopts the proximal policy optimization algorithm based on the actor-critic structure to train the visual motion strategy. The actor network maximizes the expected reward by generating actions, while constraining the policy update amplitude to ensure the stability of training. Its loss function is defined as follows:
[0080]
[0081] Among them, r t$(\theta)$ represents the probability ratio of the policy before and after the update. is the advantage estimate, and $\epsilon$ is the clipping parameter used to limit large updates.
[0082] The critic network aims to reduce the variance of policy updates by estimating the cumulative reward, and its loss function is defined as:
[0083]
[0084] where $V(s$ t ) represents the predicted state value, and $V$ target is the true state value, which is calculated as the sum of the current reward and the discounted future rewards.
[0085] The input of the visual - motor policy includes the feature representations $Z$ t-2 , $Z$ t-1 , $Z$ t of three consecutive RGB images; the internal state $s$ t of the drone, including its attitude, linear velocity, and angular velocity; and the relative position $P$ t of the drone with respect to the target. The policy network directly maps these inputs to the original control instruction values, i.e., the linear velocity $V$ t and the angular velocity $\Omega$ t . Subsequently, these values are scaled by the coefficient $\alpha$, and the parts exceeding the thresholds $V$ max and $\Omega$ max are truncated to generate the final executed velocity. Among them, each training episode terminates in the following three cases:
[0086] 1. The drone reaches the target point;
[0087] 2. The drone collides with an obstacle;
[0088] 3. The number of steps the drone runs exceeds the maximum allowed number of steps $T$;
[0089] The reward function of deep reinforcement learning is set as follows: The drone gets a reward of +10 when it reaches the target and is penalized -1 when it collides. In addition, a small penalty is incurred for each action to encourage the drone to reach the target point with as few steps as possible. At the same time, the function gives an additional reward according to the forward distance $d$, defined as $0.1\times(d$ init - $d$ current ), where $d$ init represents the Euclidean distance between the starting point and the target point, and $d$ current represents the Euclidean distance between the current position of the drone and the target point.
[0090] The present invention has conducted simulation experiments to verify the effectiveness of the method. The simulation implementation of the present invention is built on the high-fidelity simulation platform Unity, which is widely used in the field of robotics for its realistic rendering effects, customizable physics engine, and support for complex agent interactions. The present invention designs two scenarios for experimental research, one is a simple box-shaped environment for training, and the other is a complex furniture scenario for testing, as shown in the appendix Figure 2 As shown, the policy network is trained in the box-shaped environment (a) and directly tested in the furniture environment (b) without additional fine-tuning. In the training stage of cross-modal contrastive learning, the present invention collects 10,000 pairs of RGB-depth image pairs, and the resolution of each pair of images is 224×224. The present invention uses the temperature coefficient τ = 0.07 to optimize the training process, and the network finally generates a 128-dimensional feature vector, aiming to extract consistent and task-related visual feature representations from RGB and depth images.
[0091] During the training process of the visual motion policy, the settings of the training parameters include the batch size M = 1024, the experience replay buffer size D = 10240, the learning rate η =
[0092] 0.0003, the entropy regularization parameter β = 0.01 to encourage policy exploration, the policy clipping parameter ∈ = 0.2 to limit the amplitude of policy updates, and the parameter λ = 0.95 to balance the bias and variance in the advantage estimation. The architecture of the policy network consists of two fully connected layers, and each layer contains 256 hidden units. During the motion execution, the speed scaling coefficient α, the maximum linear velocity V max and the maximum angular velocity Ω max are all set to 1, and the maximum number of steps per training episode is limited to T = 5000.
[0093] To evaluate the performance of the visual motion policy, the present invention conducts comparative experiments with several benchmark methods. First, separate end-to-end policy training models based on RGB images and depth images respectively are used as comparisons. Second, a single-modal benchmark model that only uses RGB images and performs representation learning through data augmentation is selected as a control, called the single-modal RGB benchmark. In addition, a ready-made hybrid artificial potential field method is also used for comparison, which generates a collision-free trajectory through the attractive potential field of the target point and the repulsive potential field of the obstacle.
[0094] The simulation experiments are carried out separately in a normal test environment and a test environment containing unknown visual interference. In the normal test environment, the trajectories of each benchmark method are as shown in the appendix Figure 3As shown. It can be seen from the figure that the method proposed by the present invention performs equivalently to the end-to-end policy training model based on depth images in trajectory generation. However, since several other methods are only trained based on single-modal RGB images, the generalization and robustness of their models are poor. Especially when the color or texture of the obstacles changes, the performance of the model drops significantly. In addition, in unknown visual interference scenarios, the attention distributions of the present invention and the single-modal RGB benchmark are as shown in Appendices Figure 4 and 5 As shown. Figure 4 Among them, the present invention always focuses its attention on key obstacles and targets under various visual interference conditions, avoiding being interfered by irrelevant background information. In contrast, the single-modal RGB method is more sensitive to interference. When the image texture changes, its attention tends to be too scattered or shifted to interfering color blocks. Figure 5 Among them, as the scene complexity increases, the present invention can still maintain the correct focus of attention, while the single-modal RGB method is difficult to achieve this. The results show that the method of the present invention can always focus its attention on the obstacles related to the autonomous flight task, rather than the interfering elements in the background, in the face of interference such as brightness changes, hue perturbations, and random color blocks in both simulation and physical experiments. In contrast, the single-modal RGB benchmark is distracted when encountering these interferences, and the model performance fails rapidly. This further verifies the high generalization and robustness of the present invention in complex and unknown environments.
[0095] To evaluate the performance of the visual motion strategy proposed in the present invention in a real environment, the trained policy network is directly deployed in an indoor environment for testing without additional fine-tuning. However, to narrow the gap between simulation and reality, the feature extraction network is fine-tuned based on a real environment dataset. The experimental scenario settings include several cube obstacles and target points marked with QR codes. The experimental goal is to enable the Tello Edu drone equipped with a monocular RGB camera to successfully complete the visual navigation task in an indoor environment. To achieve high-precision positioning, the OptiTrack motion capture system is used in the experiment to track the position information and attitude of the drone and the target in real time. This system achieves precise positioning and tracking through reflective balls.
[0096] To comprehensively evaluate the performance of the visual motion strategy, the present invention designs a variety of experimental scenarios, including different obstacle layouts, starting point and target point configurations. To further verify the robustness of the strategy in the face of visual interference, obstacles with different colors and textures are used instead of cube obstacles with uniform color blocks in the experiment. The specific scenarios are as shown in Appendices Figure 6As shown, the drone perceives the positions of box-shaped obstacles and targets through an on-board monocular camera and successfully completes a collision-free navigation task. The experimental results show that the strategy of the present invention has strong robustness to random color interference and can stably complete tasks in diverse scenarios. In 50 trials conducted in diverse scenarios, the overall success rate of the strategy reaches 72%, indicating that its visual representation has good task adaptability and at the same time demonstrates strong robustness and transfer ability.
[0097] Embodiment 2:
[0098] A drone visual autonomous navigation device, which particularly includes:
[0099] A modeling unit for modeling the interaction between the drone and the environment as a partially observable Markov decision process. The goal of this decision process is to construct an optimal visual motion strategy so that the drone can select actions to maximize the expected value of the cumulative discounted reward;
[0100] A learning unit for constructing a cross-modal contrastive learning model, and the objective function is defined as:
[0101]
[0102] where I(Z; X RGB ) and I(Z; X Depth ) respectively represent the mutual information between the learned feature representation Z and the RGB and depth inputs;
[0103] A training unit for training the visual motion strategy by using the proximal policy optimization algorithm based on the actor-critic structure.
[0104] Embodiment 3:
[0105] An electronic device includes a processor and a memory communicatively connected to the processor and used for storing instructions executable by the processor. The processor is used to execute the above-mentioned drone visual autonomous navigation method.
[0106] Embodiment 4:
[0107] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned drone visual autonomous navigation method.
[0108] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for autonomous visual navigation of an unmanned aerial vehicle, characterized in that: The steps include: Step 1: Model the interaction between the drone and the environment as a partially observable Markov decision process, the goal of which is to construct an optimal visual-motor policy that enables the drone to choose actions to maximize the expected value of the cumulative discounted reward; Step 2: Construct a cross-modal contrastive learning model. The objective function is defined as: Among them, I(Z;X RGB ) and I(Z;X Depth ) represent the mutual information between the learned feature representation Z and the RGB and depth inputs respectively; Step 3: The visual-motor policy is trained using a proximal policy optimization algorithm based on an actor-critic structure.
2. The method for autonomous visual navigation of an unmanned aerial vehicle according to claim 1, characterized in that: In step 1, the decision process is represented by a six-tuple (S, A, P, R, O, γ), where S represents the state space of the environment, A represents the action space of the agent, and P is the state transition model, defined as P: S × A × S → [0, 1], which represents the current state s t Taking action a t Then transfer to the next state s t+1 The probability of is the reward function responsible for assigning real-valued rewards to the actions performed by the agent, O represents the observation space, and γ∈(0,1) is the discount factor used to balance the weights of immediate rewards and future rewards; The goal of the decision process is to construct an optimal visual-motor strategy π * :O→A, which enables the drone to select actions to maximize the expected value of the cumulative discounted reward, and its objective expression is:
3. The method for autonomous visual navigation of an unmanned aerial vehicle according to claim 1, characterized in that: In step 2, InfoNCE loss is used as the lower bound estimate of mutual information, which is defined as: Where τ is the temperature parameter, f sim (z RGB ,z Depth ) represents the query q RGB and the corresponding key k Depth The similarity function between the extracted feature representations.
4. The method for autonomous visual navigation of an unmanned aerial vehicle according to claim 3, characterized in that: In step 2, the cross-modal contrastive learning model training process is as follows: In each sampling iteration, a batch of original RGB images and their corresponding depth maps are randomly selected to form positive and negative sample pairs for comparative learning; Perform data enhancement operations; The feature extraction network adopts ResNet-50 structure; After obtaining the image embedding, the gradient is calculated according to the loss function, and the model parameters are updated by stochastic gradient descent; After training is complete, freeze the parameters of the feature extraction network.
5. The method for autonomous visual navigation of an unmanned aerial vehicle according to claim 1, characterized in that: In step 3, the actor network maximizes the expected reward by generating actions while constraining the strategy update amplitude to ensure the stability of training. Its loss function is defined as follows: Among them, r t (θ) represents the probability ratio of the strategies before and after the update, is the advantage estimate, ∈ is the clipping parameter used to limit large updates; The critic network reduces the variance of policy updates by estimating the cumulative reward, and its loss function is defined as: Among them, V(s t ) represents the predicted state value, V target is the true state value, calculated as the sum of the current reward and the discounted future reward.
6. The method for autonomous visual navigation of an unmanned aerial vehicle according to claim 1, characterized in that: In step 3, the policy network directly maps the input of the visual motion policy to the original control instruction value, the control instruction value is scaled by the coefficient α, and the part exceeding the threshold is truncated to generate the final execution speed.
7. A visual autonomous navigation device for an unmanned aerial vehicle, characterized in that: include: A modeling unit for modeling the interaction of the drone with the environment as a partially observable Markov decision process, the goal of which is to construct an optimal visual-motor policy that enables the drone to choose actions to maximize the expected value of the cumulative discounted reward; The learning unit is used to build a cross-modal contrastive learning model. The objective function is defined as: Among them, I(Z;X RGB ) and I(Z;X Depth ) represent the mutual information between the learned feature representation Z and the RGB and depth inputs respectively; A training unit is used to train the visual motion policy using a proximal policy optimization algorithm based on an actor-critic structure.
8. An electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, characterized in that: The processor is used to execute a drone visual autonomous navigation method as described in any one of claims 1-6 above.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for visual autonomous navigation of a drone as described in any one of claims 1 to 6 is implemented.