Visual navigation method and system of six-rotor full-drive unmanned aerial vehicle, terminal and storage medium
Through the combination of visual Transformer and long and short-term memory networks, a fusion navigation framework of dynamic visual perception and full drive control is built, which solves the problem of insufficient navigation adaptability of the six-rotor drone in high-dynamic scenarios, and achieves high-precision and robust navigation effects.
Patent Information
- Application Number
- CN202510706497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In the prior art, the six-rotor drone has insufficient visual navigation adaptability in high dynamic motion scenarios, the multi-sensor fusion complexity is high, the computing power consumption and load are limited, and the dynamic characteristics of the full-drive drone are not fully considered, resulting in insufficient navigation accuracy.
Visual Transformer is used to extract multi-scale environmental features, combine long and short-term memory networks to process timing information, and build a fusion navigation framework for dynamic visual perception and full-drive control. The global correlation modeling of environmental features is strengthened through the self-attention mechanism, and combined with the generation of control instructions of dynamic constraints, to achieve motion decoupling control.
It significantly improves the navigation accuracy and robustness of the drone in complex dynamic environments, solves the problem of feature tracking loss of traditional algorithms in high dynamic scenarios, and provides an efficient and reliable navigation solution.
Smart Images

Figure CN120255492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) perception, and particularly to a visual navigation method, system, terminal, and computer-readable storage medium for a six-rotor fully-driven UAV. Background Art
[0002] As one of the core technologies for UAV autonomous navigation, visual navigation technology has made significant progress in algorithm optimization and hardware integration in recent years. Mainstream solutions include feature point-based simultaneous localization and mapping (SLAM), deep learning end-to-end navigation, and multi-sensor fusion technology. For example, the LSD-SLAM algorithm (Large-Scale Direct Monocular SLAM, a direct-method monocular visual SLAM algorithm) realizes real-time pose estimation through the direct method and has been verified effective in quadrotor UAVs.
[0003] However, the existing technologies have significant defects. Traditional visual SLAM algorithms (such as ORB-SLAM3) rely on sparse feature point matching and are prone to feature tracking loss or map drift problems in high-dynamic motion scenarios such as rapid rotation and acceleration of six-rotor UAVs. For example, when the UAV rotates at a large angular velocity, the feature point extraction speed cannot match the motion speed, resulting in excessive positioning delay.
[0004] Current research on the navigation of six-rotor UAVs mainly focuses on path planning and control algorithms (such as model predictive control), while there is less special research on visual navigation. For example, some six-rotor UAVs use lidar for obstacle avoidance, but their visual navigation function only supports simple target recognition and cannot achieve autonomous positioning. In addition, the pose decoupling control of the fully-driven platform has not been deeply integrated with visual navigation, resulting in insufficient navigation accuracy.
[0005] Therefore, the existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of the present invention is to provide a visual navigation method, system, terminal, and computer-readable storage medium for a six-rotor fully-driven UAV, aiming to solve the problems of insufficient adaptability in high-dynamic motion scenarios, high complexity of multi-sensor fusion, limited computing power and payload in the existing visual navigation technology, and the lack of full consideration of the dynamic characteristics of six-rotor fully-driven UAVs.
[0007] To achieve the above objective, the present invention provides a visual navigation method for a six-rotor fully-driven UAV, and the visual navigation method for the six-rotor fully-driven UAV includes the following steps: Collect the pixel depth map, real-time pose quaternion, and forward speed during the flight of a six-rotor fully-driven drone. Perform invalid pixel filling and pixel value normalization on the pixel depth map to obtain a preprocessed depth map sequence. Transmit the depth map sequence, the real-time pose quaternion, and the forward speed to the ARM processor; Input the depth map sequence into a visual encoder. Divide each frame of the image into a patch sequence of pixels of a preset size through image embedding operations. Capture the global obstacle spatial correlation through a self-attention layer and output a multi-scale feature vector. Input the multi-scale feature vectors of consecutive frames into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the drone and generate an environmental state vector containing the time dimension; Input the environmental state vector, the real-time pose quaternion, and the forward speed into an MPC controller. With the goal of minimizing the trajectory tracking error and energy consumption, optimize the three-dimensional velocity command within a preset control period. Based on the six-rotor drone fully-driven dynamics model, convert the optimized three-dimensional velocity command into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle changes through a non-linear mapping function to achieve motion decoupling control.
[0008] Optionally, in the visual navigation method of the six-rotor fully-driven drone, where the step of collecting the pixel depth map, real-time pose quaternion, and forward speed during the flight of the six-rotor fully-driven drone, performing invalid pixel filling and pixel value normalization on the pixel depth map to obtain a preprocessed depth map sequence specifically includes: Collect a pixel depth map of 60×90 at a preset frequency through a depth camera carried by the six-rotor fully-driven drone, and synchronously obtain the real-time pose quaternion of the drone through the IMU and the forward speed ; Perform invalid pixel filling on the pixel depth map through an FPGA hardware module, and normalize the pixel values of the pixel depth map to the range [0, 1] to obtain a preprocessed depth map sequence.
[0009] Optionally, in the visual navigation method of the six-rotor fully-driven drone, where the step of inputting the depth map sequence into a visual encoder, dividing each frame of the image into a patch sequence of pixels of a preset size through image embedding operations, capturing the global obstacle spatial correlation through a self-attention layer, outputting a multi-scale feature vector, inputting the multi-scale feature vectors of consecutive frames into a bidirectional LSTM network, modeling the temporal dependence relationship during the high-speed movement of the drone, and generating an environmental state vector containing the time dimension specifically includes: Input the depth map sequence into the visual encoder. The pixel depth map input by the depth camera is divided into a patch sequence of 16×16 pixels through the ViT feature extraction module. Calculate the global correlation between patches through the visual encoder, and obtain the global environmental features through the self-attention mechanism of ViT; For each patch sequence, combine a learnable weight matrix to generate corresponding query vectors, key vectors, and value vectors, and calculate the original similarity score matrix: Use the softmax function to normalize the scores in the row direction to obtain a weight matrix reflecting the patch correlation strength; Weightedly sum the value vectors according to the attention weights to generate patch features containing the global environment; Aggregate all patch feature sequences through a multi-layer perceptron, and finally output a multi-scale feature vector containing global environmental features; Connect to a bidirectional LSTM network to perform temporal modeling on the multi-scale feature vectors of three consecutive frames, and output an environmental state vector containing time dependencies : ; where, represents the multi-scale feature of the first frame, represents the multi-scale feature of the second frame, represents the multi-scale feature of the third frame.
[0010] Optionally, in the visual navigation method of the six-rotor fully-driven drone, the ViT feature extraction module includes an image embedding layer, a self-attention layer, and a multi-layer perceptron.
[0011] Optionally, in the visual navigation method of the six-rotor fully-driven drone, input the environmental state vector, the real-time pose quaternion, and the forward speed into the MPC controller, and optimize the three-dimensional speed command within a preset control period with the goal of minimizing the trajectory tracking error and energy consumption. Specifically include: Construct a full-driven dynamic model of the six-rotor drone. If the mapping from the speed command to the actual acceleration is a linear relationship: ; where, represents the actual acceleration, represents the target acceleration output by the MPC controller, represents the acceleration gain matrix, represents the three-dimensional speed command, represents the forward speed; Construct a cost function with the goal of minimizing the tracking error and energy consumption: ; Among them, represents the cost function, represents the number of control cycles, represents a certain control cycle, represents the desired speed, represents the weight coefficient, represents the three-dimensional velocity command for the th control cycle, represents the desired speed for the th control cycle, represents the target acceleration output by the MPC controller for the th control cycle; , ), and the global environmental characteristics , through the MPC controller, obtain the control sequence { } for the next steps, where represents the three-dimensional velocity command for the th control cycle, represents the three-dimensional velocity command for the th control cycle, represents the three-dimensional velocity command for the th control cycle.
[0012] Optionally, for the visual navigation method of the six-rotor fully actuated drone, among them, based on the six-rotor drone fully actuated dynamics model, convert the optimized three-dimensional velocity command into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle changes through a non-linear mapping function to achieve motion decoupling control, specifically including: Calculate the thrust and counter-torque generated by each rotor of the six-rotor drone: ; ; ; Among them, represents the thrust coefficient, represents the torque coefficient, represents the th rotor's rotational angular velocity; Based on the six degrees of freedom motion of the six-rotor drone, establish a non-linear mapping relationship between the rotor rotational speeds and the fuselage acceleration and attitude angles: ; Among them, represents the non-linear mapping relationship, represents the linear acceleration in the x-axis direction of the inertial coordinate system, Represents the linear acceleration in the y-axis direction of the inertial coordinate system, Represents the linear acceleration in the z-axis direction of the inertial coordinate system, Represents the roll angular velocity, Represents the pitch angular velocity, Represents the yaw angular velocity; Combine the contributions of each rotor into the generalized force received by the UAV: ; Among them, Represents the component forces in three directions in the body coordinate system, Represents the steady-state moments of the UAV in three directions, Represents the moment of inertia, Represents the mass of the UAV; Relate the rotor speed to the generalized control force through the structure of the six-rotor UAV: ; Among them, Represents the relationship between the rotor speed and the generalized control force, Represents the distribution matrix, Represents the rotational angular velocity of the first rotor, Represents the rotational angular velocity of the second rotor, Represents the rotational angular velocity of the third rotor, Represents the rotational angular velocity of the fourth rotor, Represents the rotational angular velocity of the fifth rotor, Represents the rotational angular velocity of the sixth rotor; Convert the three-dimensional velocity command into the desired acceleration: ; Among them, Represents the desired acceleration, Represents the control period, Represents the gravitational acceleration converted to the body coordinate system; Calculate the required generalized control force through the dynamic inverse model: ; Use the pseudo-inverse of the distribution matrix to solve for the squared rotational speed of each rotor: ; Obtain the actual rotational speed by taking the square root: .
[0013] Optionally, in the visual navigation method of the six-rotor fully-driven UAV, the ARM processor is used for dynamic visual feature extraction and timing modeling operations.
[0014] In addition, to achieve the above object, the present invention further provides a visual navigation system for a six-rotor fully-driven unmanned aerial vehicle, wherein the visual navigation system of the six-rotor fully-driven unmanned aerial vehicle includes: A data acquisition and preprocessing module, configured to acquire a pixel depth map, a real-time pose quaternion, and a forward speed during the flight of the six-rotor fully-driven unmanned aerial vehicle, perform invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmit the depth map sequence, the real-time pose quaternion, and the forward speed to an ARM processor; A dynamic visual feature extraction and temporal modeling module, configured to input the depth map sequence into a visual encoder, divide each frame of image into a patch sequence of pixels of a preset size through image embedding operation, capture the global obstacle spatial association through a self-attention layer, output a multi-scale feature vector, input the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network, model the temporal dependence relationship during the high-speed movement of the unmanned aerial vehicle, and generate an environmental state vector including the time dimension; A navigation control coordination and instruction generation module, configured to input the environmental state vector, the real-time pose quaternion, and the forward speed into an MPC controller, aiming to minimize the trajectory tracking error and energy consumption, optimize the three-dimensional speed instruction within a preset control period, and based on the six-rotor unmanned aerial vehicle fully-driven dynamics model, convert the optimized three-dimensional speed instruction into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle change through a non-linear mapping function to achieve motion decoupling control.
[0015] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a visual navigation program for a six-rotor fully-driven unmanned aerial vehicle stored on the memory and executable on the processor, and when the visual navigation program for the six-rotor fully-driven unmanned aerial vehicle is executed by the processor, the steps of the above-mentioned visual navigation method for the six-rotor fully-driven unmanned aerial vehicle are implemented.
[0016] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a visual navigation program for a six-rotor fully-driven unmanned aerial vehicle, and when the visual navigation program for the six-rotor fully-driven unmanned aerial vehicle is executed by a processor, the steps of the above-mentioned visual navigation method for the six-rotor fully-driven unmanned aerial vehicle are implemented.
[0017] In the present invention, during the flight of a six-rotor fully-driven unmanned aerial vehicle (UAV), a pixel depth map, real-time pose quaternion, and forward speed are collected. The pixel depth map is subjected to invalid pixel filling processing and pixel value normalization processing to obtain a preprocessed depth map sequence. The depth map sequence, the real-time pose quaternion, and the forward speed are transmitted to an ARM processor. The depth map sequence is input into a visual encoder. Through image embedding operation, each frame of image is divided into a patch sequence of pixels of a preset size. The self-attention layer is used to capture the global obstacle spatial correlation, and a multi-scale feature vector is output. The multi-scale feature vectors of consecutive multiple frames are input into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the UAV, and an environmental state vector including the time dimension is generated. The environmental state vector, the real-time pose quaternion, and the forward speed are input into an MPC controller. With the goal of minimizing the trajectory tracking error and energy consumption, the three-dimensional speed command is optimized within a preset control period. Based on the six-rotor UAV fully-driven dynamics model, the optimized three-dimensional speed command is converted into the independent rotational speeds of 6 rotors. The body acceleration and attitude angle changes are matched through a non-linear mapping function to achieve motion decoupling control. By constructing a navigation framework that integrates dynamic visual perception and fully-driven control characteristics, the present invention uses a visual encoder to extract multi-scale environmental features, combines long short-term memory networks to process temporal information, realizes efficient feature extraction and motion trajectory prediction for image sequences under high-speed rotation and drastic attitude changes of the UAV, strengthens the global correlation modeling of environmental features through an attention mechanism, and combines the generation of control commands with dynamic constraints, significantly improving the navigation accuracy and robustness of the UAV in complex dynamic environments. Description of the Drawings
[0018] Figure 1 is a flowchart of a preferred embodiment of the visual navigation method for a six-rotor fully-driven UAV of the present invention; Figure 2 is a principle block diagram of a visual perception module in a preferred embodiment of the visual navigation method for a six-rotor fully-driven UAV of the present invention; Figure 3 is a structural diagram of a preferred embodiment of the visual navigation system for a six-rotor fully-driven UAV of the present invention; Figure 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Description of the Invention
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer and more explicit, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] Aiming at the problems of insufficient adaptability of existing visual navigation technologies in high-dynamic motion scenarios, high complexity of multi-sensor fusion, limited computing power consumption and payload, and the dynamics characteristics of six-rotor fully actuated drones not being fully considered, the present invention proposes an end-to-end visual navigation method based on Vision Transformer. By constructing a navigation framework that integrates dynamic visual perception and fully actuated control characteristics, using Vision Transformer to extract multi-scale environmental features, and combining Long Short-Term Memory (LSTM) network to process temporal information, it realizes efficient feature extraction and motion trajectory prediction for image sequences under the conditions of high-speed rotation and drastic attitude changes of the drone, and solves the problem of feature tracking loss of traditional algorithms in high-dynamic scenarios.
[0021] The present invention enhances the global correlation modeling of environmental features through the attention mechanism, combines the generation of control instructions with dynamic constraints, significantly improves the navigation accuracy and robustness of the drone in complex dynamic environments, and does not rely on high-precision inertial navigation compensation, providing an efficient and reliable navigation solution for the application of six-rotor fully actuated drones in scenarios such as emergency rescue and high-speed inspection.
[0022] The visual navigation method of the six-rotor fully actuated drone according to a preferred embodiment of the present invention, as Figure 1 and Figure 2 shown, the visual navigation method of the six-rotor fully actuated drone includes the following steps: Step S10: Collect the pixel depth map, real-time pose quaternion and forward speed during the flight of the six-rotor fully actuated drone, perform invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmit the depth map sequence, the real-time pose quaternion and the forward speed to the ARM processor.
[0023] Specifically, as Figure 2 shown, use the depth camera carried by the six-rotor fully actuated drone to collect a 60×90 pixel depth map from the start to the whole flight process at a preset frequency (for example, 30Hz), and synchronously obtain the real-time pose quaternion of the drone through the IMU (pose represents the position and attitude of an object in three-dimensional space, where the position is the coordinates of the object in three-dimensional space, represented by a triple ( ), and the attitude is the current rotation direction of the object, represented by a quaternion ( ), where is the quaternion representing the current attitude of the drone in the pose information) and the forward speed ; Then, the FPGA hardware module performs invalid pixel filling processing on the pixel depth map (each pixel value of the pixel depth map represents the distance between the corresponding point in the scene and the camera. Due to reasons such as sensor noise, occlusion, and out-of-measurement range, there may be invalid pixels in the pixel depth map, and these areas need to be filled), and normalizes the pixel values of the pixel depth map to the range [0, 1] (linearly scales the pixel values of the pixel depth map from the original range (where is the maximum measurement distance of the sensor) to [0, 1] so that all pixel values fall between 0 and 1), obtaining a preprocessed depth map sequence. After that, the preprocessed depth map sequence (3 consecutive frames), the real-time pose quaternion and the forward velocity are transmitted to the ARM processor, facilitating the use of the computing power of the ARM processor for subsequent dynamic vision feature extraction and temporal modeling operations.
[0024] Step S20: Input the depth map sequence into the vision encoder. Through image embedding operation, each frame of the image is divided into a patch sequence of preset-sized pixels. The self-attention layer captures the global obstacle spatial association, and outputs a multi-scale feature vector. The multi-scale feature vectors of consecutive multiple frames are input into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the drone, generating an environmental state vector containing the time dimension.
[0025] Specifically, the depth map sequence is input into the vision encoder (Vision Transformer). Through image embedding operation, each frame of the image is divided into a patch sequence of 16×16 pixels. The self-attention layer captures the global obstacle spatial association, and outputs a multi-scale feature vector. Then, the multi-scale feature vectors of 3 consecutive frames are input into a bidirectional LSTM network (Long Short-Term Memory Network), modeling the temporal dependence relationship during the high-speed movement of the drone, generating an environmental state vector , solving the problem of feature loss in single-frame images in dynamic scenes.
[0026] As Figure 2 shown, a vision encoder and LSTM fusion architecture is adopted to process the image sequence (60×90 pixel depth map) input by the depth camera carried by the drone. The image is divided into a patch sequence of 16×16 pixels through the ViT feature extraction module (ViT, Vision Transformer, vision encoder, including an image embedding layer, a self-attention layer, and a multi-layer perceptron), and the global environmental features are captured through the self-attention mechanism.
[0027] Among them, the implementation process of each step in the ViT feature extraction module is as follows: Three frames of pixel depth maps are stitched into a 3×60×90 tensor. Each frame of depth map is divided into 15 non-overlapping patches of 16×16 pixels. There are 45 patches in total for 3 frames. Each patch is flattened into a 256-dimensional vector. Then, the global correlation between patches is calculated through a visual encoder, and the global environmental features are obtained through the self-attention mechanism of ViT. The principle of the self-attention mechanism is described as follows: ; Among them, represents the query vector, represents the key vector, represents the value vector, represents the dimension of the key vector, represents the transpose.
[0028] A. For each 256-dimensional patch vector, it is multiplied by three learnable weight matrices respectively to generate the corresponding query vector , key vector and value vector . These vectors are used for subsequent calculation of attention scores.
[0029] B. For all and ( j = 1, 2,..., 45), the dot product is calculated. To ensure numerical stability, it is then divided by the square root of the vector dimension scaling factor (i.e., ) to obtain the original similarity score matrix: ; C. Use the function to normalize the scores in the row direction to obtain the weight matrix reflecting the patch correlation strength: ; Among them, represents the number of key vectors, represents the th key vector; D. Weighted sum the value vectors according to the attention weights to generate the patch feature containing the global environment: ; Among them, represents the j th value vector; E. Finally, all patch feature sequences are aggregated through a multi-layer perceptron, and the visual feature vector containing the global environmental features is finally output.
[0030] Access the bidirectional LSTM network to perform temporal modeling on the feature sequences of three consecutive frames and output an environmental state vector containing time dependencies. , where represents the multi-scale features of the first frame, represents the multi-scale features of the second frame, represents the multi-scale features of the third frame.
[0031] Step S30: Input the environmental state vector, the real-time pose quaternion, and the forward velocity into the MPC controller. With the goal of minimizing the trajectory tracking error and energy consumption, optimize the three-dimensional velocity command within a preset control period. Based on the full-actuated dynamics model of the six-rotor UAV, convert the optimized three-dimensional velocity command into the independent rotational speeds of the six rotors, and match the fuselage acceleration and attitude angle changes through a non-linear mapping function to achieve motion decoupling control.
[0032] Specifically, input the environmental state vector output by the visual perception module and the real-time pose ( , ) into the MPC controller (Model Predictive Control, a control system embedded in the flight control computer of the UAV in algorithm form). With the goal of minimizing the trajectory tracking error and energy consumption, optimize the three-dimensional velocity command within a 50ms control period ∈ , and then based on the full-actuated dynamics model of the six-rotor UAV, convert the three-dimensional velocity command into the independent rotational speeds of the six rotors , and through the non-linear mapping function match the fuselage acceleration and attitude angle changes to achieve motion decoupling control.
[0033] The optimization principle is as follows: A. Construct the full-actuated dynamics model of the six-rotor UAV. If the mapping from the velocity command to the actual acceleration is a linear relationship: ; where represents the actual acceleration (the derivative of velocity with respect to time), represents the target acceleration output by the MPC controller, represents the acceleration gain matrix used to adjust the tracking response speed, represents the three-dimensional velocity command, that is, the velocity in the three dimensions in the space coordinate system, represents the forward velocity (the goal is to optimize this command sequence).
[0034] B. Construct a cost function with the goal of minimizing the tracking error and energy consumption: ; Among them, represents the cost function, represents the number of control cycles, represents a certain control cycle, represents the desired speed, represents the weight coefficient, represents the th three-dimensional velocity command of the control cycle, represents the th desired speed of the control cycle, represents the th target acceleration output by the MPC controller of the control cycle.
[0035] C. In each control cycle, based on the current state ( , ) and the global environmental characteristics , obtain the control sequence { } for the next steps through the MPC controller. Among them, represents the th three-dimensional velocity command of the control cycle, represents the th three-dimensional velocity command of the control cycle, represents the th three-dimensional velocity command of the control cycle.
[0036] D. Only execute the velocity command of the first cycle, and repeat this process in the next cycle.
[0037] Then convert the velocity command into the independent rotational speeds of the 6 rotors through the full-driven thrust allocation algorithm , and the conversion and derivation process is as follows: A. Construct a dynamic model. The six-rotor UAV is a full-driven system, and calculate the thrust and anti-torque generated by each rotor of the six-rotor UAV: Generate thrust and anti-torque: ; ; Among them, represents the thrust coefficient, represents the torque coefficient, represents the th rotational angular velocity of the rotor.
[0038] Based on the six - degree - of - freedom motion of a hexacopter UAV, a non - linear mapping relationship between the rotor speeds and the body acceleration and attitude angles is established: ; Among them, represents the non - linear mapping relationship, ( i i = 1, 2, …, 6) represents the rotational angular velocity of the i - th rotor, represents the linear acceleration in the x - axis direction of the inertial coordinate system, characterizing the acceleration motion of the body along the horizontal transverse direction, represents the linear acceleration in the y - axis direction of the inertial coordinate system, characterizing the acceleration motion of the body along the horizontal longitudinal direction, represents the linear acceleration in the z - axis direction of the inertial coordinate system, characterizing the acceleration motion of the body along the vertical direction, represents the roll angular velocity, that is, the rotational angular velocity of the body about the x - axis, represents the pitch angular velocity, that is, the rotational angular velocity of the body about the y - axis, represents the yaw angular velocity, that is, the rotational angular velocity of the body about the z - axis.
[0039] B. Combine the contributions of each rotor into the generalized force received by the UAV: ; Among them, represents the component forces in three directions in the body coordinate system, represents the steady - state torques of the UAV in three directions (neglecting angular acceleration, applicable to rapid thrust distribution), represents the moment of inertia, represents the mass of the UAV.
[0040] C. Relate the rotor speeds to the generalized control forces through the structure of the hexacopter UAV: ; Among them, represents the relationship between the rotor speeds and the generalized control forces, represents the distribution matrix, represents the rotational angular speed of the first rotor, represents the rotational angular speed of the second rotor, represents the rotational angular speed of the third rotor, represents the rotational angular speed of the fourth rotor, represents the rotational angular speed of the fifth rotor, represents the rotational angular speed of the sixth rotor.
[0041] D. Convert the three - dimensional velocity command into the desired acceleration: ; Among them, represents the desired acceleration, represents the control period, represents the gravitational acceleration converted to the body coordinate system.
[0042] E. Calculate the required generalized control force through the inverse dynamics model: ; F. Use the pseudo-inverse of the distribution matrix to solve the squared rotational speed of each rotor: ; G. Obtain the actual rotational speed by taking the square root: .
[0043] The present invention adopts an ARM+FPGA heterogeneous architecture: the ARM processor runs the inference of the ViT-LSTM neural network, and the FPGA realizes the parallel acceleration of the self-attention layer and the hardware implementation of the thrust distribution algorithm. The depth map preprocessing is completed by the FPGA, the global environmental features extracted by the ViT are transmitted to the ARM through a high-speed bus, and the final control instructions are generated by the FPGA in real time.
[0044] The innovation points of the present invention are as follows: (1). End-to-end vision-control collaborative architecture (different from traditional modular navigation): For the first time, the vision Transformer is combined with the long short-term memory network to construct a dynamic vision perception module, which directly extracts the environmental state vector containing global spatial associations and temporal motion features from the depth image sequence, replacing the traditional method that relies on sparse feature point matching or manually designed features.
[0045] (2). Model predictive control (MPC) adapted to full-driven dynamics: Design an MPC closed-loop control algorithm based on the full-driven characteristics of a six-rotor, directly map the pose information output by visual navigation to the independent rotational speed commands of 6 rotors, and achieve motion decoupling through a non-linear dynamics model to break the decoupled design of traditional navigation and control modules such as "positioning first and then separately planning the trajectory".
[0046] (3). Lightweight heterogeneous computing hardware architecture: Adopt an ARM+FPGA heterogeneous computing solution, the FPGA parallelly accelerates the ViT self-attention layer and the thrust distribution algorithm, and the ARM processes high-level inferences, reducing the overall power consumption and weight, and solving the payload and endurance bottlenecks of existing visual navigation systems.
[0047] The beneficial effects brought by the present invention: (1) Improvement in dynamic environment adaptability: The global obstacle distribution is captured through the self-attention mechanism, which improves the long-distance feature recognition rate compared with traditional CNN methods such as ResNet-18; LSTM processes temporal information, and in the high-speed rotation scenario, the feature tracking success rate is improved, avoiding navigation failure caused by image blur. At the same time, the global modeling ability of ViT and the temporal memory ability of LSTM solve the problem of dynamic scene feature loss in the prior art.
[0048] (2) Optimization of full-drive control accuracy: Reduce the pose error output by visual navigation, and end-to-end learning eliminates the cumulative error caused by the decoupling of navigation and control, shortening the trajectory tracking delay when the UAV makes a high-speed turn. And the optimization of the full-drive dynamic model and MPC directly improves the feasibility of control commands, avoiding the problem of mismatch between control commands and the physical characteristics of the UAV in traditional methods.
[0049] Alternative implementation and deformation methods of the present invention: (1) Alternative sensor scheme: In addition to the RGB-D (RGB Image and Depth Image) camera, lidar can also be used to measure distance information. However, the cost of lidar is much higher than that of the depth camera, and it has a large volume and high power consumption, which is not very suitable for lightweight UAVs. On the contrary, the depth camera is structurally compact, has low power consumption, and is convenient for embedded deployment.
[0050] (2) Network structure deformation: The ViT encoder can be replaced by a traditional convolutional neural network CNN (Convolutional Neural Network). However, ViT directly captures cross-region long-distance dependencies through the self-attention mechanism, avoiding the limitations of CNN indirectly modeling global information through multi-layer stacking or global pooling. It is more effective for tasks that require global semantic understanding, such as image classification and long-tail distribution scenarios.
[0051] (3) For visual perception, ORB-SLAM3 or LSD-SLAM is used for sparse feature point localization to replace the end-to-end feature extraction of ViT-LSTM; the control module uses a classic PID controller to adjust the speed of the six-rotor, and the thrust distribution is solved through an independent kinematic model, avoiding the joint design of MPC and the dynamic mapping function thereof.
[0052] Furthermore, as Figure 3 shown, based on the above visual navigation method for a six-rotor full-drive UAV, the present invention also correspondingly provides a visual navigation system for a six-rotor full-drive UAV, wherein the visual navigation system for the six-rotor full-drive UAV includes: The data acquisition and preprocessing module 51 is used to acquire the pixel depth map, real-time pose quaternion, and forward speed during the flight of the six-rotor fully-driven drone, perform invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmit the depth map sequence, the real-time pose quaternion, and the forward speed to the ARM processor; The dynamic vision feature extraction and temporal modeling module 52 is used to input the depth map sequence into the vision encoder, divide each frame of image into a patch sequence of pixels of a preset size through image embedding operation, capture the global obstacle spatial association through the self-attention layer, output a multi-scale feature vector, input the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network, model the temporal dependence relationship during the high-speed movement of the drone, and generate an environmental state vector including the time dimension; The navigation control coordination and instruction generation module 53 is used to input the environmental state vector, the real-time pose quaternion, and the forward speed into the MPC controller, aiming to minimize the trajectory tracking error and energy consumption, optimize the three-dimensional speed instruction within a preset control period, and based on the six-rotor drone fully-driven dynamics model, convert the optimized three-dimensional speed instruction into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle change through a non-linear mapping function to achieve motion decoupling control.
[0053] Further, as Figure 4 shown, based on the above-mentioned vision navigation method and system of the six-rotor fully-driven drone, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0054] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a visual navigation program 40 of a six-rotor fully-driven unmanned aerial vehicle is stored on the memory 20, and the visual navigation program 40 of the six-rotor fully-driven unmanned aerial vehicle can be executed by the processor 10, so as to implement the visual navigation method of the six-rotor fully-driven unmanned aerial vehicle in the present application.
[0055] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the visual navigation method of the six-rotor fully-driven unmanned aerial vehicle, etc.
[0056] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, the memory 20, and the display 30 of the terminal communicate with each other through a system bus.
[0057] In one embodiment, when the processor 10 executes the visual navigation program 40 of the six-rotor fully-driven unmanned aerial vehicle in the memory 20, the steps of the visual navigation of the six-rotor fully-driven unmanned aerial vehicle as described above are implemented.
[0058] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a visual navigation program of a six-rotor fully-driven unmanned aerial vehicle, and when the visual navigation program of the six-rotor fully-driven unmanned aerial vehicle is executed by a processor, the steps of the visual navigation method of the six-rotor fully-driven unmanned aerial vehicle as described above are implemented.
[0059] In summary, the present invention provides a visual navigation method, system, terminal and storage medium for a six-rotor fully-driven unmanned aerial vehicle. The method includes: collecting a pixel depth map, a real-time pose quaternion, and a forward speed during the flight of the six-rotor fully-driven unmanned aerial vehicle, performing invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmitting the depth map sequence, the real-time pose quaternion, and the forward speed to an ARM processor; inputting the depth map sequence into a visual encoder, dividing each frame of image into a patch sequence of preset-sized pixels through image embedding operation, capturing the global obstacle spatial association through a self-attention layer, and outputting a multi-scale feature vector. Inputting the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the unmanned aerial vehicle and generating an environmental state vector including the time dimension; inputting the environmental state vector, the real-time pose quaternion, and the forward speed into an MPC controller, aiming to minimize the trajectory tracking error and energy consumption, optimizing the three-dimensional speed command within a preset control period, and based on the six-rotor unmanned aerial vehicle full-driven dynamics model, converting the optimized three-dimensional speed command into the independent rotational speeds of 6 rotors, and matching the fuselage acceleration and attitude angle change through a nonlinear mapping function to achieve motion decoupling control. By constructing a navigation framework that integrates dynamic visual perception and full-driven control characteristics, the present invention uses a visual encoder to extract multi-scale environmental features, combines long short-term memory networks to process temporal information, realizes efficient feature extraction and motion trajectory prediction for image sequences under high-speed rotation and drastic attitude changes of the unmanned aerial vehicle, strengthens the global association modeling of environmental features through an attention mechanism, and combines the generation of control instructions with dynamic constraints, significantly improving the navigation accuracy and robustness of the unmanned aerial vehicle in complex dynamic environments.
[0060] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.
[0061] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0062] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or modifications can be made according to the above description, and all such improvements and modifications shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A visual navigation method for a six-rotor fully-driven drone, characterized in that, The visual navigation method of the six-rotor fully-driven UAV includes: Collect the pixel depth map, real-time pose quaternion, and forward speed during the flight of the six-rotor fully-driven UAV, perform invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmit the depth map sequence, the real-time pose quaternion, and the forward speed to the ARM processor; Input the depth map sequence into the visual encoder. Divide each frame of image into a patch sequence of preset-sized pixels through image embedding operation, capture the global obstacle spatial correlation through the self-attention layer, output a multi-scale feature vector, and input the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the UAV, and generate an environmental state vector containing the time dimension; Input the environmental state vector, the real-time pose quaternion, and the forward speed into the MPC controller. With the goal of minimizing the trajectory tracking error and energy consumption, optimize the three-dimensional speed command within a preset control period. Based on the six-rotor UAV fully-driven dynamics model, convert the optimized three-dimensional speed command into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle change through a non-linear mapping function to achieve motion decoupling control.
2. The visual navigation method of the six-rotor fully-driven unmanned aerial vehicle according to claim 1, wherein The process of collecting the pixel depth map, real-time pose quaternion, and forward speed during the flight of the six-rotor fully-driven UAV, performing invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence specifically includes: The depth camera carried by the six-rotor fully-driven drone collects a pixel depth map of 60×90 at a preset frequency, and the real-time pose quaternion of the drone is obtained synchronously through the IMU and the forward speed ; Perform invalid pixel filling processing on the pixel depth map through the FPGA hardware module, and normalize the pixel values of the pixel depth map to the range of [0, 1] to obtain a preprocessed depth map sequence.
3. The visual navigation method of the six-rotor fully-driven drone according to claim 2, characterized in that The process of inputting the depth map sequence into the visual encoder, dividing each frame of image into a patch sequence of preset-sized pixels through image embedding operation, capturing the global obstacle spatial correlation through the self-attention layer, outputting a multi-scale feature vector, and inputting the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network to model the temporal dependence relationship during the high-speed movement of the UAV, and generating an environmental state vector containing the time dimension specifically includes: Input the depth map sequence into the visual encoder. Divide the pixel depth map input by the depth camera into a patch sequence of 16×16 pixels through the ViT feature extraction module, calculate the global correlation between patches through the visual encoder, and obtain the global environmental features through the self-attention mechanism of ViT; For each patch sequence, combine a learnable weight matrix to generate corresponding query vectors, key vectors, and value vectors, and calculate the original similarity score matrix: Normalize the row-direction scores using the softmax function to obtain a weight matrix reflecting the patch correlation strength; Perform weighted summation on the value vectors according to the attention weights to generate patch features containing the global environment; Aggregate all the patch feature sequences through a multi-layer perceptron, and finally output a multi-scale feature vector containing the global environmental features; Access the bidirectional LSTM network to perform temporal modeling on the multi-scale feature vectors of three consecutive frames, and output the environmental state vector containing time dependence : ; Among them, represents the multi-scale feature of the first frame, represents the multi-scale feature of the second frame, represents the multi-scale feature of the third frame.
4. The visual navigation method of the six-rotor fully-driven unmanned aerial vehicle according to claim 3, characterized in that, The ViT feature extraction module includes an image embedding layer, a self-attention layer, and a multi-layer perceptron.
5. The visual navigation method of the six-rotor fully-driven drone according to claim 3, characterized in that Inputting the environmental state vector, the real-time pose quaternion, and the forward velocity into the MPC controller, with the goal of minimizing the trajectory tracking error and energy consumption, and optimizing the three-dimensional velocity command within a preset control period, specifically including: Construct a full-driven dynamic model of a six-rotor UAV. If the mapping from the velocity command to the actual acceleration is a linear relationship: ; Among them, represents the actual acceleration, represents the target acceleration output by the MPC controller, represents the acceleration gain matrix, represents the three-dimensional velocity command, represents the forward velocity; Constructing a cost function with the goal of minimizing the tracking error and energy consumption: ; Among them, represents the cost function, represents the number of control cycles, represents a certain control cycle, represents the desired speed, represents the weight coefficient, represents the three-dimensional velocity command of the th control cycle, represents the desired speed of the th control cycle, represents the target acceleration output by the MPC controller in the th control cycle; In each control period, based on the current state ( , ), and the global environmental characteristics , the control sequence for the next steps is obtained through the MPC controller, where represents the three-dimensional velocity command for the -th control period, represents the three-dimensional velocity command for the -th control period, represents the three-dimensional velocity command for the -th control period, represents the three-dimensional velocity command for the -th control period.
6. The visual navigation method of the six-rotor fully-driven unmanned aerial vehicle according to claim 5, wherein, Based on the fully actuated dynamics model of the hexacopter UAV, converting the optimized three-dimensional velocity command into the independent rotational speeds of 6 rotors, and matching the fuselage acceleration and attitude angle changes through a non-linear mapping function to achieve motion decoupling control, specifically including: Calculate each rotor of the hexacopter Generate thrust and counter torque: ; ; Among them, represents the thrust coefficient, represents the torque coefficient, represents the rotational angular velocity of the Based on the 6-degree-of-freedom motion of the hexacopter UAV, establishing a non-linear mapping relationship between the rotor rotational speed and the fuselage acceleration and attitude angle: ; Among them, represents a non-linear mapping relationship, represents the linear acceleration in the x-axis direction of the inertial coordinate system, represents the linear acceleration in the y-axis direction of the inertial coordinate system, represents the linear acceleration in the z-axis direction of the inertial coordinate system, represents the roll angular velocity, represents the pitch angular velocity, represents the yaw angular velocity; Combining the contributions of each rotor into the generalized force received by the UAV: ; Among them, represents the component forces in three directions in the body coordinate system, represents the steady-state moments of the UAV in three directions, represents the moment of inertia, represents the mass of the UAV; Associating the rotor rotational speed with the generalized control force through the structure of the hexacopter UAV: ; Among them, represents the relationship between the rotor speed and the generalized control force, represents the distribution matrix, represents the angular velocity of the rotation of the first rotor, represents the angular velocity of the rotation of the second rotor, represents the angular velocity of the rotation of the third rotor, represents the angular velocity of the rotation of the fourth rotor, represents the angular velocity of the rotation of the fifth rotor, represents the angular velocity of the rotation of the sixth rotor; Convert the three-dimensional velocity command to the desired acceleration: ; wherein, represents the desired acceleration, represents the control period, represents the gravitational acceleration transformed into the body coordinate system; Calculating the required generalized control force through the dynamic inverse model: ; Using the pseudo-inverse of the distribution matrix to solve for the squared rotational speeds of each rotor: ; Obtaining the actual rotational speed by taking the square root: 。 7. The visual navigation method of the six-rotor fully-driven drone according to claim 1, characterized in that, The ARM processor is used for dynamic vision feature extraction and time-series modeling operations.
8. A vision navigation system for a six-rotor fully-driven unmanned aerial vehicle, characterized in that, The vision navigation system of the fully actuated hexacopter UAV includes: A data acquisition and preprocessing module, which is used to collect the pixel depth map, real-time pose quaternion, and forward velocity during the flight of the fully actuated hexacopter UAV, perform invalid pixel filling processing and pixel value normalization processing on the pixel depth map to obtain a preprocessed depth map sequence, and transmit the depth map sequence, the real-time pose quaternion, and the forward velocity to the ARM processor; A dynamic vision feature extraction and time-series modeling module, which is used to input the depth map sequence into a vision encoder, divide each frame of image into a patch sequence of a preset size of pixels through image embedding operations, capture the global obstacle spatial association through the self-attention layer, output a multi-scale feature vector, input the multi-scale feature vectors of consecutive multiple frames into a bidirectional LSTM network to model the time-series dependence relationship during the high-speed movement of the UAV, and generate an environmental state vector including the time dimension; A navigation control coordination and command generation module, which is used to input the environmental state vector, the real-time pose quaternion, and the forward velocity into the MPC controller, with the goal of minimizing the trajectory tracking error and energy consumption, optimize the three-dimensional velocity command within a preset control period, and based on the fully actuated dynamics model of the hexacopter UAV, convert the optimized three-dimensional velocity command into the independent rotational speeds of 6 rotors, and match the fuselage acceleration and attitude angle changes through a non-linear mapping function to achieve motion decoupling control.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a vision navigation program of the fully actuated hexacopter UAV stored on the memory and executable on the processor. When the vision navigation program of the fully actuated hexacopter UAV is executed by the processor, it realizes the steps of the vision navigation method of the fully actuated hexacopter UAV as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a visual navigation program for a six-rotor fully-driven unmanned aerial vehicle. When the visual navigation program for the six-rotor fully-driven unmanned aerial vehicle is executed by a processor, the steps of the visual navigation method for the six-rotor fully-driven unmanned aerial vehicle according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Dynamic obstacle environment navigation method and device based on visual semantic information
CN111367318A
Flight control method, device, aircraft, system, and storage medium
US20230280745A1
Cited By
Infrared focal plane array attitude estimation method and device
CN120579150A
Infrared focal plane array attitude estimation method and device
CN120579150B
Unmanned aerial vehicle visual active tracking method and system for cross-category targets
CN121884203A
A Visual Active Tracking Method and System for Cross-Category Targets in Unmanned Aerial Vehicles
CN121884203B
Visual inertial navigation fused unmanned aerial vehicle trajectory tracking control method and system
CN121995938A