Power line inspection unmanned aerial vehicle pose estimation method and device based on multi-modal data fusion
Through multimodal data fusion and adaptive weight calculation methods, the problem of low positioning accuracy and stability of VINS system in complex environments is solved, and higher positioning accuracy and robustness are achieved, ensuring the stability and reliability of power line patrol drones.
Patent Information
- Application Number
- CN202510197490.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
The existing VINS systems have low positioning accuracy and stability in complex environments, poor environmental adaptability, poor robustness, and difficult to adapt to light changes and occlusion problems, and the positioning accuracy will be reduced during long-term inspection tasks.
The position estimation method of power line patrol drone based on multimodal data fusion is adopted. By obtaining the detection data of binocular vision sensors, IMU and TOF sensors, flight altitude initialization and scale alignment of monocular odometers, the adaptive weight is calculated and the position estimation is performed based on the adaptive extended Kalman filtering algorithm, and the position estimation results of monocular and binocular odometers are fused.
It improves the positioning accuracy, efficiency and stability of the drone in complex environments, enhances environmental adaptability and robustness, reduces positioning drift, and ensures the reliability and stability of patrol tasks.
Smart Images

Figure CN120141457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power line inspection, and in particular to a method and device for estimating the pose of an unmanned aerial vehicle (UAV) for power line inspection based on multi-modal data fusion. Background Art
[0002] Applying UAVs to power line inspection tasks, especially the inspection of overhead power lines in mountainous areas, can improve the inspection efficiency and accuracy, quickly and accurately complete the inspection tasks in high-altitude or dangerous areas, and at the same time can significantly reduce costs and save a large amount of human resources. Currently, UAVs usually use a Visual Inertial Navigation System (VINS) to achieve self-positioning and attitude estimation. In the VINS system, positioning and mapping are achieved by combining an Inertial Measurement Unit (IMU) and a visual sensor (such as a monocular or binocular camera). The visual sensor provides rich environmental texture information, and the IMU provides high-frequency acceleration and angular velocity measurements. The data of the visual sensor and the IMU are fused to improve the positioning accuracy and robustness.
[0003] However, currently, the VINS system only simply fuses the data of the visual sensor and the IMU, and usually assigns fixed weights to the visual sensor and the IMU to achieve data fusion. When a UAV based on such a VINS system is applied to perform power line inspection tasks, due to the inability to dynamically adjust the weights according to environmental conditions, the positioning accuracy and stability are not high. For example, in a scene where visual feature points are sparse or the lighting is insufficient, the reliability of the visual sensor will decrease, and the integration error of the IMU data will also accumulate over time. At the same time, the VINS system based on fixed weights also has problems such as poor environmental adaptability and weak robustness, and it is difficult to adapt to changes in environmental conditions and complex scenes. Since UAVs may need to fly in various complex terrains such as complex mountainous areas, problems such as lighting changes and occlusions may occur during the execution of inspection tasks, which will further reduce the positioning accuracy and reliability. In addition, in a long-term inspection task, due to the difficulty of reducing the cumulative error through global optimization or loop detection, the VINS system based on fixed weights will further reduce the positioning accuracy, and thus it is actually difficult to quickly and accurately achieve the stable positioning of the inspection UAV in a complex environment. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: aiming at the technical problems existing in the prior art, the present invention provides a method and device for estimating the pose of an unmanned aerial vehicle for power line inspection based on multi-modal data fusion, which has a simple implementation method, low cost, high estimation accuracy and stability, and strong environmental adaptability and robustness. It can improve the positioning accuracy, efficiency and stability of the power line inspection UAV in various complex environments, and at the same time effectively enhance the environmental adaptability and robustness of the positioning.
[0005] To solve the above technical problems, the technical solution proposed by the present invention is as follows:
[0006] A method for estimating the pose of an unmanned aerial vehicle (UAV) for power line inspection based on multi-modal data fusion, the steps including:
[0007] During the operation of the UAV for power line inspection, obtain the detection data of the binocular vision sensor, IMU, and TOF sensor to obtain multi-modal detection data;
[0008] Use the TOF sensor data for flight altitude initialization and scale alignment of the monocular odometer, where the monocular odometer is an odometer based on a monocular vision sensor, and the binocular IMU odometer is an odometer based on a binocular vision sensor and IMU;
[0009] Determine the linear velocity reference value and angular velocity reference value according to the poses obtained at two consecutive moments by the binocular IMU odometer and the detection data of the IMU, and calculate the adaptive weights of the binocular vision sensor and IMU according to the determined linear velocity reference value, angular velocity reference value, and the multi-modal detection data;
[0010] Perform pose estimation based on the adaptive extended Kalman filter algorithm according to the detection data of the binocular vision sensor, IMU, and the calculated adaptive weights to obtain the pose estimation result of the binocular IMU odometer;
[0011] Fuse the pose estimation results of the monocular odometer and the binocular IMU odometer to obtain the finally corrected UAV pose.
[0012] Further, using the TOF sensor data for flight altitude initialization and scale alignment of the monocular odometer includes:
[0013] When the UAV for power line inspection starts, perform a lifting action in place, and use the ground distance measured by the TOF sensor to determine the lifting distance to initialize the flight altitude;
[0014] Align the TOF sensor with the monocular odometer and the binocular IMU odometer, calculate the ratio of the Z-axis height to the ground distance to obtain the scaling factors of the monocular odometer and the binocular IMU odometer;
[0015] Scale the data of the monocular odometer and the binocular IMU odometer respectively according to the scaling factors of the monocular odometer and the binocular IMU odometer, so that the TOF sensor data is aligned with the data of the monocular odometer and the binocular odometer in scale.
[0016] Further, the determining the linear velocity reference value and angular velocity reference value according to the poses obtained at two consecutive moments by the binocular IMU odometer and the detection data of the IMU includes:
[0017] Calculate the linear velocity and angular velocity of the odometer at the current moment using the current pose of the binocular IMU odometer and the pose at the previous moment;
[0018] Obtain the linear velocity information of the IMU at the current moment according to the acceleration information of the IMU at the current moment;
[0019] Calculate the weight coefficient of the IMU and the weight coefficient of the odometer respectively according to the linear velocity information of the IMU at the current moment;
[0020] According to the linear velocity of the odometer at the current moment and the linear velocity information of the IMU at the current moment, determine the linear velocity reference value and the angular velocity reference value respectively using the weight coefficient of the IMU and the weight coefficient of the odometer.
[0021] Further, calculate the linear velocity and angular velocity of the odometer at the current moment using the current pose of the binocular IMU odometer and the pose at the previous moment according to the following formula;
[0022]
[0023] Among them, is the linear velocity of the odometer at the current moment, and are the position information at time t and time t - 1, Δt is the unit time difference between two frames, r t and r t-1 are the rotation matrices of two frames, is the angular velocity of the odometer at time t, log(Δr) is the logarithmic mapping of the rotation matrix difference;
[0024] Integrate according to the acceleration information of the IMU at the current moment to obtain the linear velocity information of the IMU at the current moment, and the calculation formula is:
[0025]
[0026] Among them, is the linear velocity of the IMU at time t, is the linear velocity at time t - 1, Δt is the unit time difference between two frames, is the IMU acceleration at time t;
[0027] Calculate the weight coefficient of the IMU and the weight coefficient of the odometer respectively according to the following formula according to the linear velocity information of the IMU at the current moment:
[0028]
[0029] Among them, and are the weight coefficient of the IMU and the weight coefficient of the odometer respectively, and are the set comparison linear velocity and comparison angular velocity;
[0030] Use the weight coefficients of the IMU and the weight coefficients of the odometer to respectively determine the linear velocity reference value and the angular velocity reference value according to the following formula:
[0031]
[0032] where and are the linear velocity reference value and the angular velocity reference value respectively.
[0033] Furthermore, the expressions for the adaptive weights of the binocular vision sensor and the IMU are:
[0034]
[0035] w visual = 1 - w IMU ;
[0036]
[0037] where w visual is the adaptive weight of the binocular vision sensor, w IMU is the adaptive weight of the IMU, v max and ω max are the preset maximum linear velocity and maximum angular velocity respectively, α is the velocity influence factor, β is the angular velocity influence factor, ρ dep is the depth influence factor, is the average point cloud depth, d ref is the preset comparison depth, and are the linear velocity reference value and the angular velocity reference value respectively.
[0038] Furthermore, the fusion of the pose estimation results of the monocular odometer and the binocular IMU odometer to obtain the finally corrected UAV pose includes:
[0039] Map the initial pose estimation result of the monocular odometer to the body IMU coordinate system through the dynamic conversion relationship of the dual IMU to obtain the pose estimation result of the monocular odometer;
[0040] Weight the pose estimation result of the monocular odometer and the pose estimation result of the binocular IMU odometer to obtain the finally corrected UAV pose.
[0041] Furthermore, weight the pose estimation result of the monocular odometer and the pose estimation result of the binocular IMU odometer according to the following formula to obtain the finally corrected UAV pose:
[0042] Tf w(t) = w V · T(t) V w(t) + w VI · T(t) VI w(t);
[0043]
[0044] w VI w(t) = 1 - w V w(t);
[0045] where μ and θ are adjustment factors, and are the linear velocity reference value and the angular velocity reference value respectively, T V (t) is the monocular odometer pose estimation result at time t, T VI (t) is the binocular odometer pose estimation result at time t, w VI (t), w V (t) are weights.
[0046] Furthermore, the IMU includes a body IMU and a gimbal IMU. After obtaining the multi-modal detection data, it further includes aligning the multi-modal detection data in time and space, including:
[0047] In terms of time, taking the time obtained by one of the binocular vision sensor, IMU, and TOF sensor as a reference, obtaining the timestamps of the detection data of each sensor, and aligning the data according to the specified delay time;
[0048] In terms of space, performing single-body calibration and joint calibration on the binocular vision sensor and the body IMU to obtain a static displacement matrix; performing single-body calibration and joint calibration on the binocular vision sensor and the gimbal IMU to obtain a static displacement matrix; converting the body IMU data into three-axis rotation angles, forming a dynamic conversion relationship according to the dynamic rotation matrix of the dual IMUs, and measuring the relative distance between the TOF sensor and the body IMU.
[0049] A power line inspection UAV pose estimation device based on multi-modal data fusion includes a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to perform the method as described above.
[0050] A computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the method as described above.
[0051] Compared with the prior art, the advantages of the present invention are as follows: The present invention realizes the pose estimation of the UAV through multi-modal data fusion and dynamic weighting, can adapt to different environmental changes, fully fuse the data of multi-modal sensors, automatically adjust the weights by comprehensively considering the advantages and disadvantages of different sensors in a specific environment, effectively optimize the state estimation at adjacent moments, reduce the influence caused by sensor failure or noise, ensure the stable positioning of the UAV in a complex environment, effectively improve the positioning accuracy and robustness of the UAV in complex dynamic inspection tasks, reduce the positioning drift caused by sensor limitations, and at the same time obtain the finally corrected UAV pose by fusing the pose estimation results of the monocular odometer and the binocular IMU odometer, and can still support the normal operation of the system in case of any sensor failure, thereby significantly improving the reliability and stability of the inspection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 FIG. is a schematic flow chart of the implementation of the method for estimating the pose of a power line inspection UAV based on multi-modal data fusion in this embodiment.
[0053] Figure 2 FIG. is a schematic diagram of the principle for implementing pose estimation in this embodiment.
[0054] Figure 3 FIG. is a schematic diagram for comparing the test effects of using the traditional estimation algorithm and the present invention respectively for estimating the pose of a power line inspection UAV in a specific application embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0056] For ease of understanding, first, an introduction to the relevant technical background related to the present invention is given by way of example.
[0057] Currently, UAVs usually adopt the VINS system to achieve positioning and attitude estimation by combining an IMU and a vision sensor (such as a monocular or binocular camera). However, the VINS system only simply fuses the data of two sensors, namely the vision sensor and the IMU, and uses a fixed weight allocation method to achieve data fusion, which will have problems such as low positioning accuracy and stability, poor environmental adaptability, and weak robustness. With the execution of inspection tasks for a long time, the positioning accuracy will be further reduced.
[0058] Specifically, the current VINS system based on an IMU and a vision sensor will have the following problems in the process of positioning and attitude estimation:
[0059] 1. Low positioning accuracy and stability
[0060] The VINS system only fuses the data of two sensors, namely the visual sensor and the IMU, and realizes data fusion by using a fixed weight allocation method. It cannot dynamically adjust the weights of the sensors according to environmental conditions, which may lead to low positioning accuracy and stability, especially poor positioning accuracy and stability in extreme or complex inspection environments. In complex environments,
[0061] 2. Poor environmental adaptability and robustness, unable to cope with long-term flight or extreme environments
[0062] In complex environments such as low-light and dynamic environments, the VINS system is easily affected by environmental changes. The visual inertial positioning system of the UAV itself has a fixed weight and poor adaptability to environmental changes. The visual inertial positioning system with a fixed weight cannot adaptively adjust the weights of the sensors under different environmental conditions, resulting in inaccurate position estimation. For example, due to factors such as terrain undulation, vegetation coverage, and light changes, the visual sensor may encounter interference such as uneven illumination and obstacles, making the positioning accuracy unstable. During long-term power inspection tasks, the IMU may have problems with error accumulation after long-term flight. Currently, the VINS system lacks sufficient robustness to cope with environmental complexity and long-term error accumulation.
[0063] 3. Sensitivity to sensor failures
[0064] When the sensors in the VINS system fail or the data is abnormal, there will also be problems with ineffective compensation. For example, if the IMU or the visual sensor fails, the system is difficult to quickly correct itself, which may lead to a sharp drop in positioning accuracy and even the inability to complete the inspection task.
[0065] 4. Relocalization problems
[0066] In an environment where GPS is unavailable or the signal is weak, the VINS system needs to rely on the internal data of the sensors for positioning and attitude estimation. However, due to sensor errors and local environmental changes, the VINS system is prone to positioning drift or relocalization failure, which affects the continuity and accuracy of the inspection task, and further affects the positioning accuracy and stability.
[0067] Through multi-modal data fusion based on the data of binocular vision sensors, IMUs, and TOF sensors, and by adopting a dynamic weighting method to fuse the data of different sensors, the present invention can adapt to different environmental changes, fully fuse the data of multi-modal sensors, automatically adjust the weights by comprehensively considering the advantages and disadvantages of different sensors in a specific environment, effectively optimize the state estimation at adjacent moments, reduce the influence caused by sensor failures or noises, ensure the stable positioning of the UAV in complex environments, effectively improve the positioning accuracy and robustness of the UAV in complex dynamic inspection tasks, reduce the positioning drift caused by sensor limitations. At the same time, by fusing the pose estimation results of the monocular odometer and the binocular IMU odometer, the finally corrected UAV pose is obtained, which can still support the normal operation of the system in case of any sensor failure, ensure the reliability and stability of the inspection task, be applicable to various complex application scenarios such as mountain power line inspection, and can significantly improve the positioning accuracy, stability, and robustness of the UAV.
[0068] As Figure 1 shown, the steps of the power line inspection UAV pose estimation method based on multi-modal data fusion in this embodiment specifically include:
[0069] Step S01. During the operation of the line inspection UAV, obtain the detection data of the binocular vision sensor, IMU, and TOF sensor to obtain multi-modal detection data.
[0070] Specifically, the line inspection UAV is equipped with a binocular vision sensor, an airframe IMU, and a TOF sensor. During the operation of the line inspection UAV, the objects in the surrounding environment are detected by the binocular vision sensor, airframe IMU, and TOF sensor of the UAV, multi-sensor data is obtained, and multi-modal detection data is obtained. After obtaining the multi-modal detection data, operations such as filtering and denoising can be further performed on the data of each sensor to reduce errors and noises.
[0071] In complex power inspection environments such as mountain forests, GPS signals cannot cover or the signals are weak. The TOF sensor can effectively measure the flight altitude in such complex environments, provide an accurate initial value for UAV visual inertial positioning, avoid the errors that may be brought by traditional altitude measurement methods, and ensure the flight safety of the UAV in complex terrains.
[0072] Furthermore, after obtaining the multi-modal detection data, it also includes performing spatio-temporal alignment on the multi-modal detection data, that is, aligning the data of each sensor in terms of time and space, ensuring that the data of each sensor can be fused and calculated within the same time window. Through precise spatio-temporal alignment, the accuracy of the initial pose and scale can be ensured, and at the same time, the interference of error sources such as GPS can be reduced to ensure the accuracy in complex environments.
[0073] As an alternative embodiment, aligning the sensor data in terms of time and space includes:
[0074] In terms of time, taking the time obtained by one of the binocular vision sensor, IMU, and TOF sensor as a reference, obtaining the timestamps of the detection data of each sensor, and aligning the data according to the specified delay time;
[0075] The IMU includes a body IMU and a gimbal IMU. In terms of space, performing individual calibration and joint calibration on the binocular vision sensor and the body IMU to obtain a static displacement matrix; performing individual calibration and joint calibration on the binocular vision sensor and the gimbal IMU to obtain a static displacement matrix; converting the body IMU data into three-axis rotation angles (i.e., yaw angle, pitch angle, and roll angle), constructing a dynamic conversion relationship based on the dynamic rotation matrix of the dual IMUs, and measuring the relative distance between the TOF sensor and the body IMU.
[0076] For example, in terms of time, the timestamps of each data can be obtained with the time of obtaining binocular sensor data as a reference, and the data can be aligned with a delay time of 50 ms. In terms of space, performing individual calibration and joint calibration on the binocular vision sensor and the body IMU to obtain a displacement matrix. The static displacement matrix is the matrix of the fixed relative position and attitude relationship between the binocular vision sensor and the body IMU, and the static conversion relationship between the binocular vision sensor and the body IMU can be determined; performing individual calibration and joint calibration on the body IMU and the gimbal IMU, setting the gimbal to the follow mode, fixing the rotation angle at the same time, performing external parameter calibration using the gimbal monocular camera and the body IMU to obtain a static displacement matrix, converting the body IMU data into three-axis rotation angles, combining the dynamic rotation matrix of the dual IMUs to form a dynamic conversion relationship. The dynamic rotation matrix is the matrix used to describe the rotation state of the UAV in three-dimensional space in a dynamic system. The dynamic rotation matrix can be constructed through the three-axis rotation angles, and then the dynamic conversion relationship between the dual IMUs can be obtained according to the dynamic rotation matrix. Finally, the relative distance between the TOF sensor and the body IMU is measured, and then all modal data is published to the topic using ROS messages.
[0077] Step S02. Using the TOF sensor data for flight altitude initialization and scale alignment of the monocular odometer. The monocular odometer is an odometer based on a monocular vision sensor, and the binocular IMU odometer is an odometer based on a binocular vision sensor and an IMU.
[0078] In this embodiment, by using the TOF sensor data for flight altitude initialization and scale alignment of the monocular odometer, the accuracy of the initial pose and scale can be further ensured, and the interference of error sources can be reduced.
[0079] As an alternative implementation, using TOF sensor data for flight altitude initialization and scale alignment of the monocular odometer includes:
[0080] Step S201. When the line inspection UAV starts, perform a lifting action in place, and use the ground distance of the TOF sensor to determine the lifting distance to initialize the flight altitude;
[0081] Step S202. Use the TOF sensor data to align the monocular odometer and the binocular IMU odometer, calculate the ratio of the Z-axis height to the ground distance, and obtain the scaling factors of the monocular odometer and the binocular IMU odometer;
[0082] Step S203. Scale the data of the monocular odometer and the binocular IMU odometer respectively according to the scaling factors of the monocular odometer and the binocular IMU odometer, so as to align the scale of the TOF sensor data with the data of the monocular odometer and the binocular odometer.
[0083] For example, according to the ratio of the Z-axis height to the ground distance of the monocular odometer (calculated using a visual SLAM module such as ORB-SAM3) and the binocular IMU odometer, the scaling factor M of the monocular odometer and the scaling factor S of the binocular IMU odometer can be obtained respectively. Then, scale the binocular IMU odometer data by S times and the monocular odometer data by M times to align the monocular odometer and the binocular IMU odometer, and then publish it with a new ROS topic.
[0084] Step S03. Determine the linear velocity reference value and the angular velocity reference value according to the poses obtained at two consecutive moments of the binocular IMU odometer and the detection data of the IMU, and calculate the adaptive weights of the binocular vision sensor and the IMU according to the determined linear velocity reference value, angular velocity reference value and multi-modal detection data.
[0085] In this embodiment, by using the poses obtained at two consecutive moments of the binocular IMU odometer and the detection data of the IMU, the linear velocity reference value and the angular velocity reference value are determined. Furthermore, according to the determined linear velocity reference value, angular velocity reference value and multi-modal detection data, the adaptive weights of the binocular vision sensor and the IMU are calculated. The weights can be dynamically adjusted in real time according to the multi-modal sensor data (binocular vision, IMU), and the sensor fusion can be flexibly optimized, thereby improving the positioning accuracy, environmental adaptability and robustness of the system.
[0086] As an alternative implementation, the current linear velocity and angular velocity can be calculated first by using the poses obtained at two consecutive moments, and then dynamically weighted and fused with the linear velocity and angular velocity obtained by integrating the current IMU to determine the linear velocity reference value and the angular velocity reference value. The steps include:
[0087] Step S211. Calculate the linear velocity and angular velocity of the odometer at the current moment using the current pose of the binocular IMU odometer and the pose at the previous moment;
[0088] Step S212. Obtain the linear velocity information of the IMU at the current moment according to the acceleration information of the IMU at the current moment;
[0089] Step S213. Calculate the weight coefficient of the IMU and the weight coefficient of the odometer respectively according to the linear velocity information of the IMU at the current moment;
[0090] Step S214. Determine the linear velocity reference value and the angular velocity reference value respectively using the weight coefficient of the IMU and the weight coefficient of the odometer according to the linear velocity of the odometer at the current moment and the linear velocity information of the IMU at the current moment.
[0091] Specifically, the odometer data O at time t-1 t-1 and the odometer data O at time t t can be used to calculate the linear velocity and angular velocity of the odometer at time t. Then, the linear velocity information of the IMU can be obtained by integrating the acceleration information of the IMU at time t and angular velocity information The angular velocity data of the IMU can be directly read. Finally, the weight is dynamically adjusted according to the data changes of the IMU and the odometer to obtain the final linear velocity and angular velocity reference values and
[0092] For example, the current pose of the binocular IMU odometer and the pose at the previous moment can be used to calculate the linear velocity and angular velocity of the odometer at the current moment according to the following formula;
[0093]
[0094] where is the linear velocity of the odometer at the current moment, and are the position information at time t and time t-1, Δt is the unit time difference between two frames, r t and r t-1 are the rotation matrices of two frames, is the angular velocity of the odometer at time t, and log(Δr) is the logarithmic mapping of the rotation matrix difference.
[0095] Integrate according to the acceleration information of the IMU at the current moment to obtain the linear velocity information of the IMU. The calculation expression can be expressed as:
[0096]
[0097] where is the linear velocity of the IMU at time t, is the linear velocity at time t-1, and Δt is the unit time difference between two frames, is the IMU acceleration at time t.
[0098] Then, according to the IMU linear velocity information at the current moment, the weight coefficients of the IMU and the weight coefficients of the odometer can be calculated respectively according to the following formula:
[0099]
[0100] Among them, and are the weight coefficients of the IMU and the weight coefficients of the odometer respectively, and are the set comparison linear velocity and comparison angular velocity.
[0101] Finally, use the weight coefficients of the IMU and the weight coefficients of the odometer to determine the linear velocity reference value and the angular velocity reference value respectively according to the following formula:
[0102]
[0103] Among them, and are the linear velocity reference value and the angular velocity reference value respectively.
[0104] As shown in formulas (7) and (8), the larger it is, the greater the weight it occupies when calculating the linear velocity and angular velocity reference values. Preferably, when is greater than , can be set to 1, and finally the data of the IMU and the binocular vision sensor can be fully fused to accurately obtain the fused linear velocity reference value and the angular velocity reference value
[0105] Through the above steps, use the current pose and the previous pose of the binocular IMU odometer to calculate the current linear velocity and angular velocity, and perform dynamic weighted fusion with the linear velocity and angular velocity obtained by IMU integration, so as to dynamically adjust the weight according to the data changes of the IMU and the odometer to accurately obtain the final linear velocity and angular velocity reference values, that is, the velocity reference values are calculated based on the multi-modal data (IMU data and image data) at the previous time, and can dynamically adjust the velocity reference values in real time considering the data changes. Specifically, the multi-modal data in this embodiment includes the body binocular camera images and the body IMU, the gimbal camera images and the gimbal IMU, and the pose, linear velocity, angular velocity, point cloud depth and other information can be calculated using the above multi-modal data.
[0106] As an alternative implementation, based on the determined linear velocity reference value and the angular velocity reference value Furthermore, a Visual SLAM (Simultaneous Localization and Mapping) module can be used to obtain sparse point clouds by methods such as ORB-SLAM3. The depth influence factor can be dynamically calculated based on the obtained sparse point cloud depth values. Then, the adaptive weights of binocular vision and IMU can be calculated according to the changes in multi-modal data, and the weights can be dynamically adjusted according to the real-time state of the multi-modal data, enabling the system to optimize the use of visual sensors in different depth and depth-of-field environments, thereby further improving the environmental adaptability and robustness of the estimation.
[0107] For example, the average depth can be calculated using the sparse point cloud obtained by the Visual SLAM module, and the point cloud information within a specified range (such as a 30-degree field of view up and down) can be intercepted to calculate the average point cloud depth Then calculate the average point cloud depth Compare with the preset reference depth d ref to obtain the depth influence factor Furthermore, based on this depth influence factor ρ dep an adaptive weight function is established to obtain the weight w of the binocular vision sensor visual and the weight w of the IMU IMU .
[0108] As an alternative implementation, the adaptive weights of the binocular vision sensor and the IMU can be calculated using the following adaptive weight function:
[0109]
[0110] W visual = 1 - w IMU (10)
[0111] where w visual is the adaptive weight of the binocular vision sensor, w IMU is the adaptive weight of the IMU, v max and ω max are the preset maximum linear velocity and maximum angular velocity respectively, α is the velocity influence factor, β is the angular velocity influence factor, ρ dep is the depth influence factor, is the average point cloud depth, d ref is the preset reference depth, and are the linear velocity reference value and the angular velocity reference value respectively.
[0112] As shown in Equation (9), if the current linear velocity is large, the weight of the IMU will increase, and the weight of the binocular vision sensor will decrease accordingly; if the current linear velocity is small, the weight of the binocular vision sensor will increase, and the weight of the IMU will decrease. The same applies to the angular velocity. The weights of the sensors can be dynamically adjusted according to the magnitude of the current linear velocity to adapt to environmental changes. Preferably, when the angular velocity is greater than the maximum angular velocity ω max , w IMU can be configured to be 1 to increase or decrease the observation noise covariance matrix R visual and R imu , thereby indirectly affecting the update of the state covariance matrix P. By adopting the above adaptive weight function, the speed reference value will be dynamically adjusted according to the multi-modal data of the previous time, so that the adaptive weights of the binocular vision sensor and the IMU can be accurately determined according to the changes in the multi-modal data.
[0113] Step S04. Based on the detection data of the binocular vision sensor and the IMU and the calculated adaptive weights, perform pose estimation using the adaptive extended Kalman filter algorithm (AEKF) to obtain the pose estimation result of the binocular IMU odometer.
[0114] In this embodiment, after determining the adaptive weights w visual , w IMU of the binocular vision sensor and the IMU, perform pose estimation by combining the extended Kalman filter with adaptive weights, which can effectively fuse visual and IMU data in a dynamic environment, improve the anti-interference ability, and ensure the stability of the system in a high-dynamic scenario.
[0115] Specifically, when calculating the pose of the drone using the extended Kalman filter combined with adaptive weights, the adaptive weights w visual , w IMU of the binocular vision sensor and the IMU are introduced. Use the extended Kalman filter for pose estimation. Utilize the existing state estimation and prediction models, combine the new observation data, update the estimated value of the system state using the adaptive weights, and simultaneously calculate the corresponding estimation error, and output the optimized pose trajectory T VI (t) of the binocular IMU odometer. The specific calculation steps are as follows:
[0116] Step S401. Define the system model
[0117] Establish the system dynamic model of the target, including the state variables x k (position p, velocity v, attitude q, bias b a of the accelerometer, bias b g of the gyroscope), observation value z k (sensor observation value), input quantity u k, process noise w k and observation noise v k etc.
[0118] For example, the system model can be as follows:
[0119] x k = F k x k-1 + Bu k-1 + w k (11)
[0120] z k = Hx k + v k (12)
[0121] where F is the state transition matrix, B is the control input matrix, and H is the observation matrix.
[0122] Step S402. Initialize the Kalman filter
[0123] Set the initial state estimate and the state covariance matrix P 0 , where the initial state estimate can be obtained from the initial observations of the IMU, and the state covariance matrix P 0 represents the uncertainty of the initial state and is a diagonal matrix.
[0124] Step S403. Prediction
[0125] Based on the system model, use the state estimate at the previous time to make a prediction to obtain the state prediction value at the current time and the state covariance P k|k-1 :
[0126]
[0127] P k|k-1 = FP k-1|k-1 F T + Q (14)
[0128] In the above formula, Q is the process noise covariance matrix.
[0129] Step S404. Update
[0130] Compare the sensor observations with the prediction results and calculate the Kalman gain K k , which determines the weight between the predicted value and the observed value and the degree of correction to the state estimate.
[0131] Weighing the predicted value and the observed value can be expressed as:
[0132] Kk = P k|k-1 H T (HP k|k-1 H T + R) -1 (15)
[0133] In the above formula, R is the covariance matrix of the observation noise.
[0134] Step S405. Correction
[0135] The predicted state estimate is corrected to an estimate closer to the observed value through the Kalman gain, and at the same time, the error covariance matrix is updated to reflect the accuracy of the corrected state estimate.
[0136] The Kalman gain for correcting the state estimate and state covariance can be expressed as:
[0137]
[0138] P k|k = (I - K k H)P k|k-1 (17)
[0139] Step S406. Repeat prediction and update.
[0140] At each time step, prediction and update are iteratively performed, and the state estimate is continuously corrected using the latest observation data. Through the above steps, the Kalman filter can fuse the observation data of multiple sensors, comprehensively consider the characteristics of each sensor and the observation error, so as to obtain an accurate target state estimate and dynamically update the reliability of the estimate (state covariance P k|k ), and finally output the accurate state variable x k , and obtain the pose trajectory data of the binocular IMU odometer.
[0141] As Figure 2 shown, in this embodiment, the pose at time T2 is calculated using the poses at time T1 and T2. At the same time, there is an observed velocity of the IMU at time T2. The reference velocity at time T2 is determined using the dynamic reference weight. The adaptive weight function is calculated by combining the depth influence factor at time T2, and the observation weights of the binocular and IMU are determined to calculate the pose at T3. Among them, time T1 is the initial value and is at the origin. The pose at T2 is calculated using the extended Kalman filter with an average weight. After obtaining the dynamic weight and adaptive weight function, the pose at time T3 is calculated using the adaptive extended Kalman filter (AEKF). Subsequently, the AEKF can be used for pose calculation.
[0142] Step S05. Fuse the pose estimation results of the monocular odometer and the binocular IMU odometer to obtain the finally corrected pose of the drone.
[0143] In this embodiment, the trajectory is independently estimated by the monocular system and the visual inertial system, and then the final result is obtained through trajectory alignment, correction, and weighted fusion, which can not only ensure the flexibility of the system but also give full play to the respective advantages of the monocular odometer and the binocular IMU odometer.
[0144] As an alternative embodiment, fusing the pose estimation results of the monocular odometer and the binocular IMU odometer to obtain the finally corrected UAV pose includes:
[0145] Step S501. Map the initial pose estimation result of the monocular odometer to the body IMU coordinate system through the dynamic conversion relationship of the dual IMUs to obtain the pose estimation result of the monocular odometer;
[0146] Step S502. Weight the pose estimation result of the monocular odometer and the pose estimation result of the binocular IMU odometer to obtain the finally corrected UAV pose.
[0147] As an alternative embodiment, the weight w V (t) of the monocular odometer and the weight w VI (t) of the binocular IMU odometer can be calculated according to the flight environment and flight state, so that at high speeds, the weight of the binocular IMU odometer is increased, and at low speeds and in a stationary state, the weight of the monocular odometer is increased. Then, the weighted average method is used to fuse the two trajectories to finally obtain the corrected pose trajectory T final (t).
[0148] For example, the weight w V (t) of the monocular odometer and the weight w VI (t) of the binocular IMU odometer can be determined according to the following formula:
[0149]
[0150] w VI (t) = 1 - w V (t) (19)
[0151] where μ and θ are adjustment factors, and are the linear velocity reference value and the angular velocity reference value respectively.
[0152] Furthermore, the pose estimation result of the monocular odometer and the pose estimation result of the binocular IMU odometer are weighted to obtain the finally corrected UAV pose, that is:
[0153] T f (t) = w V (t) · T V (t) + wVI (t)·T VI (t) (20)
[0154] wherein, T V (t) is the monocular odometer pose estimation result at time t, and T VI (t) is the binocular odometer pose estimation result at time t.
[0155] By fusing the pose trajectories of the monocular odometer and the binocular IMU odometer in the above manner to correct the final estimation result, the normal operation of the system can still be supported in the case of the failure of any trajectory calculation.
[0156] To verify the effect of the present invention, in a specific application embodiment, the extended Kalman filter method that combines traditional IMU high-frequency information and binocular vision and the above method of the present invention are respectively used to implement the pose estimation of the power inspection unmanned aerial vehicle. As Figure 3 shown, it can be seen from the results that the traditional extended Kalman filter method is prone to tracking failure when the visual data weakens, and thus performs a relocalization operation, causing the trajectory to terminate and return to the initial value. However, due to the introduction of adaptive weights in the present invention to realize the fusion of multi-source sensor data, compared with the traditional fixed-weight scheme, the fusion ratio of IMU and visual information can be dynamically adjusted according to the actual situation. When the performance of the visual sensor is poor, the system will automatically reduce the weight of the visual information and increase the influence of the IMU data. When the visual conditions are good, the system will reduce the weight of the IMU data and reduce the interference of high-frequency data fluctuations. Therefore, it can effectively reduce the positioning interruption caused by relocalization and ensure the continuity of the unmanned aerial vehicle positioning. At the same time, by combining the binocular odometer and the monocular odometer, fusing the pose trajectories of the dual systems to correct the final pose result, the accuracy and robustness of the attitude estimation can also be improved, and the positioning accuracy and robustness of the unmanned aerial vehicle in the power inspection task in a complex environment can be optimized.
[0157] In summary, the present invention can dynamically adjust the weights of binocular vision and IMU according to the changes of multi-modal sensor data, adaptively increase the weights of relatively reliable sensors and reduce the influence of unreliable sensors according to the data quality, thereby optimizing the positioning accuracy. At the same time, using the extended Kalman filter algorithm combined with adaptive weights to process the fused data for pose estimation can effectively handle the non-linear characteristics of sensor data. And through the fusion of multi-sensor data, error accumulation can be reduced. By combining the monocular odometer with the binocular IMU odometer, the pose trajectories of the unmanned aerial vehicle can be jointly optimized in different flight environments and states, and when the calculation of one of the trajectories fails, it can ensure that the system can still operate normally, optimize the state estimation of the unmanned aerial vehicle in the power inspection task in a complex environment, improve the accuracy, stability and robustness of the pose estimation, and ensure that the unmanned aerial vehicle can fly stably in various complex terrains and efficiently complete the inspection task of power equipment.
[0158] This embodiment further provides a pose estimation device for an unmanned aerial vehicle (UAV) for power line inspection based on multi-modal data fusion, including a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to perform the method as described above.
[0159] It can be understood that the above method of this embodiment can be executed by a single device, such as a computer or a server, etc., or can also be applied to a distributed scenario where multiple devices cooperate with each other to complete. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps of the above method of this embodiment, and the multiple devices interact with each other to complete the above method. The processor can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., and is used to execute relevant programs to implement the above method of this embodiment. The memory can be implemented in forms such as a read-only memory (ROM), a random access memory (RAM), a static storage device, and a dynamic storage device, etc. The memory can store an operating system and other application programs. When implementing the above method of this embodiment through software or firmware, the relevant program codes are stored in the memory and are called and executed by the processor.
[0160] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0161] Those skilled in the art should understand that the above embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1Apparatus for the functions specified in one or more boxes. These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction apparatus that implements the process Figure 1 One process or multiple processes and / or boxes Figure 1 Apparatus for the functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 One process or multiple processes and / or boxes Figure 1 One or more boxes.
[0162] The above are only the preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for estimating the posture of a UAV for power line inspection based on multimodal data fusion, characterized in that the steps include: During the operation of the line inspection drone, the detection data of the binocular vision sensor, IMU and TOF sensor are obtained to obtain multimodal detection data; Use TOF sensor data to initialize the flight altitude and align the scale of the monocular odometer, where the monocular odometer is an odometer based on a monocular vision sensor, and the binocular IMU odometer is an odometer based on a binocular vision sensor and an IMU; Determine a linear velocity reference value and an angular velocity reference value according to the position and posture obtained at two moments before and after the binocular IMU odometer and the detection data of the IMU, and calculate the adaptive weights of the binocular vision sensor and the IMU according to the determined linear velocity reference value and angular velocity reference value and the multimodal detection data; Performing pose estimation based on the detection data of the binocular vision sensor and the IMU and the calculated adaptive weights based on the adaptive extended Kalman filter algorithm to obtain a pose estimation result of the binocular IMU odometer; The pose estimation results of the monocular odometer and the binocular IMU odometer are fused to obtain the final corrected UAV pose.
2. The method for estimating the posture of a power line inspection drone based on multimodal data fusion according to claim 1 is characterized in that: Using TOF sensor data to initialize flight altitude and align the scale of the monocular odometer includes: When the line inspection drone is started, it will perform a pull-up action on the spot, and use the TOF sensor's distance to the ground to determine the pull-up distance to initialize the flight altitude; Align the TOF sensor with the monocular odometer and the binocular IMU odometer, calculate the ratio of the Z-axis height and the distance to the ground, and obtain the scaling factor of the monocular odometer and the binocular IMU odometer; The data of the monocular odometer and the binocular IMU odometer are scaled according to the scaling factors of the monocular odometer and the binocular IMU odometer, respectively, so that the TOF sensor data is scale-aligned with the data of the monocular odometer and the binocular odometer.
3. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to claim 1 is characterized in that: Determining the linear velocity reference value and the angular velocity reference value according to the position and posture obtained at two moments before and after the binocular IMU odometer and the detection data of the IMU includes: Use the current position and the previous position of the binocular IMU odometer to calculate the linear velocity and angular velocity of the odometer at the current moment; Obtain the current IMU linear velocity information based on the current IMU acceleration information; Calculate the weight coefficient of the IMU and the weight coefficient of the odometer respectively according to the IMU linear velocity information at the current moment; According to the linear velocity of the odometer at the current moment and the linear velocity information of the IMU at the current moment, the linear velocity reference value and the angular velocity reference value are determined respectively using the weight coefficient of the IMU and the weight coefficient of the odometer.
4. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to claim 3 is characterized in that: Use the current position and the previous position of the binocular IMU odometer to calculate the linear velocity and angular velocity of the odometer at the current moment according to the following formula; in, is the linear speed of the odometer at the current moment, and is the position information at time t and time t-1, Δt is the unit time difference between the two frames, r t and r t-1 is the rotation matrix of the two frames, is the angular velocity of the odometer at time t, log(Δr) is the logarithmic mapping of the rotation matrix difference; According to the current acceleration information of IMU, the linear velocity information of IMU is obtained by integrating it. The calculation expression is: in, is the IMU linear velocity at time t, is the linear velocity at time t-1, Δt is the unit time difference between two frames, is the IMU acceleration at time t; According to the current IMU linear velocity information, the weight coefficient of the IMU and the weight coefficient of the odometer are calculated according to the following formulas: in, and They are the weight coefficients of IMU and odometer respectively, and is the set comparison linear velocity and comparison angular velocity; The linear velocity reference value and the angular velocity reference value are determined respectively by using the weight coefficient of the IMU and the weight coefficient of the odometer according to the following formulas: in, and Linear velocity reference value and angular velocity reference value respectively.
5. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to claim 1 is characterized in that: The expressions of the adaptive weights of the binocular vision sensor and the IMU are: In visual =1-in IMU ; Among them, w visual is the adaptive weight of the binocular vision sensor, w IMU is the adaptive weight of IMU, v max and ω max are the preset maximum linear velocity and maximum angular velocity respectively, α is the velocity influence factor, β is the angular velocity influence factor, ρ dep is the depth impact factor, is the average point cloud depth, d ref is the preset contrast depth, and are the linear velocity reference value and angular velocity reference value respectively.
6. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to any one of claims 1 to 5, characterized in that: The pose estimation results of the monocular odometer and the binocular IMU odometer are fused to obtain the final corrected UAV pose, including: The initial pose estimation result of the monocular odometer is mapped to the body IMU coordinate system through the dynamic conversion relationship of the dual IMU to obtain the monocular odometer pose estimation result; The pose estimation result of the monocular odometer is weighted with the pose estimation result of the binocular IMU odometer to obtain the final corrected UAV pose.
7. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to claim 6 is characterized in that: The pose estimation result of the monocular odometer and the pose estimation result of the binocular IMU odometer are weighted according to the following formula to obtain the final corrected UAV pose: T f (t)=w V (t)·T V (t)+w VI (t)·T VI (t); w VI (t)=1-w V (t); Among them, μ and θ are adjustment factors, and are the linear velocity reference value and angular velocity reference value respectively, T V (t) is the monocular odometer pose estimation result at time t, T VI (t) is the stereo odometry pose estimation result at time t, w VI (t), w V (t) is the weight.
8. The method for estimating the posture of a UAV for power line inspection based on multimodal data fusion according to any one of claims 1 to 5, characterized in that: IMU includes body IMU and gimbal IMU. After obtaining multi-modal detection data, it also includes aligning the multi-modal detection data in time and space, including: In terms of time, the time obtained by one of the binocular vision sensor, IMU and TOF sensor is used as the benchmark to obtain the timestamp of each sensor's detection data, and the data is aligned according to the specified delay time; In space, the binocular vision sensor and the body IMU are calibrated individually and jointly to obtain a static displacement matrix. The binocular vision sensor and the gimbal IMU are calibrated individually and jointly to obtain a static displacement matrix. The body IMU data is converted into a three-week rotation angle, and a dynamic conversion relationship is constructed according to the dynamic rotation matrix of the dual IMU to measure the relative distance between the TOF sensor and the body IMU.
9. A power line inspection drone posture estimation device based on multimodal data fusion, comprising a processor and a memory, wherein the memory is used to store a computer program, characterized in that: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Unmanned aerial vehicle flight path positioning method and system based on multi-source information fusion
CN121252823A
Camera pose regression estimation system and method based on long-time-sequence arbitrary point tracking
CN121259094A
Tracking type scanning method and system based on monocular and binocular fusion
CN121353346A