Target pose estimation method based on monocular vision and inertial navigation fusion

CN117968640BActive Publication Date: 2026-09-29XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410265306.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2026-09-29
Estimated Expiration
2044-03-08

AI Technical Summary

Technical Problem

[0006]本发明的目的是针对上述现有技术的不足,提出一种基于单目视觉与惯导融合的目标位姿估计方法,旨在解决计算机视觉算法复杂度高、耗时长,特征点检测难度大,导致的飞行器与地面可识别目标间相对位姿估计实时性低和准确性差的问题

Benefits of technology

[0015]第一,本发明使用单目相机通过拍摄的单帧图像就能对飞行器进行相对位姿估计,克服了现有技术中计算机视觉算法运行时间较长的问题,使得本发明能以更高的实时性完成飞行器的相对位姿估计。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117968640B_ABST
    Figure CN117968640B_ABST
Patent Text Reader

Abstract

The application discloses a target pose estimation method based on monocular vision and inertial navigation fusion, and the implementation steps are as follows: obtaining internal parameters of a calibrated monocular camera, shooting a target image by using the monocular camera and preprocessing the image, detecting a ground identifiable target from the preprocessed image, extracting two feature point pixel coordinates in the ground identifiable target, obtaining angle data of aircraft inertial navigation and calculating a rotation matrix of the ground identifiable target, and simultaneously solving relative poses between the camera and the ground identifiable target by using a distance formula and two sets of coordinate conversion formulas corresponding to the two feature points. The application utilizes inertial navigation attitude prior information, and only needs to obtain two feature points of the ground identifiable target in the target image, so that the relative poses between the camera on the aircraft and the ground identifiable target can be solved, the detection requirement of feature points for pose estimation application is reduced, and the real-time performance and accuracy of the relative pose estimation of the aircraft are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target tracking technology, and further relates to a target pose estimation method based on monocular vision and inertial navigation fusion in the field of visual navigation technology. This invention can be used in terminal guidance to guide the aircraft to a precise landing by estimating the relative pose between the aircraft's camera and a identifiable ground target. Background Technology

[0002] Visual navigation technology refers to the use of visual information to guide mobile devices such as robots, aircraft, and autonomous vehicles in their navigation within their environment. Compared to traditional satellite navigation and inertial navigation, visual navigation technology offers advantages such as strong autonomy, passive operation, and high guidance accuracy. This has made visual-guided aircraft landing technology an important research area both domestically and internationally. It primarily involves using cameras mounted on the aircraft to acquire images of the vicinity of the landing point, employing computer vision algorithms to estimate the aircraft's position and orientation relative to the landing point, and combining this with other sensors to guide the aircraft's precise landing.

[0003] In vision-based navigation, pose estimation is a core problem. Pose estimation refers to determining the orientation and position of a camera or ground-based identifiable target in three-dimensional space, directly affecting the accuracy and reliability of navigation. The most common pose estimation method is based on 2D keypoint projection to 3D. Among these, the classic method is PnP (Perspective-n-Point). In PnP, 'n' represents n feature points. The PnP method estimates the pose between the camera and the ground-based identifiable target given the coordinates of n 3D points and their 2D projection positions.

[0004] The Shanghai Institute of Mechanical and Electrical Engineering disclosed a visual navigation and positioning method for low-speed unmanned aerial vehicles (UAVs) in its patent application "Visual Navigation and Positioning Method and System for Low-Speed ​​Unmanned Aerial Vehicles" (Patent Application No. 2023102865823, Publication No. CN 116429098 A). The specific steps of this method are as follows: images are acquired using a visual sensor and preprocessed; feature points are extracted from the preprocessed images to obtain their image coordinates; the attitude transformation matrix information in the inertial frame is obtained based on the UAV's position, attitude, and altitude information measured by the inertial measurement unit and altimeter; the attitude and altitude information output by the flight control system are subscribed to using the attitude transformation matrix; and the spatial position of the corresponding pixels in the world's three-dimensional coordinate system is calculated using a visual sensor model. The drawback of this method is that acquiring the attitude transformation matrix information requires at least two adjacent image frames to obtain the corresponding rotation matrix and translation vector. Acquiring two image frames and processing each frame separately takes longer, resulting in low real-time positioning performance.

[0005] In his paper "Research on Monocular Vision Pose Measurement Technology for Cooperative Targets in Space" (Master's Thesis, University of Chinese Academy of Sciences, 2018), Lü Yaoyu proposed a method for aircraft pose estimation based on cooperative targets. This method designs and identifies cooperative targets, quickly obtains feature points using target recognition algorithms, locates these feature points using the encoded information designed within the cooperative targets, and finally solves the aircraft pose using the PnP algorithm. The method has limitations. The PnP algorithm used for pose estimation has a limitation on the number of feature points; it requires a minimum of three. This limitation increases the difficulty and time required for feature point detection, affecting the real-time performance and accuracy of aircraft pose estimation. Furthermore, the PnP method typically requires iterative solutions, resulting in high computational complexity and long algorithm execution time, which also impacts the real-time performance of aircraft pose estimation. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of the existing technology by proposing a target pose estimation method based on the fusion of monocular vision and inertial navigation. This method aims to solve the problems of low real-time performance and poor accuracy in relative pose estimation between aircraft and identifiable ground targets, which are caused by the high complexity and long processing time of computer vision algorithms and the difficulty in feature point detection.

[0007] The idea behind this invention is to integrate aircraft inertial navigation data, eliminating the need for multi-camera or multi-frame image acquisition of attitude information. Instead, a single-camera system can estimate the relative pose between the aircraft and a identifiable ground target by capturing a single frame, thus solving the problem of long execution times in existing computer vision algorithms. After preprocessing and detecting the target image, this invention obtains the pixel coordinates of two feature points on the identifiable ground target. Using only these two feature point coordinates, the relative pose between the aircraft and the target can be calculated. This overcomes the challenge of requiring at least three feature points for relative pose estimation using computer vision algorithms alone, reducing the feature point detection requirements for pose estimation applications. In terms of pose calculation, this invention integrates the angle data of the aircraft's inertial navigation, namely the pitch angle, yaw angle and roll angle, to calculate the rotation matrix of the ground-identifiable target and measure the spatial positional relationship between two feature points. By combining the coordinate system transformation formula and the distance formula, a unique relative pose can be quickly calculated without the need for other complex calculation methods such as iteration. This overcomes the problems of high computational complexity and long algorithm running time in traditional PnP methods when solving pose.

[0008] The implementation steps of this invention are as follows:

[0009] Step 1: Obtain the internal parameters of the calibrated monocular camera;

[0010] Step 2: Use a monocular camera to capture target images and preprocess the images, then detect identifiable ground targets from the preprocessed images;

[0011] Step 3: Extract the pixel coordinates of two feature points from the identifiable target on the ground;

[0012] Step 4: Obtain the angle data of the aircraft's inertial navigation and calculate the rotation matrix of the ground-identifiable target;

[0013] Step 5: Combine the distance formula and the two sets of coordinate transformation formulas corresponding to the two feature points to calculate the relative pose between the camera and the identifiable ground target.

[0014] The invention has the following advantages compared to existing technologies:

[0015] First, the present invention uses a monocular camera to estimate the relative pose of an aircraft by capturing a single frame image, which overcomes the problem of long running time of computer vision algorithms in the prior art, enabling the present invention to complete the relative pose estimation of the aircraft with higher real-time performance.

[0016] Secondly, this invention only requires two feature points in the image of the ground-identifiable target to calculate the relative pose of the target. This overcomes the problem that relying solely on computer vision algorithms requires at least three feature points to estimate the relative pose between the aircraft and the ground-identifiable target. It reduces the detection requirements of feature points for pose estimation applications, enabling this invention to improve the real-time performance and accuracy of relative pose estimation of the aircraft in the terminal guidance phase.

[0017] Third, in terms of pose calculation, this invention integrates the angle data of the aircraft's inertial navigation, which can quickly calculate a unique relative pose using only two feature points. This overcomes the shortcomings of the traditional PnP method, which has high computational complexity and long algorithm running time when solving pose. This invention can save valuable computing resources on the aircraft, shorten the pose calculation time, and improve the calculation speed of relative pose estimation. Attached Figure Description

[0018] Figure 1 This is a flowchart of the present invention;

[0019] Figure 2 This is a diagram showing the results of the simulation experiment of this invention. Detailed Implementation

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] Reference Figure 1 The implementation steps of the present invention will be further described below.

[0022] Step 1: Obtain the internal parameters of the calibrated monocular camera.

[0023] The calibration of the monocular camera can be performed using any one of the following methods: Tsai two-step method, Zhang's calibration method, active vision camera calibration method, camera self-calibration method, zero-loss machine calibration method, calibration based on genetic algorithm, calibration based on neural network, and calibration based on error backpropagation.

[0024] In an embodiment of the present invention, Zhang's calibration method is used to calibrate the monocular camera on the aircraft. A standard checkerboard pattern for camera calibration is placed in an appropriate position, and 20 images with complete checkerboard patterns are taken by changing the camera angle.

[0025] Twenty captured images were input into the cameracalibrator app program included with MATLAB (Matrix Laboratory) software. After inputting the checkerboard grid specifications, calibration was performed. If the calibration pixel error was no greater than 0.5 pixels, the calibration was considered successful. The internal parameters of the monocular camera after calibration were obtained from the output of the MATLAB software.

[0026] The internal parameters include: the focal length of the calibrated monocular camera f = 15mm, the horizontal pixel offset dx, the vertical pixel offset dy, and the pixel coordinates of the image center u0 = 431 and v0 = 204.5.

[0027] Step 2: Use a monocular camera to capture target images and preprocess the images, then detect recognizable ground targets from the preprocessed images.

[0028] In this embodiment of the invention, Autodesk 3ds Max 2012 simulation software is used to simulate the terminal guidance process of the aircraft, and MATLAB software is used to preprocess the target image, detect identifiable ground targets, and extract the pixel coordinates of two feature points from the identifiable ground targets. Figure 2 As shown.

[0029] Figure 2 (a) is a scene simulation image of a ground-identifiable target captured by a camera on an aircraft, obtained from Autodesk 3ds Max 2012 simulation software. In the coordinate system of the aircraft's camera, the camera coordinates of the first and second feature points of the ground-identifiable target are (125.14, -118.58, -307.37) and (133.82, -114.29, -258.32), respectively. Figure 2 (b) is a target image with identifiable ground targets captured by a simulated camera in the software.

[0030] The preprocessed image can be preprocessed using any one of the following methods: grayscale conversion, binarization, histogram equalization, filtering, image segmentation, image enhancement, geometric transformation, and edge detection, to preprocess the target image captured by the camera on the aircraft.

[0031] In the embodiments of the present invention, image segmentation, edge detection, binarization, and filter methods are sequentially used to preprocess the target image captured by the monocular camera. Figure 2 (c) is the result image obtained after performing image segmentation on the target image using the Otsu's method in MATLAB software. Figure 2 (d) is the result image obtained after edge detection of the target image after image segmentation using the LoG (Laplacian of Gaussian) operator. Figure 2 (e) is... Figure 2 (d) The result image after binarization. Figure 2 (f) is a... Figure 2 (e) The result after filtering.

[0032] The detection of identifiable ground targets from the preprocessed image can be achieved using any one of the following methods: edge detection, corner detection, key point detection, interest point detection, texture analysis, shape matching, and image segmentation. This method can detect identifiable ground targets from the preprocessed image during the terminal guidance process of the aircraft.

[0033] In embodiments of the present invention, an edge detection algorithm is used to find the edges in the preprocessed image, extract all the closed shapes formed by the detected edges, calculate the area and perimeter of the closed shapes, and calculate the roundness of each closed shape according to the formula.

[0034] The formula for calculating roundness is as follows:

[0035] e = (4 * Π * s) / c 2

[0036] Where e represents the roundness of the closed shape, s represents the area of ​​the closed shape, and c represents the perimeter of the closed shape. The roundness threshold is set to 0.9, and closed shapes with a roundness greater than 0.9 are identified as circles.

[0037] Figure 2 (g) is the result image of detecting identifiable targets on the ground from the preprocessed image.

[0038] Step 3: Extract the pixel coordinates of two feature points from the identifiable target on the ground.

[0039] The extraction of pixel coordinates of two feature points in identifiable ground targets can be achieved using any one of the following methods: fast corner detection, Harris corner detection, or Hough transform.

[0040] In an embodiment of the present invention, the Hough transform method is used to obtain the pixel coordinates of two feature points in the ground identifiable target. The pixel coordinates of the first feature point are (u1, v1) = (247.09, 361.56) and the pixel coordinates of the second feature point are (u2, v2) = (286.47, 341.45).

[0041] Step 4: Obtain the angle data of the aircraft's inertial navigation and calculate the rotation matrix of the ground-identifiable target.

[0042] The calculation of the rotation matrix of the ground-identifiable target refers to the fact that, during terminal guidance, the aircraft's rotation matrix is ​​equivalent to the rotation matrix of the ground-identifiable target. The expression for the rotation matrix of the ground-identifiable target is as follows:

[0043]

[0044] Where R represents the rotation matrix of the ground-identifiable target, α represents the pitch angle of the aircraft, β represents the yaw angle of the aircraft, and γ represents the roll angle of the aircraft.

[0045] In an embodiment of the present invention, the inertial navigation angle data of the aircraft at the time of image capture are α = 5°, β = 10°, and γ = 60°. Therefore, the rotation matrix R of the ground-identifiable target is:

[0046]

[0047] Step 5: Solve the relative pose between the camera and the identifiable ground target by combining the distance formula and the two sets of coordinate transformation formulas corresponding to the two feature points.

[0048] The distance formula is as follows:

[0049] (X w1 -X w2 ) 2 +(Y w1 -Y w2 ) 2 +(Z w1 -Z w2 ) 2 =(X c1 -X c2 ) 2 +(Y c1 -Y c2 ) 2 +(Z c1 -Zc2 ) 2

[0050] Among them, X w1 Y w1 Z w1 These represent the locations of the first identifiable feature point of the ground in the target image in the world coordinate system, specifically X and X. w Axis, Y w Axis, Z w The coordinate values ​​of the axis, X w2 Y w2 Z w2 These represent the second feature point of the identifiable ground target in the target image, located in the X coordinate system. w Axis, Y w Axis, Z w The coordinate values ​​of the axis; X c1 Y c1 Z c1 These represent the locations of the first identifiable feature point of the ground in the target image, in the camera coordinate system, at X and Y respectively. c Axis, Y c Axis, Z c The coordinate values ​​of the axis, X c2 Y c2 Z c2 These represent the locations of the second feature point of the identifiable ground target in the target image, in the camera coordinate system, at X and Y. c Axis, Y c Axis, Z c The coordinate values ​​of the axis.

[0051] In an embodiment of the present invention, two feature points in a ground-identifiable target are located at X in the world coordinate system. w Axis, Y w Axis, Z w The coordinate differences of the axes are as follows: X w1 -X w2 =0mm, Y w1 -Y w2 =0mm, Z w1 -Z w2 =50mm.

[0052] The coordinate transformation formula is as follows:

[0053]

[0054]

[0055] Where u and v represent the horizontal and vertical coordinates of any point in space in the pixel coordinate system, respectively; dx and dy represent the horizontal and vertical pixel offsets of the calibrated monocular camera, respectively; f represents the focal length of the calibrated monocular camera; u0 and v0 represent the horizontal and vertical coordinates of the image center of the calibrated monocular camera in the pixel coordinate system, respectively; and X... c Y c Z c These represent the positions of any point in space corresponding to the pixel coordinate system in the camera coordinate system, specifically the X and Y coordinates. c Axis, Y c Axis, Z c The coordinate values ​​of the axis, X w Y w Z w These represent the positions of any point in space corresponding to the pixel coordinate system in the world coordinate system, specifically the X and X coordinates. w Axis, Y w Axis, Z w The coordinate values ​​of the axis, where T represents the translation vector of the identifiable target on the ground.

[0056] In this embodiment of the invention, two feature points in the identifiable ground target are taken as two spatial points. The pixel coordinates (u1, v1) of the first feature point and (u2, v2) of the second feature point are substituted into the above formula to obtain two sets of coordinate transformation formulas. By combining the distance formula with the two sets of coordinate transformation formulas corresponding to the two feature points, the translation vector T of the identifiable ground target can be eliminated, and the camera coordinates of the first feature point in the identifiable ground target can be finally calculated as (X...). c1 Y c1 Z c1 The camera coordinates of the second feature point among the ground-identifiable targets are (X) = (125.14, -118.58, -307.37) and (X). c2 Y c2 Z c2 = (133.82, -114.29, -258.32), which means obtaining the relative pose between the camera and the recognizable target on the ground.

[0057] Comparing the camera coordinates of the two feature points in the finally obtained ground-identifiable target with the camera coordinates of the two feature points in the ground-identifiable target set in the simulation software, it can be seen that, under the premise of obtaining accurate pixel coordinates of the two feature points in the ground-identifiable target, the two sets of camera coordinates are completely consistent with zero error. The relative pose between the camera and the ground-identifiable target is accurately calculated. Therefore, it can be shown that the present invention can quickly and accurately calculate the camera coordinates of the two feature points in the aircraft camera coordinate system by fusing aircraft inertial navigation data and using two feature points, that is, obtain the accurate relative pose between the camera and the ground-identifiable target.

Claims

1. A target pose estimation method based on monocular vision and inertial navigation fusion, characterized in that, The method involves obtaining the internal parameters of a calibrated monocular camera, capturing a target image using the calibrated camera, preprocessing the target image and detecting ground-identifiable targets, extracting the pixel coordinates of two feature points from the ground-identifiable target, fusing inertial navigation data, calculating the rotation matrix of the ground-identifiable target, and solving the relative pose between the camera and the ground-identifiable target using the distance formula and coordinate transformation formula. The steps of this method are as follows: Step 1: Obtain the internal parameters of the calibrated monocular camera; Step 2: Use a monocular camera to capture target images and preprocess the images, then detect identifiable ground targets from the preprocessed images; Step 3: Extract the pixel coordinates of two feature points from the identifiable target on the ground; Step 4: Obtain the angle data of the aircraft's inertial navigation and calculate the rotation matrix of the ground-identifiable target; Step 5: Combine the distance formula and the two sets of coordinate transformation formulas corresponding to the two feature points to calculate the relative pose between the camera and the identifiable ground target.

2. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, The monocular camera calibration described in step 1 uses any one of the following methods to calibrate the monocular camera used on the aircraft: Tsai two-step method, Zhang's calibration method, active vision camera calibration method, camera self-calibration method, zero-loss machine calibration method, calibration based on genetic algorithm, calibration based on neural network, and calibration based on error backpropagation.

3. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, The internal parameters mentioned in step 1 include: the focal length of the calibrated monocular camera, the horizontal pixel offset, the vertical pixel offset, and the pixel coordinates of the image center.

4. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, The preprocessing of the image in step 2 uses any one of the following image processing methods: grayscale conversion, binarization, histogram equalization, filtering, image segmentation, image enhancement, geometric transformation, and edge detection, to preprocess the target image captured by the camera on the aircraft.

5. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, Step 2 describes the detection of identifiable ground targets from the preprocessed image using any one of the following methods: edge detection, corner detection, key point detection, interest point detection, texture analysis, shape matching, and image segmentation. This method is used to detect identifiable ground targets from the preprocessed image during the aircraft's terminal guidance process.

6. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, Step 3 involves extracting the pixel coordinates of two feature points from identifiable ground targets. This can be done using any one of the following methods: fast corner detection, Harris corner detection, or Hough transform.

7. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, The calculation of the rotation matrix of the ground-identifiable target in step 4 is equivalent to the rotation matrix of the ground-identifiable target during terminal guidance. The expression for the rotation matrix of the ground-identifiable target is as follows: ; Where R represents the rotation matrix of the identifiable ground target. Indicates the pitch angle of the aircraft. Indicates the yaw angle of the aircraft. This indicates the roll angle of the aircraft.

8. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 1, characterized in that, The distance formula mentioned in step 5 is as follows: ; in, , , These represent the locations of the first feature point of the identifiable ground target in the target image in the world coordinate system. axis, axis, The coordinate values ​​of the axis. , , These represent the locations of the second feature points of the ground-identifiable targets in the target image in the world coordinate system. axis, axis, The coordinate values ​​of the axis; , , These represent the locations of the first feature point of the identifiable ground target in the target image in the camera coordinate system. axis, axis, The coordinate values ​​of the axis. , , These represent the locations of the second feature points of the identifiable ground target in the target image in the camera coordinate system. axis, axis, The coordinate values ​​of the axis.

9. The target pose estimation method based on monocular vision and inertial navigation fusion according to claim 7, characterized in that, The coordinate transformation formula mentioned in step 5 is as follows: ; ; in, , These represent the horizontal and vertical coordinates of any point in space in the pixel coordinate system, respectively. , These represent the horizontal and vertical pixel offsets of the monocular camera after calibration, respectively. This indicates the focal length of the monocular camera after calibration. , These represent the horizontal and vertical coordinates of the image center of the calibrated monocular camera in the pixel coordinate system. , , These represent the positions of any point in space corresponding to the pixel coordinate system in the camera coordinate system. axis, axis, The coordinate values ​​of the axis. , , These represent the positions of any point in space corresponding to the pixel coordinate system in the world coordinate system. axis, axis, The coordinate values ​​of the axis, where T represents the translation vector of the identifiable target on the ground.

Citation Information

Patent Citations

  • Visual navigation positioning method and system for low-speed unmanned aerial vehicle

    CN116429098A

  • Monocular camera pose measurement method based on inertial measurement unit and point-line features

    CN110375732A

  • Monocular simultaneous localization and mapping pose solving method fused with inertial measurement unit

    CN110375738A