Monocular vision pose tracking method

By using four infrared reflective ball targets and monocular cameras, the constraint relationship between the three-dimensional coordinates of the ball target center and pixel point coordinates are established, the coordinates of the ball target center are optimized, and the pose is calculated through singular value decomposition, the automation and anti-interference problems of monocular visual pose estimation are solved, and efficient pose tracking is achieved.

CN120147426APending Publication Date: 2025-06-13CENT SOUTH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510304651.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art relies on manual matching feature points in monocular visual pose estimation, and the depth estimation method is sensitive to illumination changes and external interference, making it difficult to achieve efficient automated pose tracking.

Method used

Four infrared reflective ball targets are used as pose tracking targets, and the constraint relationship between the three-dimensional coordinates of the spherical target center and the pixel point coordinates are established through the monocular camera imaging principle, and the coordinates of the spherical target center are optimized and solved, and the six-degree-of-freedom space pose is calculated by singular value decomposition.

Benefits of technology

The automation of monocular visual pose tracking is realized, the steps of manually matching feature points are avoided, the positioning accuracy and anti-interference ability are improved, and it is suitable for real-time pose tracking in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147426A_ABST
    Figure CN120147426A_ABST
Patent Text Reader

Abstract

The invention provides a monocular vision pose tracking method, which comprises the following steps: according to a camera imaging principle, establishing a constraint relation between three-dimensional coordinates of the center of a ball target and coordinates of pixel points on an imaging contour, optimally solving the three-dimensional coordinates of the center of the ball target, realizing spatial positioning of a single ball target, and forming a tracking target by a plurality of ball targets. After the center position of each ball target is captured in real time, the six-degree-of-freedom space pose of the tracking ball target is obtained according to singular value decomposition, and for positioning errors caused by ball target contour extraction, ball target diameter deviation and the like, a calibration method based on feature points is provided to calibrate internal parameters of a system, and compared with a traditional PnP method, the method has the advantage that the positioning accuracy is improved. The positioning process does not need to pre-specify the corresponding relation between the feature points and the image points, does not depend on the pose relation of other feature points, does not relate to other auxiliary constraints, and is easy to realize automatically.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a monocular vision pose tracking method. Background Art

[0002] The positioning and measurement of objects have wide applications in industrial automation, target detection and tracking. Traditionally, lidar is often used to obtain high-precision target depth information. However, due to its high price, it is currently mostly used in the technology R & D and testing stages, and there is still a certain distance from large-scale market applications. In recent years, with the rapid development of artificial intelligence technology, vision has gradually become a research hotspot, but some drawbacks have also emerged. Among them, the depth estimation based on binocular vision is limited by the baseline length, resulting in poor matching between the device volume and the vehicle platform; the depth estimation range based on RGB-D (red-green-blue-depth) is short, and its capabilities are limited in practical applications. While a monocular camera has the advantages of low price, rich information content, small volume, etc., and can effectively overcome many deficiencies of the above sensors. Therefore, it is of great research significance to use a monocular camera to obtain depth information.

[0003] Monocular vision pose estimation is a process of solving the target pose from the constraints between the image information and the spatial information established by the imaging principle. The image information is usually called "features", and the frequently used features include points, lines and circles. For the pose measurement method based on feature points, the high-precision and fast pose estimation algorithm (Efficient Perspective-N-Point, EPNP) proposed by Lepetie and Moreno in 2009 is considered to be one of the most efficient camera pose estimation algorithms at present. This algorithm does not require iterative solution, the time complexity is O(n), and it has strong anti-interference ability. Only 3 pairs of coplanar (4 pairs are required for non-coplanar) 3D-2D matching points are needed to obtain the accurate pose. However, when performing pose estimation by this method, the operator needs to manually match the corresponding relationship between the feature points and the image points (or the matching of object points and image points) according to the observed image before the pose solution can be carried out. There is less research on the method based on circle features. However, the duality problem exists in the solution of a single circle feature, and auxiliary functions are required to eliminate this duality.

[0004] With the wide application of deep learning, recent work has adopted deep neural networks (DNNs) to estimate depth maps from single images, including unsupervised learning methods and supervised learning methods. Unsupervised learning methods use geometric constraints from adjacent frames as supervision to estimate monocular depth, thus avoiding the large amount of manual work of collecting ground truth. However, the training loss of unsupervised learning largely depends on the geometric consistency and reconstruction accuracy between consecutive frames, so the accuracy of depth estimation is easily affected by illumination changes and blurred images during fast driving. Supervised depth estimation methods predict the pixel-level depth of the input image. They learn special information from the environment and motion based on geometric principles and target features to improve the quality of the depth map, including semantic information, continuous conditional random fields, geometric relationships, and structure from motion. However, they require a large amount of real depth and environmental annotations, which hinders their generalization ability. Other work has tried to synthesize data from structure from motion or generate sparse training data, but they are sensitive to external interference, especially in dynamic scenes, and cannot cover as many scenarios as possible. Summary of the Invention

[0005] The present invention proposes a monocular vision pose tracking method, which uses four infrared reflective ball targets to form a pose tracking target to achieve the pose tracking of monocular vision. The positioning process does not depend on the pose relationship of other feature points and does not involve other auxiliary constraints, making it easy to automate. First, according to the imaging principle of the monocular camera, the constraint relationship between the three-dimensional coordinates of the ball target center and the pixel coordinates on the imaging contour is established, and the three-dimensional coordinates of the ball target center are optimized and solved to achieve the spatial positioning of a single ball target. A calibration method based on feature points is proposed for the ball target positioning error caused by parameter errors such as image contour extraction and ball target diameter. According to the imaging ellipse equation, the relationship between the ball target center and the system internal parameters is deduced, and the accurate coordinates of the ball target center are obtained through the PnP method, and the system parameters are calibrated using this as the calibration reference. After forming a tracking target with multiple ball targets and capturing the positions of the centers of each ball target in real time, the six-degree-of-freedom spatial pose of the tracking ball target is obtained according to the singular value decomposition.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] Step 1: Determine the ellipse equation of the ball target on the image plane according to the imaging principle of the ball target in the camera;

[0008] Step 2: Obtain the approximate circle radius of the ball target contour on the image and estimate the initial value P of the ball center coordinates 0 ;

[0009] Step 3: Construct an error function to optimize and update the ball target center coordinates;

[0010] Step 4: For the case where the error exceeds the allowable value, a calibration method based on feature points is proposed to calibrate the system and improve the accuracy of ball target contour fitting and positioning;

[0011] Step 5: Use multiple ball targets to form a tracking target. After calculating the center coordinates of each ball target, calculate the target pose through singular value decomposition to achieve the pose tracking of the target.

[0012] Using the above monocular vision pose tracking method, a tracking target is composed of 4 infrared reflective ball targets. The center coordinates of the ball are located according to the imaging contour of each ball target to achieve the pose tracking of the target. The positioning process does not depend on the pose relationship of other feature points and does not involve other auxiliary constraints, making it easy to achieve automation. Brief Description of the Drawings

[0013] Figure 1 is the flowchart of monocular vision pose tracking;

[0014] Figure 2 is the imaging diagram of the ball target in the camera;

[0015] Figure 3 is the tracking target diagram. Detailed Embodiment

[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0017] The process of the monocular vision pose tracking method is shown in Figure 1 , and the specific steps are as follows:

[0018] Step 1: According to the imaging principle of the ball target in the camera, determine the elliptical equation of the ball target on the image plane;

[0019] The position of the perspective center of the pinhole imaging model relative to the image plane is determined. Define the camera coordinate system O with the optical center as the origin c -X c Y c Z c ; The relationship between the pixel coordinate system and the camera coordinate system is expressed in matrix form as Equation (1):

[0020]

[0021] In the formula, f is the camera focal length, d x and d y are the pixel sizes of the camera in the u-axis and v-axis directions respectively; c x and c y are the u-axis and v-axis coordinates of the optical axis center in the pixel coordinate system respectively; f, d x , d y , cx and c y The internal parameters of the camera, such as x and c, can be obtained through the Zhang's camera calibration method;

[0022] Assume that the coordinates of the center of a sphere with radius R in the camera coordinate system are P 0 = [a, b, c] T ; The image of the sphere on the image plane is a circle with radius r. The projection model of the sphere is a cone, and the half vertex angle θ = arc sin(R / |P 0 |). There is:

[0023]

[0024] The imaging contour of the sphere inside the camera is as shown in Figure 2 . The projection model of the sphere is a cone. For any point P = [x, y, z] on the surface of the cone, it should satisfy |P Figure 2 × P| = |P 0 ||P| sin θ, that is: 0 ||P| sin θ, that is:

[0025]

[0026] On the image plane, the imaging pattern of the sphere is an ellipse, and its equation is:

[0027] (b 2 + c 2 - R 2 )x 2 + (a 2 + c 2 - R 2 )y 2 - 2abxy - 2acfx - 2bcfy = (R 2 - a 2 - b 2 )f 2

[0028] Step 2: Calculate the approximate radius of the circular contour of the sphere target on the image, and estimate the initial value P 0 ;

[0029] Through image contour detection methods such as the Canny operator, the points c i = [u i , v i of the sphere contour on the pixel plane can be extracted. The set of N points on the contour is C = {c T}, where i = 1, 2,..., N. Converted to the image plane, the corresponding point coordinates g i = [x i , y i . There is: i T There is: ​

[0030]

[0031] Currently, the pixels of most cameras are square, i.e., d x = d y . Let:

[0032]

[0033] Divide both sides of Equation (4) by d x 2 . Then Equation (4) can be transformed into:

[0034]

[0035] The centroid coordinates c of the contour 0 = [u 0 , v 0 T There is:

[0036]

[0037] On the image plane, the approximate circular radius ρ of the sphere contour has:

[0038]

[0039] Estimate the approximate value of the object distance c of the small ball according to the following formula:

[0040]

[0041] Furthermore, the center coordinates P of the sphere can be estimated 0 .

[0042]

[0043] Step 3: Construct an error function to optimize and update the equation and the center coordinates of the spherical target, and complete the fitting of the spherical target ellipse equation; construct the error function E(a, b, c):

[0044]

[0045] Among them:

[0046]

[0047] Calculate the gradient of the error function:

[0048]

[0049] Among them:

[0050]

[0051] ​Calculate the Hessian matrix H:

[0052]

[0053] Where:

[0054]

[0055] Iteratively update the estimated value of the center coordinate P of the sphere obtained from Equation (11) until 0 it is less than the allowable error ξ or the number of iterations reaches the maximum number of iterations limit k max max .

[0056]

[0057] Step 4: For the case where the error exceeds the allowable value, propose a calibration method based on feature points to calibrate the system and improve the accuracy of sphere target contour fitting and positioning;

[0058] Due to errors in parameters such as the camera internal parameters, the radius of the target sphere, and the coordinates of the sphere target contour points, there are certain errors in the calculated center coordinates of the sphere target; since the number of sphere target contour points is large and the errors of each contour point are different during each detection, it is impossible to calibrate the coordinate errors of the sphere target contour points one by one. A feasible way is to identify the measurement error to the camera internal parameters and the diameter of the sphere target. For example, if the imaged contour of the sphere target is detected to be smaller, the radius R of the sphere target can be reduced to ensure the detection accuracy. Therefore, the present invention proposes a calibration method based on feature points. By obtaining the actual coordinates of the sphere target through the PnP method, using this as the calibration reference, establishing the relationship between the center coordinates of the sphere target and the internal parameters of the monocular tracking system, compensating the system internal parameters, and improving the accuracy of sphere target contour fitting and positioning;

[0059] Equation (7) can be written in the following form:

[0060]

[0061] Where:

[0062]

[0063] It can be deduced that the short semi - axis l of the ellipse formed by the imaging of the sphere target on the pixel plane a is:

[0064]

[0065] According to the center coordinate P of the sphere target optimized in Step 3 0 , combined with Equation (1) and Equation (6), the projection coordinate O of the center of the sphere target on the pixel plane can be obtained w = [u w , v w ​T There are:

[0066]

[0067] From Equation (11) and Equation (21), the relationship between the center coordinates of the spherical target and the internal parameters of the system can be established:

[0068]

[0069] The internal parameters of the system are summarized as S = [R, f x , c x , c y T .

[0070] Differentiating Equation (23) gives:

[0071]

[0072] where J is the Jacobian matrix corresponding to the center point of the spherical target:

[0073]

[0074] Discretizing Equation (27) gives:

[0075] δP = J·δS

[0076] First, obtain the center coordinates of n spherical targets through the PnP method, denoted as p 1 , p 2 , …… p n . As a calibration reference, Equation (26) can be extended to:

[0077]

[0078] (K T K) -1 K T represents the pseudo-inverse of matrix K. The symbols in Equation (27) are summarized as follows:

[0079]

[0080] Step 5: Use multiple spherical targets to form a tracking target. After measuring the center coordinates of each spherical target, calculate the pose of the target through singular value decomposition to achieve pose tracking of the target;

[0081] In the tracking target described in the present invention, 4 reflective spherical targets are used as the positioning reference, and the pose of the tracking target is described by the coordinate system O w -XYZ formed by the 4 centers of the spheres. The coordinates of the 4 centers of the spheres in this coordinate system are respectively W P 1 = [0, L 1 , 0]​T , W P 2 =[L 2 ,0,0] T , W P 3 =[0,L 3 ,0] T , W P 4 =[L 4 ,0,0] T ,where L 1 , L 2 , L 3 and L 4 are not equal to each other, as Figure 3 shown.

[0082] The coordinates of the centers of the 4 balls in the camera coordinate system measured by the vision camera are respectively C P 1 , C P 2 , C P 3 , C P 4 . It is necessary to find a set of Euclidean transformations R, t such that:

[0083] C P i =R· W P i +t

[0084] Calculate the centroid - removed coordinates of the two sets of points:

[0085]

[0086] Define the matrix:

[0087]

[0088] M is a 3×3 matrix. Perform singular - value decomposition on it to get:

[0089] M=UΣV T

[0090] where Σ is a diagonal matrix composed of singular values, and the diagonal elements are arranged from large to small, and U and V are orthogonal matrices. When W is full - rank, R is:

[0091] R=UV T

[0092] Furthermore, t can be obtained:

[0093] t= C P - R· W P

[0094] The homogeneous transformation matrix composed of R and t C T w That is the pose description of the tracking target in the camera coordinate system.

[0095]

Claims

1. A monocular vision posture tracking method, characterized in that: The following steps are included: Step 1, according to the imaging principle of the ball target in the camera, determine the ellipse equation of the ball target on the image plane; Step 2, find the approximate circle radius of the ball target outline on the image and estimate the initial value P0 of the ball center coordinate; Step 3, construct an error function to optimize and update the center coordinates of the ball target; Step 4: When the error exceeds the allowable value, a calibration method based on feature points is proposed to calibrate the system and improve the accuracy of ball target contour fitting and positioning; Step 5, a tracking target is formed by using multiple spherical targets, and after the center coordinates of each spherical target are calculated, the target pose is calculated by singular value decomposition to achieve target pose tracking.

2. The method according to claim 1, characterized in that The center coordinates of the ball target are optimized and solved according to the imaging pattern of the ball target on the camera to realize the spatial positioning of the ball target in the camera coordinate system.

Citation Information

Cited By

  • Slope displacement monitoring method and monitoring system based on spherical target

    CN120609274A