Three-view relative pose solving method, system and equipment under planar motion, medium and product
By collecting views on moving objects and constructing linear equation systems, and using feature matching points for solving, the problem of solving difficulties and high computational complexity of traditional three-view pose estimation methods under plane motion is solved, and efficient pose solution is achieved.
Patent Information
- Application Number
- CN202510374515.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-17
AI Technical Summary
The application of traditional three-view pose estimation method under plane motion is limited by problems such as difficulty in solving and high computational complexity, especially when there is a mismatch situation, the problem of solution degradation is prone to occur.
By setting up a monocular camera on a moving object, the views of different moments are collected, and a linear system of equations is constructed using projection matrix and triple-focus tensor, and feature matching points are extracted for solving. When the number of groups of feature matching points is three groups, the singular value decomposition method is used for the solution; when it is one group, the yaw angle is measured in combination with the inertial measurement unit.
This method reduces the number of feature matching points in the pose solution process, reduces the calculation amount, improves the solution efficiency, and is suitable for scenarios with high real-time requirements.
Smart Images

Figure CN120160635A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of pose solution, and particularly to a method, system, device, medium and product for solving the relative pose of three views under planar motion. Background Art
[0002] In moving platforms such as unmanned vehicles and unmanned aerial vehicles, using vision sensors for pose estimation is a key technology. Traditional two-view relative pose estimation algorithms are relatively mature, but they are prone to the problem of solution degradation in the presence of mismatches. Traditional three-view pose estimation methods usually require a large number of feature matching points, and there are problems of difficult solution and high computational complexity, which limit their applications. Summary of the Invention
[0003] The purpose of the present application is to provide a method, system, device, medium and product for solving the relative pose of three views under planar motion, which can improve the solution efficiency of the relative pose of an object.
[0004] To achieve the above purpose, the present application provides the following solutions:
[0005] In the first aspect, the present application provides a method for solving the relative pose of three views under planar motion, including:
[0006] During the planar motion of a moving object, use a monocular camera installed on the moving object to collect views at different times, obtaining a first view, a second view and a third view;
[0007] Construct a linear equation system of the trifocal tensor according to the projection matrix and trifocal tensor of the monocular camera when the moving object is in planar motion;
[0008] Extract a preset number of groups of feature matching points from the first view, the second view and the third view;
[0009] When the number of groups of feature matching points is three, use the feature matching points to solve the linear equation system of the trifocal tensor;
[0010] When the number of groups of feature matching points is one, use an inertial measurement unit to measure the yaw angle of each view; and solve the linear equation system of the trifocal tensor based on the yaw angle and the feature matching points;
[0011] Determine the relative pose between the first view, the second view and the third view according to the solution result.
[0012] Further, extracting a preset number of groups of feature matching points from the first view, the second view and the third view specifically includes:
[0013] Extract a preset number of feature matching points in the first view, the second view, and the third view by using the Scale-Invariant Feature Transform (SIFT) algorithm.
[0014] Further, the projection matrix is:
[0015]
[0016] where t k is the translation vector of the k-th view, and P k is the projection matrix of the monocular camera corresponding to the k-th view, θ k is the yaw angle of the k-th view, and the subscript y in them indicates that the monocular camera rotates around the Y-axis, is the cosine value of the yaw angle of the k-th view, is the sine value of the yaw angle of the k-th view, is the translation displacement of the k-th view along the x-axis, is the translation displacement of the k-th view along the z-axis.
[0017] Further, the linear equations of the trifocal tensor are:
[0018]
[0019] where, are all known coefficients, and Q1 to Q 12 are all unknowns.
[0020] Further, when the preset number of feature matching points is three, use the feature matching points to solve the linear equations of the trifocal tensor, specifically including:
[0021] Based on the feature matching points, use the singular value decomposition method to solve the linear equations of the trifocal tensor.
[0022] Further, the relative pose includes a translation vector and a yaw angle; taking the coordinate system of the first view as the world coordinate system, the yaw angle of the second view is the yaw angle of the second view relative to the first view, and the translation vector of the second view is the translation vector of the second view relative to the first view; the yaw angle of the third view is the yaw angle of the third view relative to the first view, and the translation vector of the third view is the translation vector of the third view relative to the first view;
[0023] The expressions for the translation vector t2 and the yaw angle θ2 of the second view are:
[0024]
[0025] where, is the cosine value of the yaw angle of the second view; is the sine value of the yaw angle of the second view;
[0026] The expressions for the translation vector t3 and the yaw angle θ3 of the third view are respectively:
[0027]
[0028] where, is the cosine value of the yaw angle of the third view; is the sine value of the yaw angle of the third view.
[0029] In a second aspect, the present application provides a three-view relative pose solving system under planar motion, including:
[0030] An acquisition module, configured to collect views at different times by using a monocular camera disposed on a moving object during the planar motion of the moving object, so as to obtain a first view, a second view, and a third view;
[0031] A construction module, configured to construct a linear equation system of the trifocal tensor according to the projection matrix and the trifocal tensor of the monocular camera when the moving object performs planar motion;
[0032] An extraction module, configured to extract a preset number of groups of feature matching points in the first view, the second view, and the third view;
[0033] A first solving module, configured to solve the linear equation system of the trifocal tensor by using the feature matching points when the number of groups of feature matching points is three;
[0034] A second solving module, configured to measure the yaw angle of each view by using an inertial measurement unit when the number of groups of feature matching points is one; and solve the linear equation system of the trifocal tensor based on the yaw angle and the feature matching points;
[0035] A pose determination module, configured to determine the relative pose between the first view, the second view, and the third view according to the solution result.
[0036] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned three-view relative pose solving method under planar motion.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned three-view relative pose solving method under planar motion is implemented.
[0038] In a fifth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method for solving the relative pose of three views under planar motion described above.
[0039] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0040] The present application provides a method, a system, a device, a medium and a product for solving the relative pose of three views under planar motion. By constructing a linear equation system of trifocal tensors based on the projection matrix and trifocal tensors of a monocular camera when a moving object performs planar motion; and using the monocular camera arranged on the moving object to collect views at different times to obtain a first view, a second view and a third view, and extracting a preset number of groups of feature matching points in the first view, the second view and the third view. When the number of groups of feature matching points is three, the linear equation system of trifocal tensors is solved using the feature matching points; when the number of groups of feature matching points is one, the yaw angles of each view are measured, and the linear equation system of trifocal tensors is solved using the yaw angles and the feature matching points. The relative pose between the first view, the second view and the third view is determined using the solution result. This method uses three groups or one group of feature matching points during the solution process, reducing the number of feature matching points in the pose solution process. Based on this, the computational amount in the pose solution process is greatly reduced, thereby improving the solution efficiency. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 It is an application environment diagram of a method for solving the relative pose of three views under planar motion in an embodiment of the present application;
[0043] Figure 2 It is a schematic flowchart of a method for solving the relative pose of three views under planar motion provided by an embodiment of the present application;
[0044] Figure 3 It is a schematic diagram of relative pose by the three-point method;
[0045] Figure 4 It is a schematic diagram of triangulation principle;
[0046] Figure 5 It is a comparison diagram of probability density of rotation error by the three-point method;
[0047] Figure 6 It is a comparison diagram of the probability density of the translation error of the three-point method;
[0048] Figure 7 It is a comparison diagram of the rotation error of the three-point method varying with the image noise;
[0049] Figure 8 It is a comparison diagram of the translation error of the three-point method varying with the image noise;
[0050] Figure 9 It is a comparison diagram of the error of the three-point method varying with the camera vibration;
[0051] Figure 10 It is a comparison diagram of the probability density of the translation error of the one-point method;
[0052] Figure 11 It is a comparison diagram of the translation error of the one-point method varying with the image noise;
[0053] Figure 12 It is a comparison diagram of the translation error of the one-point method varying with the angular noise;
[0054] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0056] To make the purpose, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0057] The method for solving the relative pose of the three views under planar motion provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send views at different times to the server 104. After receiving the views at different times, for the views at different times, the server 104, based on the views at different times, obtains a first view, a second view, and a third view; extracts the feature matching points in the first view, the second view, and the third view; constructs a linear equation system of trifocal tensors according to the projection matrix of the monocular camera when the moving object performs planar motion and the feature matching points; solves the linear equation system of the trifocal tensors to obtain the relative poses between the first view, the second view, and the third view. The server 104 can feedback the views at different times to the terminal 102. In addition, in some embodiments, the method for solving the relative poses of three views under planar motion can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly process the views at different times, or the server 104 can obtain the views at different times from the data storage system and process the views at different times.
[0058] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0059] In an exemplary embodiment, as Figure 2 shown, a method for solving the relative poses of three views under planar motion is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in
[0060] as an example, the following steps 201 to step 206 are included.
[0061] Step 201, during the process of the moving object performing planar motion, use a monocular camera set on the moving object to collect views at different times, and obtain a first view, a second view, and a third view.
[0062] Step 202: Construct a linear equation system of the trifocal tensor based on the projection matrix and trifocal tensor of the monocular camera when the moving object performs planar motion.
[0063] Step 203: Extract a preset number of groups of feature matching points in the first view, second view, and third view.
[0064] Specifically, the scale-invariant feature transform algorithm is used to extract the feature matching points in the first view, second view, and third view. The feature matching points are corresponding points in the first view, second view, and third view.
[0065] Step 204: When the number of groups of feature matching points is three, use the feature matching points to solve the linear equation system of the trifocal tensor.
[0066] Step 205: When the number of groups of feature matching points is one, use the inertial measurement unit to measure the yaw angle of each view; and based on the yaw angle and the feature matching points, solve the linear equation system of the trifocal tensor.
[0067] Step 206: Determine the relative poses between the first view, second view, and third view according to the solution results.
[0068] Implementing the above steps 201 to 206 can reduce the computational amount when solving the relative pose, and thus improve the solving efficiency of the relative pose.
[0069] The projection matrix of the monocular camera when the moving object performs planar motion in step 202 is:
[0070]
[0071] where, t k is the translation vector of the k-th view, P k is the projection matrix of the monocular camera corresponding to the k-th view, θ k is the yaw angle of the k-th view, and the subscript y in them indicates that the monocular camera rotates around the Y axis, is the cosine value of the yaw angle of the k-th view, is the sine value of the yaw angle of the k-th view, is the translation displacement of the k-th view along the x-axis, is the translation displacement of the k-th view along the z-axis.
[0072] The derivation process of the projection matrix of the monocular camera when the moving object performs planar motion is:
[0073] When a monocular camera mounted on a moving object moves in three-dimensional space, it has three degrees of freedom of rotation. Considering rotation around the X, Y, and Z axes respectively, the rotation matrix R for rotation around the X axis x and the rotation matrix R for rotation around the Y axis y and the rotation matrix R for rotation around the Z axis z are respectively:
[0074]
[0075] where α, θ, and β are the pitch angle, yaw angle, and roll angle of the monocular camera respectively. The rotation matrix R k is the product of the three matrices R x , R y , and R z . The translation vector t k can be expressed as:
[0076]
[0077] where t k is the translation vector of the k-th view, and t x , t y , t z are the translation displacements of the monocular camera on the X, Y, and Z axes respectively.
[0078] In this embodiment, normalized projection matrices and normalized image coordinates are adopted. The projection matrix P corresponding to the k-th view of the monocular camera k can be expressed as:
[0079]
[0080] When the moving object is in planar motion, there is only one degree of freedom of rotation left. Assuming that the moving object can only rotate around the Y axis, the rotation matrix R k is:
[0081]
[0082] where θ k is the yaw angle of the k-th view; is the cosine value of the yaw angle of the k-th view; is the sine value of the yaw angle of the k-th view.
[0083] And the translation displacement on the Y axis during planar motion is always 0. The translation vector t k can be expressed by formula (2); it is deduced that the projection matrix P of the monocular camera under planar motion k can be expressed by formula (1).
[0084] The specific process of establishing the linear equations of the trifocal tensor is as follows:
[0085] As Figure 3 shown, assuming that the coordinate system of the first view Figure X 1 is used as the world coordinate system, then R1 = I 3×3 , t1 = (0, 0, 0) T , R k and t k are respectively the rotation matrix and translation vector of the k-th view relative to the first view Figure X 1. C1, C2, and C3 in Figure 3 are respectively the different positions of the monocular cameras for obtaining the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3. Let the projection matrix corresponding to the k-th view of the monocular camera be expressed as: P1 = [I|0], P2 = [A|a4], P3 = [B|b4], where k takes values 1, 2, and 3 respectively. Here, A and B are 3×3 matrices, and the vectors a i and b i are respectively the i-th columns of the A and B matrices, with i = 1, 2, 3. a4 and b4 are respectively the epipoles generated by the monocular camera for the second view Figure X 2 and the third view Figure X 3. Then the trifocal tensor T i can be expressed as:
[0086]
[0087] Combined with formula (1), the expression of the trifocal tensor {T1 T2 T3} under planar motion can be obtained:
[0088]
[0089] where,
[0090]
[0091] Meanwhile,
[0092]
[0093] Considering a set of corresponding points in the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3 The relationship of this set of corresponding points is as follows:
[0094]
[0095] where, x 1 , x2 , x 3 are the image coordinates corresponding to the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3 respectively. If is the i-th coordinate of x 1 , formula (16) can be transformed into:
[0096]
[0097] Among them, are all known coefficients, and Q1 to Q 12 are all unknowns.
[0098] In step 204, when the number of groups of feature matching points is three, the linear equations of the trifocal tensor are solved by using the Singular Value Decomposition (SVD) method based on the three groups of feature matching points. At this time, all the unknowns of the linear equations of the trifocal tensor are Q1 to Q 12 .
[0099] Based on Q1 to Q 12 , the relative poses between the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3 are obtained.
[0100] Specifically, formula (18) contains 9 equations. However, by analyzing the rank of the equations in formula (18), it is found that only 4 equations are independent. When there are 3 groups of corresponding points, 12 independent equations can be constructed. The system of equations composed of these 12 equations can be expressed as Hq = 0, where q = [Q1 Q2 … Q 12 T ; each element in H is known, where n represents the n-th equation and m represents the m-th Q value. Therefore, [Q1 Q2 … Q 12 can be solved by SVD decomposition T , and then can be solved from formula (16) to determine the rotation matrix R k , the translation vector t k , and the relative poses between the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3 are obtained.
[0101] (1) The yaw angle θ2 of the second view:
[0102]
[0103] (2) Translation vector t2 of the second view:
[0104]
[0105] where is the cosine value of the yaw angle of the second view; is the sine value of the yaw angle of the second view.
[0106] (3) Yaw angle θ3 of the third view:
[0107]
[0108] (4) Translation vector t3 of the third view:
[0109]
[0110] where is the cosine value of the yaw angle of the third view; is the sine value of the yaw angle of the third view.
[0111] In step 205, when the number of groups of feature matching points is one, the yaw angles of each view are measured using the inertial measurement unit; the yaw angle θ k of the k-th view can be obtained. When θ k is known, and are both known quantities. Considering a set of corresponding points Figure X 1 in the first view Figure X 2 in the second view and Figure X 3 in the third view from the trifocal tensor {T1 T2 T3} represented by [Q1 Q2 … Q 12 T , only is unknown. At this time, the unknowns in the linear equations of the trifocal tensor are
[0112] In the actual application process, using the measurement information of the inertial measurement unit (IMU) in the Inertial Navigation System (INS), the yaw angle θ k of the k-th view can be obtained. Therefore, based on the premise that the rotation matrix is known, a linear 1-point algorithm is proposed. Using the known yaw angle θ k , only one set of feature matching points is required to solve all the unknowns in the linear equations of the trifocal tensor.
[0113] Equation (18) can be expressed in the following form:
[0114]
[0115] This equation is in the form of Gt = 0, and the rank of G is 3. Let |t| = 1, so only one set of feature matching points is needed to solve for the value, and then the first view Figure X 1, the second view Figure X 2, and the third view Figure X 3. The relative pose between them can be obtained.
[0116] In the actual application process, especially when dealing with real image data, to ensure the feasibility of the algorithm, it is often necessary to correct the rotation matrix and translation vector of the monocular camera first to make it satisfy the planar motion constraint. Before correction, the rotation matrix between the first view and the second view of the monocular camera is R, and the translation vector is t. After correction, the rotation matrix is R', and the translation vector is t'.
[0117]
[0118] It is easy to see that R = R z R x R', assuming the correction matrix has the following formula:
[0119]
[0120] where R p is the correction rotation matrix, and t p is the correction translation vector.
[0121] Let's assume R p = R z R x , then we can solve for t p = t - R p t'. Thus, the correction matrix J can be obtained.
[0122] After solving the correction matrix J, the image coordinates can be corrected so that the transformation relationship between the image coordinates is represented by the corrected rotation matrix R' and translation vector t'. Assume the coordinates of the feature matching points in the first view are X c1 , and the coordinates in the second view are X c2 , then we can deduce:
[0123]
[0124] Substitute (25) into (26):
[0125]
[0126] That is, the corrected image coordinate X c2 ' = J -1 X c2 .
[0127] When correcting the image coordinates, since the image coordinates used in the dataset are normalized coordinates and therefore, the corrected image coordinates are actually and the depth information λ2 of the feature matching points is unknown. Therefore, the triangulation method is used to recover the depth information.
[0128] Triangulation refers to the method of observing the same point P from different positions and inferring the coordinates of this point through the coordinates of the homologous points on different images. As Figure 4 shown, O1 and O2 represent different positions. Assume are the normalized coordinates of two homologous points, and the coordinates of point P in the world coordinate system are From the previous derivation, we know that:
[0129]
[0130] Decompose T1 into row vectors T 11 , T 12 , T 13 , we have:
[0131]
[0132] That is,
[0133]
[0134] In the system of equations in formula (30), substitute λ1 = T 13 P into the first two equations:
[0135]
[0136] Similarly, decompose T2 into row vectors T 21 , T 22 , T 33 , we have:
[0137]
[0138] The system of equations (31)(32) has four equations in the form of EP = 0, and point P has 3 unknowns. Using the least squares method, P can be calculated.
[0139] This application proposes a three-view relative pose estimation method based on the planar motion assumption, including two solution strategies: the three-point method (solving the linear equations of the trifocal tensor using three sets of feature matching points) and the one-point method (solving the linear equations of the trifocal tensor using the yaw angle and one set of feature matching points), which are used to solve the three-view relative pose estimation problem. The three-point method establishes a linear equation system based on the trifocal tensor through three sets of feature matching points, and can linearly solve the relative pose between the three views, with high accuracy and high robustness; while the one-point method combines with the inertial navigation system, obtains the yaw angle using the measurement information of the inertial navigation system, and uses the known yaw angle. Only one set of corresponding points is required to solve the linear equations of the trifocal tensor, and then the pose can be determined, significantly reducing the computational complexity and being applicable to scenarios with high real-time requirements. Compared with the traditional pose solution algorithms, the method in this application not only improves the accuracy and robustness of the three-view pose solution, but also further improves the computational efficiency through the one-point method, and has high engineering application value.
[0140] Compare the methods proposed in this embodiment - Our-3pt (three-point method), Our-1pt (one-point method) with the ordinary normalization solution Hartley-7pt of the three views disclosed in the related technology, and some classic and mature two-view algorithms (Hartley-8pt and Nister-5pt).
[0141] The specific experimental results of the three-point method are as follows:
[0142] 1. Without considering the influence of image noise, the experimental results of numerical stability are as Figure 5 and Figure 6 shown. All algorithms are executed 5000 times. The smaller the rotation error, the higher the probability density, indicating better stability; the smaller the translation error, the higher the probability density, indicating better stability. It can be seen that the stability of Our-3pt is the best, followed by the stability of Nister-5pt and Hartley-8pt, and the stability of Hartley-7pt is the worst. Generally speaking, the Our-3pt algorithm proposed in this embodiment has complete advantages in terms of stability.
[0143] 2. In order to measure the accuracy and robustness of the method proposed in this embodiment under different image noise levels, Gaussian noise of 0 to 2 pixels is used in the simulation. The experimental results are as Figure 7 and Figure 8As shown, it can be seen that the Our-3pt method is significantly more accurate than the other three methods when solving the translation vector. It not only has the smallest average error but also the smallest error fluctuation. For the solution of the rotation matrix, Nister-5pt performs the best, but the accuracy of Our-3pt is almost similar to that of Nister-5pt, and it has the advantage of solving more parameters with fewer point correspondences. Hartley-8pt and Hartley-7pt are easily affected by noise, and the error increases rapidly as the noise increases.
[0144] 3. Since it is difficult to achieve ideal planar motion, during the motion process, there may be small vibrations of the monocular camera on the y-axis. Consider the influence of these small vibrations on the accuracy of the algorithm. When selecting the value range of this small vibration, the modulus |t| of the distance between the monocular camera at the current moment and the first moment is used as the measurement unit, and it is set to 0 to 1%|t|. In Figure 9 it can be seen that both the rotation accuracy and translation accuracy of Our-3pt are significantly better than those of Hartley-7pt; the rotation accuracy of Our-3pt is better than that of Hartley-8pt when the vibration is lower than 0.5%|t|, and the translation accuracy is better than that of Hartley-8pt in the range of 0 to 1%|t|; although the rotation accuracy of Our-3pt is not as good as that of Nister-5pt, it is also within a reasonable range, and the translation accuracy of Our-3pt is better than that of Nister-5pt when the vibration is lower than 0.6%|t|.
[0145] The specific experimental results of the one-point method are as follows:
[0146] 1. Without considering the influence of image noise, the experimental results of numerical stability are as Figure 10 shown. Similar to the three-point method, all algorithms are executed 5000 times. It can be seen that when calculating the translation vector, the stability of Our-1pt is the best, followed by the stability of Nister-5pt, and the stability of Hartley-7pt is the worst.
[0147] 2. To measure the accuracy and robustness of the method proposed in this embodiment under different image noise levels, Gaussian noise of 0 to 2 pixels is used in the simulation. The experimental results are as Figure 11 shown. As the image noise increases, the errors of all algorithms increase, but the error growth of Our-1pt is relatively gentle and relatively small, indicating that it has strong anti-interference ability to image noise.
[0148] 3. Although it is assumed that the monocular camera only rotates around the y-axis, in fact, there may be certain angular noises on the x-axis and z-axis of the monocular camera. The noise around the x-axis is the pitch angle noise, and the noise around the z-axis is the roll angle noise. To study the influence of angular noise on the accuracy of the algorithm, the angular noise is set to 0° to 1°. The image noise is set to 1 pixel and remains unchanged. The experimental results are as Figure 12 shown. The error of Our-1pt increases with the increase of angular noise, but the error value is always lower than that of other algorithms, indicating that it has strong robustness in the face of angular noise and can maintain a high estimation accuracy.
[0149] Based on the same inventive concept, the embodiment of the present application also provides a three-view relative pose solving system for planar motion for implementing the three-view relative pose solving method under the above-mentioned involved planar motion. The implementation solutions provided by this system to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the three-view relative pose solving system for planar motion provided below can refer to the limitations on the three-view relative pose solving method for planar motion in the above text, and will not be repeated here.
[0150] In an exemplary embodiment, a three-view relative pose solving system for planar motion is provided, including:
[0151] An acquisition module, configured to collect views at different times by using a monocular camera disposed on a moving object during the planar motion of the moving object, so as to obtain a first view, a second view, and a third view.
[0152] An extraction module, configured to extract a preset number of groups of feature matching points from the first view, the second view, and the third view;
[0153] A first solving module, configured to solve the linear equations of the trifocal tensor by using the feature matching points when the number of groups of feature matching points is three;
[0154] A second solving module, configured to measure the yaw angle of each view by using an inertial measurement unit when the number of groups of feature matching points is one; and solve the linear equations of the trifocal tensor based on the yaw angle and the feature matching points;
[0155] A pose determination module, configured to determine the relative pose among the first view, the second view, and the third view according to the solving result.
[0156] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store views at different times. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a method for solving the relative pose of three views under planar motion.
[0157] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0158] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are realized.
[0159] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0160] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0162] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0163] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0164] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.
[0165] In this article, specific examples are used to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for solving relative pose of three-view images under planar motion, characterized in that: The method for solving the relative pose of three-view images under planar motion includes: During the planar motion of the moving object, a monocular camera arranged on the moving object is used to collect views at different times to obtain a first view, a second view and a third view; According to the projection matrix of the monocular camera and the trifocal tensor when the moving object moves in a plane, a linear equation group of the trifocal tensor is constructed; Extracting a preset number of groups of feature matching points from the first view, the second view, and the third view; When the preset number of groups of feature matching points is three, solving the linear equation group of the trifocal tensor using the feature matching points; When the preset number of groups of feature matching points is one group, the yaw angle of each view is measured by using an inertial measurement unit; and the linear equation group of the trifocal tensor is solved based on the yaw angle and the feature matching points; The relative positions among the first view, the second view and the third view are determined according to the solution results.
2. The method for solving relative pose of three-view images under planar motion according to claim 1, characterized in that: Extracting a preset number of feature matching points from the first view, the second view, and the third view specifically includes: A scale-invariant feature transformation algorithm is used to extract a preset number of groups of feature matching points in the first view, the second view, and the third view.
3. The method for solving relative pose of three-view images under planar motion according to claim 1, characterized in that: The projection matrix is: Among them, t k is the translation vector of the kth view, P k is the projection matrix of the monocular camera corresponding to the kth view, θ k is the yaw angle of the kth view, and The subscript y in the figure indicates that the monocular camera rotates around the Y axis. is the cosine value of the yaw angle of the kth view, is the sine value of the yaw angle of the kth view, is the translation displacement of the kth view along the x-axis, is the translation displacement of the kth view along the z-axis.
4. The method for solving relative pose of three-view images under planar motion according to claim 1, characterized in that: The linear equations of the trifocal tensor are: in, All are known coefficients, Q1~Q 12 All are unknown.
5. The method for solving relative pose of three-view images under planar motion according to claim 4, characterized in that: When the preset number of groups of feature matching points is three, solving the linear equation group of the trifocal tensor using the feature matching points specifically includes: Based on the feature matching points, a singular value decomposition method is used to solve the linear equations of the trifocal tensor.
6. The method for solving relative pose of three-view images under planar motion according to claim 5, characterized in that: The relative posture includes a translation vector and a yaw angle; taking the coordinate system of the first view as the world coordinate system, the yaw angle of the second view is the yaw angle of the second view relative to the first view, and the translation vector of the second view is the translation vector of the second view relative to the first view; The yaw angle of the third view is the yaw angle of the third view relative to the first view, and the translation vector of the third view is the translation vector of the third view relative to the first view; The expressions of the translation vector t2 and yaw angle θ2 of the second view are: in, is the cosine value of the yaw angle θ2 of the second view; is the sine value of the yaw angle θ2 of the second view; The expressions of the translation vector t3 and yaw angle θ3 of the third view are: in, is the cosine of the yaw angle of the third view θ3; is the sine of the yaw angle of the third view θ3.
7. A system for solving relative poses of three-view images under planar motion, characterized in that: The three-view relative pose solving system under the plane motion comprises: A collection module, used for collecting views at different times by using a monocular camera arranged on the moving object during the planar motion of the moving object, to obtain a first view, a second view and a third view; A construction module, used for constructing a linear equation system of a trifocal tensor according to a projection matrix of a monocular camera and a trifocal tensor when a moving object performs planar motion; An extraction module, used to extract a preset number of feature matching points in the first view, the second view and the third view; A first solving module, used for solving the linear equation group of the trifocal tensor by using the feature matching points when the number of groups of feature matching points is three; A second solving module is used for measuring the yaw angle of each view by using an inertial measurement unit when the number of groups of feature matching points is one group; and solving the linear equation group of the trifocal tensor based on the yaw angle and the feature matching points; The posture determination module is used to determine the relative postures among the first view, the second view and the third view according to the solution result.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for solving the relative posture of three views under planar motion as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for solving the relative posture of three views under planar motion described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for solving the relative posture of three views under planar motion described in any one of claims 1 to 6 is implemented.