Camera Rotation Estimation Method for Indoor Environments Based on the Manhattan Assumption

By detecting the dominant plane and calculating the Manhattan coordinate system in an indoor environment, combining the optimization of the cost function and the solution of the rotation matrix, the accuracy and dependence problems of camera rotation estimation in an indoor environment are solved, and efficient and stable rotation tracking is achieved.

CN114463406BActive Publication Date: 2025-07-01BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210090415.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-07-01
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently estimate the rotational motion of the camera in indoor environments, especially in scenarios with low texture and complex occlusions, and has a high dependence on environmental structure.

Method used

By detecting the dominant plane on the depth image of the indoor environment, and calculating the structure of the Manhattan coordinate system on the Gaussian sphere, nonlinear optimization is performed using the cost function, and finally solving the rotational motion of the camera using the rotation matrix.

Benefits of technology

The accuracy of camera rotation estimation in indoor environments is improved, the dependence on environmental structure is reduced, and stable and efficient camera rotation tracking is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463406B_ABST
    Figure CN114463406B_ABST
Patent Text Reader

Abstract

A method for estimating the rotational motion of a camera in an indoor scene based on the Manhattan structure assumption, belonging to the field of computer vision. In this invention, (1) a new method is adopted for plane fitting. Three points are randomly selected from the three-dimensional point cloud data of the indoor environment depth image to generate a plane model. Through sequential probability ratio testing and LO-RANSAC optimization of the model, an accurate dominant plane is obtained; (2) lines are detected from the RGB image, and the lines are projected onto the Gaussian sphere to obtain the projected great circle and the normal vector of the great circle. At the same time, the dominant plane is also projected onto the Gaussian sphere in the form of a normal vector and cross-multiplied with the normal vector of the line projection to determine the structure of the Manhattan coordinate system; (3) a cost function is proposed to optimize the structure of the Manhattan coordinate system to obtain a more stable and accurate result; (4) the rotational motion of the camera is indirectly calculated through the rotation matrix between the Manhattan coordinate system and the camera coordinate system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention designs a method for estimating camera rotation in an indoor environment based on the Manhattan assumption. This method extracts plane and line features in the indoor environment, determines the structure of the Manhattan coordinate system using only a single plane and line, non-linearly optimizes the estimation result using constraint conditions, and then solves the rotation motion of the camera through the rotation matrix between the camera and the Manhattan coordinate system, improving the accuracy of camera rotation estimation in the indoor environment and the adaptability of the algorithm to the environment. Background Art

[0002] The key issue in indoor positioning and navigation is to determine the position of the camera from the current environment. Some indoor robots use lidar sensors to obtain environmental data. With the progress of optical technology and the improvement of computing power, vision-based methods have attracted people's interest. Compared with the relatively expensive laser technology, in some civilian fields, vision-based positioning technology can greatly reduce costs. During the positioning process, the rotation motion of the camera has a greater impact on positioning than the translational motion of the camera. Although the indoor environment has complex occlusions and low texture, the objects and structures in the environment usually exhibit a high degree of organization in the form of orthogonal and parallel planes. The introduction of the Manhattan assumption makes it more convenient to estimate the rotation of the camera in the indoor scene.

[0003] Currently, the estimation methods for the Manhattan coordinate system in indoor scenarios are mainly divided into two categories: One is to estimate the orthogonal vanishing points in the two-dimensional image domain. This method finds the vanishing points through the perspective relationship in the image plane. However, due to the lack of necessary depth information, it is difficult to impose strict constraints on the orthogonality of the vanishing points, and the efficiency is also relatively low. The other method is to estimate three orthogonal main directions in the three-dimensional domain. By making assumptions about the Manhattan results in space and combining the necessary depth information for search and verification, relatively satisfactory results can be obtained. However, these studies require prior information of the camera, and the estimation cost for a single image is also relatively high. In recent years, a method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption has been proposed. For example, the Manhattan frame probabilistic hybrid algorithm proposed by Straub et al. (J. Straub, O. Freifeld, G. Rosman, J. J. Leonard and J. W. Fisher, “The Manhattan Frame Model—Manhattan World Inference in the Space of Surface Normals,” IEEE Trans on Pattern Analysis and Machine Intelligence, vol. 40, no. 1, pp. 235-249, Jan. 2018) derives a simple MAP inference algorithm and an algorithm based on Gibbs sampling using the surface normal distribution of the scene. This algorithm uses Metropolis-Hastings to split / merge the surface normal vectors and adjusts the number of MFs to best capture the surface normal distribution of the scene. However, this method has a high dependence on the environment. When there are not enough orthogonal planes or lines in the environment, it will cause the algorithm estimation to fail.

[0004] To solve the limitations of the current research algorithms and improve the estimation efficiency, the present invention proposes a method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption. Using the spatial position relationship between a single plane and a single line existing in the indoor environment, the structure of the Manhattan coordinate system is calculated on the Gaussian sphere. Secondly, a cost function is used to optimize the estimation result of the Manhattan coordinate system, and the direction of the camera is solved through the rotation matrix between the Manhattan coordinate system and the camera coordinate system, and the rotation of the camera is stably tracked in a continuous video sequence. Summary of the Invention

[0005] The present invention proposes a method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption, which solves the problems of low texture and complex occlusions in the indoor environment that affect the accuracy of camera rotation estimation, reduces the dependence of the algorithm on the indoor environment structure, and improves the drift-free rotation estimation accuracy of a 3-degree-of-freedom camera. Specifically, first, the present invention detects the dominant plane of the point cloud data of the depth image, and then detects straight lines in the RGB image. On the Gaussian sphere, the structure of the Manhattan coordinate system is estimated using only a single plane and a single straight line. Then, using the constraint relationship between the vanishing point and the straight line on the image plane, a cost function is proposed to optimize the calculation result of the Manhattan coordinate system. Finally, the direction of the camera is obtained using the rotation matrix between the camera and the Manhattan coordinate system.

[0006] To achieve the above object, the present invention provides the following solution:

[0007] A method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption, the method comprising:

[0008] Step 1: Detect the dominant plane of the depth map of the indoor environment;

[0009] Step 2: Calculate the structure of the Manhattan coordinate system on the Gaussian sphere;

[0010] Step 3: Perform non-linear optimization using the cost function;

[0011] Step 4: Solve the camera direction using the rotation matrix.

[0012] The detection of the dominant plane of the depth map of the indoor environment is specifically as follows:

[0013] For the point cloud data of the depth image, three data points are randomly selected and plane fitting is performed on them. Substitute into the spatial equation of the plane to solve for the undetermined coefficients. First, the model is verified for degeneracy to ensure that the three selected data points are not collinear, and then a sequential probability ratio test is performed on the plane model:

[0014]

[0015] where λ is the result of the test for all data points, n is the number of all data, p(x r |H b ) is the probability that the x r -th data point does not conform to the model, b is bad, p(x r |H g ) is the probability that the x rThe probability that a data point conforms to the model. Let g be good. For each data point, the probability of conforming to the model and the probability of not conforming to the model are tested in sequence. Finally, if λ is greater than the given threshold of 4.4, the model is eliminated, and a new sample is selected from the point cloud data for fitting. If λ is less than or equal to the given threshold of 4.4, the next step is continued. For the model that passes the sequential probability ratio test, Lo-RANSAC optimization is continued. Using the inliers that conform to the model returned from the previous sequential probability ratio test, resampling is performed to generate a plane model. Set the number of iterations to 10 to 20 times, and then select the optimal local result as the improved result. Repeat the above process iteratively, with an iteration upper limit of 1 million times, until the proportion of inliers reaches the set threshold of 0.2, and finally obtain the current dominant plane. Project the dominant plane onto the Gaussian sphere in the form of a normal vector, and use the Mean Shift algorithm to continuously track the surface normal vectors around the normal vector VP1 of the dominant plane.

[0016] The structure for calculating the Manhattan coordinate system on the Gaussian sphere specifically includes the following steps:

[0017] Step 1): For the input RGB image, use the LSD algorithm for line detection. Take the line data in the image as samples, and select one line for calculation each time.

[0018] Step 2): Project the selected line onto the Gaussian sphere. The projection on the Gaussian sphere is a great circle, represented by the normal vector V0 of the great circle. The normal vector of the dominant plane is one axis VP1 of the Manhattan coordinate system. Take the cross product of the normal vector of the dominant plane and the normal vector of the great circle to obtain the second axis VP3 of the Manhattan coordinate system. According to the orthogonality of the coordinate system and vector cross product, take the cross product of the two obtained axes again to get the last coordinate axis VP2 of the Manhattan coordinate system.

[0019] Step 3): Perform a sequential probability ratio test and Lo-RANSAC optimization on the Manhattan coordinate system model obtained in Step 2. Repeat Steps 1 and 2 above to find the Manhattan coordinate system model supported by the most line data points, and stop when the proportion of inliers reaches the set threshold of 0.8. In the continuous video sequence estimation, once the Manhattan coordinate system structure is determined, it remains unchanged.

[0020] The use of the cost function for nonlinear optimization means using the average orthogonal distance between the vanishing point and the associated line in the image plane to constrain the structure of the Manhattan coordinate system. Three main factors are considered: the average orthogonal distance between the endpoints of the line detected in the RGB image and the vanishing point, the distance from the vanishing point to the midpoint of the line, and the length of the line. The orthogonal distance d from one endpoint of the i-th line to the vanishing point i,1 can be expressed as:

[0021]

[0022] Among them, A i,1 , B i,1 , C i,1 are the coefficients of the straight line equation formed by the vanishing point and the end point of the straight line. The average orthogonal distance d i,k of the two end points of the i-th straight line is:

[0023] d i,k =(d i,1 +d i,2 ) / 2 (3)

[0024] Use the following cost function to optimize the calculation result of the Manhattan coordinate system:

[0025]

[0026] Among them, ω1, ω2, and ω3 are 0.8, 0.1, and 0.1 respectively. Among them, N k (k∈{2,3}) is the number of line segments associated with VP2 and VP3 respectively. d i,k is the average orthogonal distance between the i-th line segment and the k-th corresponding vanishing point. τ is a pixel point. d vpo is the distance between the vanishing point and the midpoint of the straight line, l i is the length of the i-th line segment associated with the vanishing point, max(l i ) is the maximum value among the lengths of N k line segments. Minimize the cost function through the Levenberg-Marquardt algorithm to find the largest consistent line set, and optimize the calculation result of the Manhattan coordinate system.

[0027] The method of solving the camera direction using the rotation matrix mentioned above means that in each frame, the rotation matrix between the camera and the Manhattan coordinate system is obtained based on the Manhattan coordinate system. The direction of the Manhattan coordinate system is always fixed relative to the world coordinate system, so as to solve the rotation of the camera relative to the world coordinate system.

[0028] Beneficial effects:

[0029] The present invention proposes a method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption. First, the dominant plane is detected in the depth image of the indoor environment through the proposed method; secondly, the structure of the Manhattan coordinate system is calculated on the Gaussian sphere, and the orthogonality of the Manhattan coordinate system is ensured by using the property of vector cross product. Then, the direction of the Manhattan coordinate system is nonlinearly optimized using the cost function. Finally, the rotation motion of the camera is solved using the rotation matrix. This method only uses a single plane and a single straight line in the environment, reduces the dependence on the environmental structure, and improves the estimation accuracy, and has good robustness and stability. Brief description of the drawings

[0030] Figure 1 is the flow chart of the camera rotation estimation method in the indoor environment based on the Manhattan assumption provided by the present invention;

[0031] Figure 2 is the schematic diagram of the processing flow of the embodiment of the camera rotation estimation method in the indoor environment based on the Manhattan assumption provided by the present invention; Detailed implementation manners

[0032] The object of the present invention is to provide a camera rotation estimation method in the indoor environment based on the Manhattan assumption, which is used to solve the problem that the rotation estimation of the camera often fails in the indoor environment with occlusion and low texture. By extracting the plane and straight line features in the environment, calculating the direction of the Manhattan coordinate system on the Gaussian sphere, and taking this as a benchmark, calculate the rotation matrix of each frame in the subsequent frames to obtain the rotation motion of the camera. The camera rotation estimation method in the indoor environment based on the Manhattan assumption of the present invention can not only stably estimate the rotation of the camera, but also effectively reduce the dependence on the environmental structure while ensuring efficiency and accuracy, and can stably and accurately estimate and track the rotation motion of the camera even in a complex environment.

[0033] The present invention will be described in detail below with reference to the accompanying drawings. It should be noted that the described embodiments are only intended to facilitate the understanding of the present invention and do not impose any limitation on it.

[0034] Figure 1 is the flow chart of the camera rotation estimation method in the indoor environment based on the Manhattan assumption provided by the present invention; Figure 2 is the schematic diagram of the processing flow of the embodiment of the camera rotation estimation method in the indoor environment based on the Manhattan assumption provided by the present invention.

[0035] The camera rotation estimation method in the indoor environment based on the Manhattan assumption provided by the present invention specifically includes:

[0036] Step 1: Detect the dominant plane of the depth image of the indoor environment;

[0037] In order to make the detected dominant plane more accurate, the present invention adds two tests in the process of plane fitting. First, sample and fit the point cloud data, and then perform degradation verification on the generated model to ensure that the three selected points are not collinear. Perform sequential probability ratio test and Lo-RANSAC optimization on the plane model. Repeat this process until the proportion of the inliers returned exceeds a given threshold. In the present invention, it is 0.2. This method can quickly and accurately detect the dominant plane in the current environment and track it in the subsequent frames.

[0038] Step 2: Calculate the structure of the Manhattan coordinate system on the Gaussian sphere;

[0039] In the present invention, the LSD algorithm is used to detect straight lines in an RGB image, and the detected straight line information is used as a sample. By using the method for dominant plane detection in step 1, each time a straight line is selected and projected onto a Gaussian sphere, and the resulting projection is a great circle. The cross product of the normal vector V0 of the great circle and the normal vector VP1 of the dominant plane is taken to obtain the second coordinate axis VP3 of the Manhattan coordinate system. Based on the orthogonality of the coordinate system, the cross product of VP1 and VP3 can be used to obtain VP2. For a determined Manhattan coordinate system model, a sequential probability ratio test and Lo-RANSAC optimization are performed on it. The above steps are repeated until the inlier ratio of the straight lines supporting the Manhattan model reaches 0.8, at which point the iteration stops.

[0040] Step 3: Optimize the calculation result using a cost function;

[0041] After solving the structure of the Manhattan coordinate system, according to the constraint relationship between the vanishing point and the straight line in the image space, a cost function is proposed to nonlinearly optimize the calculation result. The Levenberg-Marquardt algorithm is used to minimize the cost function to find the largest consistent line set.

[0042] In the embodiment of the present invention, through the impact analysis of each element, the weights ω1, ω2, and ω3 are respectively assigned values of 0.8, 0.1, and 0.1.

[0043] Step 4: Solve for the camera rotation using a rotation matrix;

[0044] In each frame of a continuous video sequence, the rotation matrix between the camera coordinate system and the Manhattan coordinate system is solved, so as to solve for the rotation of the camera coordinate system relative to the world coordinate system.

[0045] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any transformation and substitution that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for estimating camera rotation in an indoor environment based on the Manhattan assumption, characterized in that, It includes the following steps: Step 1: Detection of the dominant plane: Perform plane fitting on the point cloud data of the indoor environment depth image, detect the plane model supported by the most data points, project the plane onto the Gaussian sphere in the form of a normal vector, use the normal vector of the plane as one coordinate axis of the Manhattan coordinate system, and continuously track the normal vector of the dominant plane throughout the process using the Mean Shift algorithm; Step 2: Calculate the orthogonal directions of the Manhattan coordinate system on the Gaussian sphere; Use the LSD algorithm to perform line detection on the RGB image, project the line onto the Gaussian sphere to obtain a circle, perform a cross product of the normal vector V0 of the circle and the normal vector VP1 of the dominant plane in Step 1 to obtain the second coordinate axis VP3, and then perform a cross product of the two obtained coordinate axes to obtain the third coordinate axis VP2; Step 3: Perform nonlinear optimization using the cost function; Propose a cost function to optimize the estimation result of the Manhattan coordinate system, and use the Levenberg-Marquardt algorithm to minimize the cost function to obtain the largest consistent line set; Step 4: Solve the camera rotation using the rotation matrix; Once the Manhattan coordinate system is determined, it remains unchanged. In each frame, solve the rotation matrix between the camera coordinate system and the Manhattan coordinate system, thereby solving the rotation of the camera coordinate system relative to the world coordinate system; The specific steps of Step 3 are as follows: The cost function consists of three elements: the average orthogonal distance d between the line segment endpoints and the vanishing point i,k , the distance d between the vanishing point and the midpoint of the line segment vpo , and the line segment length l i ; The average orthogonal distance between the endpoints d1 and d2 of the i-th line segment on the image plane and the vanishing point is; d i,k = (d i,1 + d i,2 ) / 2 (2) where d i,1 is; Among them, A i,1 , B i,1 , C i,1 are the coefficients of the linear equation formed by the vanishing point and the end point of the line; d i,2 and d i,1 are calculated in the same way, and the following cost function is proposed; Among them, ω1, ω2 and ω3 are 0.8, 0.1 and 0.1 respectively; among them, N k is the number of line segments associated with VP2 and VP3, respectively, where k∈{2,3}; d i,k is the average orthogonal distance between the i-th line segment and the k-th corresponding vanishing point; τ is a pixel; d vpo is the distance between the vanishing point and the midpoint of the line, l i is the length of the i-th line segment associated with the vanishing point, max(l i ) is N k The longest value among the line segments; use the Levenberg-Marquardt algorithm to solve the maximum consistent line set that minimizes the cost function.

2. The method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption according to claim 1, wherein The specific steps of the dominant plane detection described in Step 1 are as follows: Perform plane fitting on the point cloud data of the input depth image. First, randomly select three points as samples. First, ensure that the three randomly selected points are not collinear, substitute them into the spatial equation of the plane to determine the coefficients, and perform a sequential probability ratio test on the tentative plane model; where λ is the calculation result of all data, n is the number of all data points, p(x r |H b ) is the probability that the x r -th data point does not conform to the model, b represents bad, p(x r |H g ) is the probability that the x r -th data point conforms to the model, g represents good. The probability that each data point conforms to the model and the probability that it does not conform to the model are tested in turn. Finally, if λ is greater than the given threshold 4.4, the model is eliminated and a new sample is selected for fitting. If λ is less than or equal to the given threshold 4.4, the result is continued to be optimized by Lo-RANSAC; Use the inliers that conform to the model returned from the sequential probability ratio test in the previous step to resample and generate a plane model. Set 10 to 20 iterations, and then select the optimal local result as the improved result; Repeat all the above processes until the proportion of inliers reaches the set threshold 0.2, and finally obtain the current dominant plane; Project the dominant plane onto the Gaussian sphere, which is represented by the normal vector VP1. In order to align the scene in the subsequent frames with the dominant plane detected in the first frame, the Mean Shift algorithm is used to track the surface normal vectors around the normal vector of the dominant plane.

3. A method for estimating camera rotation in an indoor environment based on the Manhattan assumption according to claim 1, characterized in that, The specific steps of calculating the orthogonal directions of the Manhattan coordinate system on the Gaussian sphere described in Step 2 are as follows: For the RGB image, use the LSD algorithm to perform line detection on it, use the detected line information as a sample, and use the same method as the dominant plane detection. Each time, select a line and project it onto the Gaussian sphere, and the projection obtained is a great circle; Perform a cross product of the normal vector V0 of the great circle and the normal vector VP1 of the dominant plane to obtain the second coordinate axis VP3 of the Manhattan coordinate system. Based on the orthogonality of the coordinate system, perform a cross product of VP1 and VP3 to obtain VP2. For the determined Manhattan coordinate system model, perform a sequential probability ratio test and Lo-RANSAC optimization on it. Repeat the above steps, and the upper limit of the iteration times is 1 million times. Stop the iteration when the inlier ratio of the line is higher than 0.

8.

4. A method for estimating the rotation of a camera in an indoor environment based on the Manhattan assumption according to claim 1, characterized in that, The specific steps of solving the camera rotation using the rotation matrix described in Step 4 are as follows: The direction of the Manhattan coordinate system is always fixed relative to the world coordinate system. In each frame, based on the Manhattan coordinate system, find the rotation matrix between the camera and the Manhattan coordinate system, thereby solving the rotation of the camera relative to the world coordinate system.

Citation Information

Patent Citations

  • Scene reconstruction method based on Manhattan hypothesis

    CN107292956A

  • Indoor three-dimensional reconstruction method based on panorama

    CN110782524A