A method for estimating the surface normal vector of a 3D model based on an end-to-end convolutional model

Through the method based on end-to-end convolution model, the traditional normal vector estimator has been solved in accuracy and speed, and efficient and accurate normal vector estimation is achieved, which is suitable for tasks such as autonomous driving, three-dimensional reconstruction and semantic segmentation in computer vision.

CN116188732BActive Publication Date: 2025-07-22TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211574413.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-07-22
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing normal vector estimation methods cannot meet the high-demand tasks in terms of accuracy and speed. The traditional methods consume time, have high data volume requirements and unstable effects. The training complexity of deep learning methods is high and the universality is low.

Method used

Using an end-to-end convolution model method, we calculate three-dimensional coordinates by obtaining depth or disparity information, define convolution templates in horizontal and vertical directions, use gradient matrix to estimate part of normal vectors, and combine neighboring points and three-dimensional coordinates to estimate the complete normal vector using multiple normal vector estimators.

Benefits of technology

It realizes high-precision and fast normal vector estimation, with a calculation speed of about 500 times and a significant improvement in accuracy, and is suitable for a variety of tasks in computer vision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188732B_ABST
    Figure CN116188732B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model, which includes the following steps: obtaining depth information or calculating depth information according to the obtained disparity information, and calculating the three-dimensional coordinates of each point of the three-dimensional model according to the depth information and known camera parameters; defining convolution templates for differentiating the matrix in the horizontal and vertical directions, and calculating the gradient matrices of the depth information in the horizontal and vertical directions; estimating partial normal vectors based on the gradient matrices; calculating neighborhood points in multiple directions based on the depth information and convolution matrices; combining the three-dimensional coordinates, partial normal vectors and neighborhood points, and using an end-to-end normal vector estimator to estimate the complete normal vector. Compared with the prior art, the present invention overcomes the problems of low operation accuracy, long time consumption and unstable effect of traditional normal vector estimators, and realizes high-precision and fast normal vector estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of normal vector estimation under computer vision, and particularly to a method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model. Background Art

[0002] In the field of computer vision research, an intelligent agent attempts to extract information related to intelligent decision-making from images or multi-dimensional data to achieve corresponding target tasks, and the surface normal vector of a three-dimensional model is a very important information feature that needs to be extracted in computer vision. With continuous development, the intelligent agent can already estimate the surface normal vector information based on image information to achieve feature extraction of the image, and the extracted information helps the intelligent agent to complete large tasks such as semantic segmentation, object recognition, and three-dimensional reconstruction.

[0003] Currently, the mainstream traditional normal vector estimators have many defects and cannot meet the requirements of high-demand tasks in terms of accuracy and speed.

[0004] First, normal vector estimators using the local point cloud modeling idea, such as PlaneSVD, PlanePCA, VectorSVD, LeastSquareSVD, and QuadSVD, use three-dimensional point clouds as the data type and complete the normal vector estimation by optimizing the model. The main time-consuming of this type of method lies in the SVD singular value decomposition of large matrices, which not only requires a relatively high data volume of three-dimensional point clouds, but also has a huge preprocessing step for constructing an undirected graph, and the operation speed and accuracy cannot reach a relatively high level.

[0005] Second, normal vector estimators using the weight averaging method, such as the AreaWeighted and AngleWeighted methods, assign different weights to the neighborhood small triangular planes of data points and perform weighted averaging on them to obtain the final normal vector. This method is even more time-consuming and has very high requirements for the data volume, and the effect of normal vector estimation is unstable. At the same time, since the normal vector of a point on the surface is calculated by the cross product of two tangent vectors of this point, when the tangent vectors of this point are linearly correlated, this algorithm may exhibit the situation of a degenerate normal vector.

[0006] Finally, deep learning neural networks are also used for normal vector estimation, and this method also has some drawbacks: the time complexity of network training is high, the network universality is low, and it is difficult to obtain the ground truth and manual annotation. Therefore, this method cannot be well promoted and applied in most task scenarios. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model to achieve fast and accurate estimation of the normal vector.

[0008] The object of the present invention can be achieved by the following technical solutions:

[0009] A method for estimating the surface normal vector of a 3D model based on an end-to-end convolutional model, comprising the following steps:

[0010] S1. Obtain depth information or calculate depth information according to the obtained disparity information, and calculate the three-dimensional coordinates of each point of the 3D model according to the depth information and known camera parameters;

[0011] S2. Define convolutional templates for differentiating the matrix in the horizontal and vertical directions, and calculate the gradient matrices of the depth information in the horizontal and vertical directions;

[0012] S3. Estimate partial normal vectors based on the gradient matrices;

[0013] S4. Calculate neighborhood points in multiple directions based on the depth information and the convolutional matrix;

[0014] S5. Combine the three-dimensional coordinates, partial normal vectors, and neighborhood points, and use an end-to-end normal vector estimator to estimate the complete normal vector.

[0015] The convolutional templates are as follows:

[0016]

[0017] where G x is the convolutional template in the horizontal direction, and G y is the convolutional template in the vertical direction, and the elements satisfy:

[0018] a0 = ∑b n +... + b1 + a0 + a1 +... + a n

[0019] The value of a0 is 0, and a1....a n take integers 1, 2,..., n in sequence, and b1....b n take the corresponding opposite numbers in sequence.

[0020] The gradient matrices are as follows:

[0021]

[0022] where T x is the gradient matrix of the depth information in the horizontal direction, and T y is the gradient matrix of the depth information in the vertical direction, and z is the depth information.

[0023] The method for estimating partial normal vectors based on the gradient matrices is:

[0024]

[0025]

[0026] Among them, n x and n y are the estimated normal vectors in the x and y directions, f x and f y are the focal lengths of the camera at the pixel points, and u and v represent the horizontal and vertical positions of the two-dimensional matrix respectively.

[0027] The convolution matrix is a square matrix composed only of elements 0 and 1, where there is only 1 element 1, and the positions of element 1 in different convolution matrices are not repeated.

[0028] Calculating the neighborhood points in multiple directions based on the depth information and the convolution matrix specifically means multiplying the depth information by the convolution matrix to obtain the neighborhood points, and the number of convolution matrices is the same as the number of neighborhood points.

[0029] The normal vector estimator includes an end-to-end convolutional PCA algorithm, an end-to-end convolutional SVD algorithm, an end-to-end convolutional area weight averaging algorithm, and an end-to-end convolutional angle weight averaging algorithm.

[0030] The end-to-end convolutional PCA algorithm and the end-to-end convolutional SVD algorithm use the idea of local point cloud modeling to optimize the model composed of neighborhood points using the least squares method to obtain the normal vector:

[0031]

[0032] Among them, n x and n y are the normal vectors in the x and y directions estimated in step S3, (x, y, z) are the three-dimensional coordinates of each point of the three-dimensional model, and k is the total number of points.

[0033] The end-to-end convolutional area weight averaging algorithm and the end-to-end convolutional angle weight averaging algorithm assign different weights to the neighborhood small triangular planes of the data points and perform weighted averaging to obtain the estimated:

[0034]

[0035] Among them, is the weight, n zi is the estimated value of the normal vector in the z direction under a neighborhood small triangular plane of a data point, and i can take values from 1 to 4 or 8 respectively according to the number of neighborhood points.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] (1) The present invention uses structured depth sensor data such as depth information or disparity information containing explicit neighborhood relationships, and uses convolution operations to extract corresponding neighborhood points. The data at the input end is simple and easy to obtain, and the preprocessing operation of the input end data is also very simple, with little time consumption.

[0038] (2) Based on the convolution template, the present invention can directly calculate the normal vector using the gradient of depth or disparity, with high precision and simple calculation.

[0039] (3) The four end-to-end normal vector estimators adopted by the present invention are estimators with high precision and high speed, with high parallel efficiency in calculation, and significantly superior to most existing normal vector estimators in the comprehensive index of speed and precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of the method of the present invention;

[0041] Figure 2 is a diagram of the normal vector estimation operation process;

[0042] Figure 3 is a schematic diagram of the neighborhood composition when estimating the normal vector using the weighted average method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0044] This embodiment provides a method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolution model, as Figure 1 shown, including the following steps:

[0045] S1. Obtain depth information z or calculate the depth information according to the obtained disparity information d p and calculate the three-dimensional coordinates (x, y, z) of each point of the three-dimensional model according to the depth information and the known camera parameters.

[0046] The depth map and the disparity map are inversely proportional, that is: This step can be completed after obtaining any one of the structured depth data of depth or disparity and the camera parameters.

[0047] This step requires obtaining depth information or disparity information, which is stored in a two-dimensional discrete matrix, and also requires obtaining the internal parameter matrix K of the perspective camera:

[0048]

[0049] If the depth z is known, the three-dimensional coordinates are calculated as follows:

[0050]

[0051]

[0052] where u and v refer to the pixel position matrices of the coordinates of the 3D object projected onto the 2D photo in the horizontal and vertical directions under the perspective camera model; is the image center; f x and f y are the focal lengths of the camera in pixels.

[0053] If the parallax d p is known, then only by taking the reciprocal of d p can the depth z be obtained, and then the x-axis and y-axis coordinates can be calculated using the above three-dimensional coordinate calculation formula.

[0054] In this embodiment, given the depth information of the three-dimensional model and the camera internal parameter matrix, the three-dimensional coordinates of each point of the model are calculated. Further, in this embodiment, the depth information is converted into a two-dimensional discrete matrix of size 640×480 and stored.

[0055] S2. Define the convolution templates for differentiating the matrix in the horizontal and vertical directions, and calculate the gradient matrices T x and T y in the horizontal and vertical directions of the depth information.

[0056] This step uses the idea of fractional calculus to implement the design of the convolution template for the discrete two-dimensional matrix.

[0057] Denote the first-order derivative operator of f(x) in the direction of increasing x as Then the nth-order derivative operator is Using the derivative operator, the nth derivative of f(x) can be calculated as:

[0058]

[0059] In practical applications, it is necessary to consider both the directions of increasing and decreasing x, and the final differential operator is:

[0060]

[0061] From this, the convolution template is obtained as:

[0062]

[0063] where G x is the convolution template in the horizontal direction, G y is the convolution template in the vertical direction, and the elements satisfy:

[0064] a0 = ∑b n +...+b1+a0+a1+...+a n

[0065] In this embodiment, the value of a0 is 0, and a1....a n take integer values 1, 2,..., n in sequence, and b1....b n take the corresponding opposite numbers in sequence.

[0066] The sizes of these two convolution templates can be 3×3, 5×5, etc., and are selected according to the image ratio.

[0067] The gradient matrix is:

[0068]

[0069] where T x is the gradient matrix of the depth information in the horizontal direction, and T y is the gradient matrix of the depth information in the vertical direction, and z is the depth information.

[0070] In this embodiment, the selected convolution template is a 3×3 kernel to calculate the gradients of the reciprocal of the depth matrix in the horizontal and vertical directions:

[0071]

[0072] S3. Estimate partial normal vectors based on the gradient matrix:

[0073]

[0074]

[0075] where n x and n y are the estimated normal vectors in the x and y directions, f x and f y are the focal lengths of the camera at the pixel points, and u and v represent the horizontal and vertical positions of the two-dimensional matrix respectively.

[0076] S4. Calculate neighborhood points in 8 directions based on the depth information and the convolution matrix.

[0077] The convolution matrix is a square matrix composed of only elements 0 and 1, where there is only 1 element 1, and the positions of the element 1 in different convolution matrices do not repeat.

[0078] The specific calculation of neighborhood points in multiple directions based on the depth information and the convolution matrix is to multiply the depth information by the convolution matrix to obtain the neighborhood points, and the number of convolution matrices is the same as the number of neighborhood points.

[0079] In this step, neighborhood points in multiple directions are obtained through convolution of matrices using different convolutional kernels, which improves the speed of the normal vector estimation process.

[0080] The above steps S1 - S4 are the end - to - end convolutional model proposed by the present invention for realizing the estimation of the normal vector, and S5 is the process of estimating and thus improving the normal vector.

[0081] S5. Combine the three - dimensional coordinates, partial normal vectors, and neighborhood points, and use an end - to - end normal vector estimator to estimate the complete normal vector.

[0082] The normal vector estimator includes the end - to - end convolutional PCA algorithm (EToEConvPCA), the end - to - end convolutional SVD algorithm (EToEConvSVD), the end - to - end convolutional area - weighted averaging algorithm (EToEConvArea), and the end - to - end convolutional angle - weighted averaging algorithm (EToEConvAngle).

[0083] The end - to - end convolutional PCA algorithm and the end - to - end convolutional SVD algorithm use the idea of local point cloud modeling to optimize the model composed of neighborhood points using the least - squares method to obtain the normal vector. Both of these methods use the neighborhood to fit a plane. For simplicity in calculation, the parameter d is omitted through mathematical formulas. EToEConvPCA uses matrix calculation to subtract the centroid of the data from the mean of the neighborhood points to omit the parameter d, and EToEConvSVD uses matrix k0 minus the center point of the original data to omit the parameter d.

[0084]

[0085]

[0086] The plane model is optimized using the least - squares method, and the final optimization equation is:

[0087]

[0088] For perform Lagrangian solution, and through mathematical derivation, the calculation formula for n z is obtained:

[0089]

[0090] where n x and n y are the normal vectors in the x - direction and y - direction estimated in step S3, (x, y, z) are the three - dimensional coordinates of each point in the three - dimensional model, and k is the total number of points.

[0091] The end-to-end convolutional area weight averaging algorithm and the end-to-end convolutional angle weight averaging algorithm assign different weights to the neighborhood small triangular planes of data points and perform weighted averaging to obtain the estimated n z :

[0092]

[0093] where w Δi is the weight and n zi is the estimated value of the normal vector in the z direction under a neighborhood small triangular plane of a data point. In this embodiment, i = 8.

[0094] Take two points in the neighborhood and the original center point to form a triangle in the order of upper left -> up -> upper right -> right -> lower right -> down -> lower left -> left -> upper left, and calculate the normal vector n zi of this triangular plane and assign different weights w Δi .

[0095] In this embodiment, n x and n y obtained in step S3 and the coordinate information of the neighborhood small triangular plane are used to estimate n zi :

[0096]

[0097] EToEConvArea takes the areas of different triangular planes as weights:

[0098] w Δi = l ai ·l bi

[0099] EToEConvAngle takes the inclination angles of the normal vectors obtained from different triangular planes as weights:

[0100] w Δi = arccos(l ai × z bi )

[0101] where l ai , l bi refer to two vectors l Δi = [x ai , y ai , z ai, z ai T , l bi = [x bi , y bi , z bi T . ​​

[0102] The operation process diagram based on the above steps is as Figure 2 shown. The schematic diagram of the selection operation for the neighborhood based on the weighted average method is as Figure 3 shown.

[0103] The test results show that: these four end-to-end normal vector estimators are approximately 500 times faster than the mentioned traditional algorithms in terms of speed. Among them, since 8 values in the convolution kernel k0 of EToEConvSVD are 0, the convolution operation can be completed quickly, achieving further optimization in terms of time; in terms of accuracy, these four normal vector estimators have excellent effects. Compared with the traditional algorithms, they can better handle the surface normal vector estimation at the model edges.

[0104] After the normal vector is estimated by using the method of the present invention, it can be applied in the field of computer vision, such as autonomous driving, 3D reconstruction, semantic segmentation, feasible region detection, etc., and can efficiently and accurately provide normal vector information to assist in the completion of tasks.

[0105] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model, characterized in that, It includes the following steps: S1. Obtain depth information or calculate depth information based on the obtained parallax information, and calculate the three-dimensional coordinates of each point of the three-dimensional model according to the depth information and known camera parameters; S2. Define the convolution templates for differentiating the matrix in the horizontal and vertical directions, and calculate the gradient matrix of the depth information in the horizontal and vertical directions; S3. Estimate partial normal vectors based on the gradient matrix; S4. Calculate neighborhood points in multiple directions based on the depth information and the convolution matrix; S5. Combine the three-dimensional coordinates, partial normal vectors, and neighborhood points, and use an end-to-end normal vector estimator to estimate the complete normal vector.

2. The 3D model surface normal vector estimation method based on an end-to-end convolutional model according to claim 1, wherein The convolution template is: Among them, G x is the convolution template in the horizontal direction, and G y is the convolution template in the vertical direction, where the elements satisfy: a0 = ∑b n +...+b1+a0+a1+...+a n 。 3. A method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model according to claim 2, characterized in that The value of a0 is 0, and a1…a n take the integer values 1, 2, …, n in sequence, and b1…b n take the corresponding opposite numbers in sequence.

4. The three-dimensional model surface normal vector estimation method based on an end-to-end convolutional model according to claim 2, wherein The gradient matrix is: Among them, T x is the gradient matrix of the depth information in the horizontal direction, and T y is the gradient matrix of the depth information in the vertical direction, where z is the depth information.

5. The three-dimensional model surface normal vector estimation method based on an end-to-end convolutional model according to claim 4, characterized in that, The method for estimating partial normal vectors based on the gradient matrix is: Among them, n x and n y are the normal vectors in the x and y directions obtained by estimation, f x and f y are the focal lengths of the camera at the pixel points, and u and v represent the horizontal and vertical positions of the two-dimensional matrix respectively.

6. The 3D model surface normal vector estimation method based on an end-to-end convolutional model according to claim 1, wherein The convolution matrix is a square matrix composed only of elements 0 and 1, where there is only 1 element 1, and the positions of element 1 in different convolution matrices do not repeat.

7. The three-dimensional model surface normal vector estimation method based on an end-to-end convolution model according to claim 1, characterized in that Specifically, calculating neighborhood points in multiple directions based on the depth information and the convolution matrix means multiplying the depth information by the convolution matrix to obtain neighborhood points, and the number of convolution matrices is the same as the number of neighborhood points.

8. The three-dimensional model surface normal vector estimation method based on an end-to-end convolutional model according to claim 1, characterized in that The normal vector estimator includes an end-to-end convolutional PCA algorithm, an end-to-end convolutional SVD algorithm, an end-to-end convolutional area weight averaging algorithm, and an end-to-end convolutional angle weight averaging algorithm.

9. A method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model according to claim 8, characterized in that The end-to-end convolutional PCA algorithm and the end-to-end convolutional SVD algorithm use the idea of local point cloud modeling to optimize the model composed of neighborhood points using the least squares method to obtain the normal vector n z : where n x and n y are the normal vectors in the x - direction and y - direction estimated in step S3, (x, y, z) are the three - dimensional coordinates of each point of the three - dimensional model, and k is the total number of points.

10. A method for estimating the surface normal vector of a three-dimensional model based on an end-to-end convolutional model according to claim 8, characterized in that, The end-to-end convolutional area weight averaging algorithm and the end-to-end convolutional angle weight averaging algorithm assign different weights to the neighborhood small triangular planes of data points and perform weighted averaging to obtain the estimated n z : where, w Δi is the weight, and n zi is the estimated value of the z-direction normal vector under a small triangular plane in the neighborhood of a data point.

Citation Information

Patent Citations

  • Indoor scene identifying method based on point cloud fragment division

    CN102930246A

  • Point cloud feature point detection method and cloud point feature extraction method

    CN108010116A