Human body three-dimensional reconstruction method based on binocular stereo vision and point cloud
By adopting the cost calculation method of median and gradients, outlier value removal and hierarchical stereo matching optimization strategy in binocular stereo vision three-dimensional reconstruction technology, the problem of stereo matching instability under uneven noise and lighting conditions is solved, and a more efficient and accurate three-dimensional reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510265500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-04
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The existing three-dimensional reconstruction technology based on binocular stereo vision is poorly performed under uneven noise and light conditions, resulting in unstable stereo matching results, affecting the effect and speed of three-dimensional reconstruction.
The cost calculation based on median value and the cost calculation fusion of gradient Census are used, and outliers are eliminated in the cost aggregation stage to improve the robustness and calculation speed of the stereo matching algorithm. Through the hierarchical stereo matching optimization strategy, the disparity map is constructed from coarse to thin, and the reconstruction effect and speed are improved.
The stability and accuracy of the stereo matching results are significantly improved under uneven noise and light conditions, the speed and effect of three-dimensional reconstruction are improved, and resource waste is reduced.
Smart Images

Figure CN120219463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, especially the three-dimensional reconstruction technology based on binocular vision. Specifically, the present invention proposes a new method for three-dimensional reconstruction of the human body based on binocular stereo vision. This technology has better stereo matching results in the presence of noise and uneven illumination, aiming to improve the three-dimensional reconstruction effect and speed based on binocular stereo vision. This technology has potential applications in fields such as virtual reality and augmented reality.
[0002] The three-dimensional reconstruction technology based on binocular vision has great development prospects in various industries. By mainly using two inexpensive cameras, the three-dimensional reconstruction of spatial objects can be achieved, which enables obtaining a three-dimensional model of spatial objects with very little cost in practical applications and saves a large amount of expenses in many fields. Therefore, it has extremely great research significance to realize the three-dimensional reconstruction of the real human body through binocular vision technology. It can greatly reduce the cost of human body reconstruction, and by using cameras to perform three-dimensional reconstruction of the human body, it can well solve the problems of human privacy and human health and safety. Background Art
[0003] In recent years, with the continuous rapid development of the Internet, applications that help humans better carry out clothing, food, housing, and transportation through the Internet have been continuously developed, all aiming to enable people to obtain higher-quality services in a shorter time and at a closer distance. Online shopping is one of them. Through online shopping, it is convenient and fast to purchase clothes, but there will be a problem that the purchased clothes do not fit because the definition standards of clothing sizes are different, which has led to an increase in the return rate during online shopping. To solve this problem, virtual fitting applications have emerged. It uses existing computer technology to realize the fitting behavior of the human body in a virtual scene. Through virtual fitting, true home shopping can be realized. And an important research part in virtual fitting is to realize the three-dimensional reconstruction of the real human body.
[0004] One way is to achieve three-dimensional reconstruction of the human body through millimeter waves, and then perform virtual fitting through the reconstructed model. The advantages of this method are high accuracy and the ability to obtain a relatively realistic two-dimensional human body model, but there are some problems. First is the privacy problem. The three-dimensional model of the human body realized through millimeter waves is a three-dimensional nude model of the human body, and a complete reconstruction of all parts of the human body is achieved, so the privacy and security of people cannot be guaranteed. Then there is the problem of human health and safety. There are different opinions about the harm of millimeter waves to the human body. This method makes users face millimeter waves directly, which will arouse users' doubts about human health and safety, so it will also make it difficult for this method to be implemented. Finally, there is the price problem. The equipment using millimeter waves is relatively expensive, and the cost of human body reconstruction is high.
[0005] With the continuous development of computer technology, the method of three-dimensional scene reconstruction through computer stereo vision technology has been continuously applied in various fields, especially in the field of three-dimensional reconstruction of the human body, such as animation movie production, VR virtual reality, virtual fitting, etc. Binocular stereo vision technology is one of its important research directions. Stereo matching algorithms such as SGM, PatchMatch, and AD-Census algorithm have their own advantages and disadvantages in terms of efficiency and effect. The improved stereo matching technology of the present invention has strong robustness to noise and illumination, and improves the calculation speed of stereo matching through a hierarchical strategy and parallel optimization. Summary of the Invention
[0006] The present invention takes the pictures collected by the left and right cameras as input, and finally obtains the three-dimensional reconstructed human body model through steps such as stereo rectification and stereo matching. By fusing the cost calculation based on median and the cost calculation of gradient Census, and realizing outlier rejection in the cost aggregation stage, the stereo matching algorithm proposed by the present invention has better stereo matching results when there are noise and uneven illumination. The calculation speed of the stereo matching algorithm is improved through a hierarchical stereo matching optimization strategy.
[0007] The present invention performs three-dimensional reconstruction of the human body based on binocular stereo vision. First, the left and right cameras at the same height are calibrated. Because in the process of converting a two-dimensional picture to a three-dimensional object, the internal and external parameters of the camera are required to calculate the depth information. In order to satisfy the epipolar constraint under parallel views during stereo matching, the stereo rectification projection matrices corresponding to the two cameras need to be obtained, and then the parameters obtained in camera calibration are applied to our stereo matching method. The stereo matching method obtains the final disparity map through cost calculation, cost aggregation, disparity calculation, and disparity optimization, and generates a three-dimensional point cloud based on the disparity map for surface reconstruction.
[0008] Based on the above invention idea, the present invention proposes a method for three-dimensional reconstruction of the human body based on binocular stereo vision and point cloud, which specifically includes the following steps:
[0009] S1, taking pictures of the checkerboard calibration board at different angles by a binocular camera to obtain the internal and external parameter matrices of the camera and the stereo rectification projection matrix, and performing distortion correction on the input left and right human body images according to the matrix parameters.
[0010] S2, performing stereo matching on the input pictures that have completed distortion correction. Our invention is based on a cost calculation method based on median and gradient, and proposes an optimization method for outlier rejection in cost aggregation.
[0011] S3, constructing a disparity map in a coarse-to-fine manner through a hierarchical stereo matching strategy to optimize the operation speed of stereo matching.
[0012] S4. Obtain the three-dimensional point cloud data of the human body from the disparity map obtained by stereo matching, and reconstruct the surface of the three-dimensional point cloud data of the human body through the greedy projection triangulation algorithm.
[0013] A three-dimensional reconstruction method of the human body based on binocular stereo vision and point cloud. The purpose of step S1 is to first obtain the internal and external parameter matrices of the binocular camera used for three-dimensional reconstruction. For the input left and right human body images, use the internal and external parameter matrices to correct the distortion of the images caused by the camera. This step specifically includes the following sub-steps:
[0014] S11. Calculate and obtain the internal and external parameter matrices of the camera and the stereo rectification projection matrix. First, obtain the internal and external parameter matrices of the camera through the calibration principle of the monocular camera respectively.
[0015]
[0016] A represents the internal parameter matrix, N represents the external parameter matrix, (X W , Y W , Z W ) represents the coordinate value in the world coordinate system, (u, v) represents the coordinate value in the image pixel coordinate system, Z c represents the value of the Z-axis in the camera coordinate system, f represents the focal length of the camera, 1 / d x , 1 / d y respectively represent the contraction ratios of the image physical coordinate system to the U-axis and V-axis of the image pixel coordinate system, (u0, v0) is the moving distance, and θ represents the final pixel coordinate system and the angle deviating from the normal situation.
[0017] As Figure 1 shown, on the left is the checkerboard calibration board required for calibration, and on the right is the binocular camera for obtaining images. Take multiple pictures of the checkerboard calibration board through the binocular camera to obtain multiple pairs of pictures containing the checkerboard calibration board at different angles. Detect the corner points in the calibration board pictures through the corner point detection algorithm. As Figure 2 shown, find the corner points on the checkerboard, and these positions are equivalent to the coordinate points in the image pixel coordinate system. Consider the checkerboard plane as the plane where Z = 0 in the world coordinate system, then take the first corner point in the upper left corner of the checkerboard as the coordinate origin, and finally according to the side length of each black and white square on the checkerboard. Obtain the coordinate values of each corner point on the checkerboard in the world coordinate system. As Figure 3 shown in the mapping relationship diagram of the corner point detection algorithm, point P is the corner point on the checkerboard calibration board. Let the coordinate of P in the camera coordinate system O1-X1Y1Z1 of the left camera be X tc , and the coordinate of P in the camera coordinate system O2-X2Y2Z2 of the right camera be X rc, according to the transformation relationship between the world coordinate system and the camera coordinate system, it is mapped to P1 of the left camera and P2 of the right camera. After monocular camera calibration, the external parameter matrix [R1 t1] of the left camera and the external parameter matrix [R2 t2] of the right camera are obtained. As shown in the equation group:
[0018]
[0019] The position relationship R and T between the two cameras are obtained through the two external parameter matrices. It can be obtained from Equation (3):
[0020]
[0021] R is the rotation matrix of the camera, and T is the translation matrix of the camera. So far, we have obtained the internal and external parameter matrices of the camera and the position relationship of the binocular camera. Then, the distortion correction is completed after preprocessing the human body images collected by the binocular camera. First, we perform Gaussian blur on the collected input images to remove noise points, and then enhance the sharpness through Unsharpen Mask. Through preprocessing, we obtain two images with slightly less noise and better quality. Then, the image distortion is restored through distortion correction and epipolar correction.
[0022] A method for three-dimensional reconstruction of the human body based on binocular stereo vision and point cloud. The purpose of step S2 is to perform stereo matching after processing the left and right images collected by the binocular camera in S1. Find the corresponding points of the pixel points in the left image in the right image, that is, search for all points on the same row in the right image, compare the similarity, and determine the point with the highest similarity as the corresponding point. The disparity is calculated through the determination of the corresponding points. Therefore, the disparity map can be finally obtained through stereo matching. This step includes the following sub-steps:
[0023] S21, measure the similarity of the pixel values of the left and right images through cost calculation. Inspired by the AD-Census algorithm, our invention combines the advantages of the AD (Absolute Differences) algorithm and the Census algorithm. The cost value under the AD algorithm is obtained by calculating and comparing the sizes of individual pixel points. Then, our invention improves the cost calculation method under the Census algorithm, with the median pixel in the window as the core of the calculation and introduces gradient calculation as one of the bases for measuring the cost value. First is the first part of the cost calculation under the AD algorithm:
[0024]
[0025] where C AD (p,q) is the cost value of point p in the left image and point q in the right image under the AD algorithm, represents the gray value of p in the i-th channel of the left image, Let \(q\) represent the grayscale value of the \(i\)-th channel in the right figure. There are a total of three channels. Finally, calculate the average of the absolute values of the differences between the three channels of \(p\) and \(q\). The present invention calculates the cost of the second part based on binary codes. By taking the median of the pixels within the window as the comparison value, the corresponding binary code is generated. The calculation formula is as follows:
[0026]
[0027] Equation (6) represents the cost calculation method of point \(p\) in the picture based on the median binary code, where \(t\in N\). p It represents the calculation of the grayscale values within the window of point \(P\), while is a bit concatenation operator. That is, for each pixel within the window, calculate the corresponding bit value, and then connect them into a binary code through this symbol. For each pixel in the window, the corresponding bit value is obtained through the function \(\zeta\), where \(I\) mid (p) represents the median of the grayscale values of all pixels within the window centered on \(p\). \(I(t)\) represents the grayscale value of each pixel within the window. By comparing their magnitudes, the bit value corresponding to the pixel in the window is obtained. Connecting all the bit values together gives the binary code corresponding to the window of point \(p\).
[0028] S22. Introduce gradients for cost calculation. Perform gradient calculation by using the Sobel operator. The Sobel convolution factors are divided into convolution factors in the \(x\)-direction and \(y\)-direction, corresponding to Figure 4 the gradients in the two directions. Finally, fuse the gradients in the two directions to obtain the final gradient value. The specific gradient calculations in the \(x\)-direction and \(y\)-direction are shown in Equation (8). The gradient in the \(x\)-direction at \((u, v)\) is \(G\) x (u, v), and \(G\) y (u, v) represents the gradient in the \(y\)-direction. Average the gradient values in these two directions to obtain the final gradient value, as shown in Figure (9).
[0029]
[0030] Obtain the cost value based on the AD algorithm, the cost value based on the median binary code, and the cost value based on gradient calculation in the above manner. Fuse the cost values of the three parts to obtain the final cost value. The cost value is the Hamming distance of the binary code. The final formulas for generating the cost value are shown in Equations (10) and (11).
[0031] C cg C(p, q)=hamming(TP(p), TP(q)) (10)
[0032] C cd C(p, q)=hamming(GT(p), GT(q)) (11)
[0033] C cg (p, q) represents the median-based grayscale cost value, C cd (p, q) represents the cost value of the gradient. The hamming function represents the number of different positions in the binary code, that is, the Hamming distance. The AD algorithm has a different measurement scale for the cost values of the above two algorithms, and the range of the cost values is also different. It is necessary to normalize the cost values to narrow the range to between [0, 1], and then fuse the cost values through different parameter coefficients. As shown in (12), the value is normalized, where λ is the control parameter and c is the cost value under different algorithms.
[0034]
[0035] Finally, the fusion method of the three cost calculation methods is shown in Equation (13), where λ cg represents the control parameter calculated based on the median, λ cd represents the control parameter calculated based on the gradient, λ AD represents the control parameter of the AD algorithm. A more stable cost calculation algorithm is obtained through cost fusion.
[0036] S23. After calculating the cost values of each pixel with other pixels, the pixels with similar colors around the pixel are found through cost aggregation. A Gaussian distribution is constructed through the cost value data in the cross domain. First, calculate the mean and variance among them. The calculation formulas are shown in (14) and (15), where μ represents the mean of the cost values in the cross domain, n nums represents the number of elements, N p represents the cross domain S. C(t) represents the cost value corresponding to each element in S, and σ represents the variance in the cross domain. Substitute the data in the cross domain into equations (14) and (15) to obtain the corresponding mean and method values.
[0037]
[0038] After obtaining the mean and variance, the corresponding Gaussian distribution is obtained. The existing noise points are selectively removed through the Gaussian distribution. The specific principles for removal are as follows:
[0039] 1) When the number of elements in the cross domain satisfies n num <α, no outlier removal is performed.
[0040] 2) When the number of elements in the cross domain satisfies β > n num ≥α, the cost values greater than μ + 3σ and less than μ - 3σ are regarded as the cost values of noise points, and these points are removed. The remaining ones are the correct cost value points.
[0041] 3) When the number of elements in the cross domain satisfies n num ≥β, the cost values greater than μ + σ and less than μ - σ are judged as noise points and removed.
[0042] Among them, α and β are control thresholds. By setting these two values, it is judged whether to remove the outlier cost values in the cross domain and which values to remove.
[0043] The third step is to calculate the final result. After removing the outliers in the cross domain, a new cross domain S is obtained, and the cost values in the remaining cross domain are all correct cost values. By calculating the corresponding average value of these remaining cost values, it is the cost value of the final pixel. The specific formula is shown in (16). C agg (P) represents the final cost value of point P after cost aggregation in the cross domain of the cross. n′ num represents the number of remaining elements after removing the noise in the proposed cross domain. N′ P represents the new cross domain S'.
[0044]
[0045] By introducing an optimization method of outlier removal in the process of cost aggregation in the cross domain of the cross, the influence of noise in the picture on the final stereo matching can be suppressed to a certain extent, and a more stable stereo matching algorithm can be obtained.
[0046] A method for three-dimensional reconstruction of the human body based on binocular stereo vision and point cloud. The purpose of step S3 is to improve the stereo matching speed and does not require setting the disparity range search. The hierarchical stereo matching strategy is not limited to the stereo matching of the current left and right image pairs, but through the way of image downsampling, multiple pairs of left and right image pairs of different scales are obtained. Generally, the Gaussian downsampling method is used for downsampling, and hierarchical left and right image pairs will be obtained. The upper-layer image is one-fourth smaller than the lower-layer image, and the length and width of the image are each reduced by half. Generally, three samplings are set, and four layers of Gaussian pyramid image pairs will be obtained. The bottom one is the original image pair, and the fourth layer is the left and right image pairs after three Gaussian downsamplings. The hierarchical stereo matching strategy is to first perform stereo matching on the left and right image pairs of the fourth layer, then infer the disparity search range of the left and right image pairs of the third layer through the stereo matching disparity map of the fourth layer, and then perform stereo matching on the left and right image pairs of the third layer until the stereo matching of the original-size left and right image pairs of the first layer is performed to obtain the final disparity map; through the coarse-to-fine stereo matching method, the stereo matching result of the coarse left and right image pairs can be applied to the relatively fine stereo matching, that is, using the disparity map of the small-scale image to infer the disparity search range of the large image. The present invention proposes a hierarchical stereo matching optimization strategy. This step specifically includes the following sub-steps:
[0047] S31. Build the Gaussian pyramids for the left and right images. The Gaussian pyramids are built from the original left and right image pair, that is, perform three Gaussian downsamplings. Each downsampling is to one - quarter of the original size, with the length and width each being half of the original length and width. The smallest - sized one is at the fourth layer, and the original image is at the first layer, as Figure 6 shown.
[0048] S32. Perform stereo matching on the left and right image pair at the fourth layer. For the stereo matching at the fourth layer, its algorithm is the same as the traditional stereo matching algorithm. First, set the disparity search range for the left and right image pair at the first layer. At this time, the disparity search range for each pixel in the picture is the same, and generally, the search range is set to [0, width * 20%], that is, the minimum disparity is 0, and the maximum disparity is 20% of the picture width. Finally, obtain the disparity map D 4 and the confidence map C 4 . The confidence map indicates the credibility of each disparity value in the disparity map. For those with high credibility, the disparity range of the next layer can be made smaller, otherwise, it can be made larger.
[0049] S33. Calculate the disparity range for each pixel in the left and right image pair I n at the n - th layer based on the disparity map and confidence map obtained from the previous layer, where n ≤ 3. For each pixel at the n - th layer, the disparity search range is determined by the disparity map D n+1 and the confidence map C n+1 from the previous layer. The specific determination method is shown in equations (17), (18), (19), and (20). The disparity search range at the pixel position (i, j) is [2d b -t, 2d b +t], where d b is the disparity value at the position (i / 2, j / 2) in the previous layer, and t takes 16 or 32, which is determined according to the credibility c b of the disparity value. If the credibility is greater than the threshold t c , then take 16, otherwise take 32. Calculate the disparity range for each pixel through equation (17), and then perform stereo matching on the left and right image pair. Finally, obtain the disparity map of this layer. Since the disparity range of each pixel is different, and the maximum range is 64 pixels and the minimum range is 32 pixels, it can shorten the time of stereo matching and reduce the memory occupancy.
[0050]
[0051] S34. Determine the disparity search range and perform stereo matching for each layer (except the fourth layer) in the same way as S33 until the last layer, and obtain the disparity map of the last layer, which is the final result. The confidence map in the hierarchical stereo matching strategy is calculated as shown in Equation (21). represents the minimum cost value corresponding to the disparity search range at the position (i, j) in the left image. represents the second smallest cost value. Through this equation, the confidence of the corresponding disparity value can be obtained, and the range of the confidence is [0, 1]. The greater the difference between the minimum cost value and the second smallest cost value, the more credible the corresponding disparity value is, and vice versa.
[0052]
[0053] Through the hierarchical stereo matching strategy, the memory occupancy of the cost space can be reduced, and the time complexity of stereo matching can be lowered. At the same time, the anti-noise performance of the algorithm can be improved through Gaussian filtering. Therefore, in the high-layer low-resolution left and right image pairs of the Gaussian pyramid, the accuracy of stereo matching will be improved due to the reduction of noise, and finally a more reliable stereo matching result can be obtained.
[0054] A method for 3D human body reconstruction based on binocular stereo vision and point cloud. The purpose of step S4 is to obtain the disparity map of the human body through the stereo matching algorithm based on median and gradient proposed in the present invention, and according to the internal and external parameter matrices of the camera, convert the disparity map into a 3D point cloud map, and finally realize the surface reconstruction of the human body through the method of triangulation. This step specifically includes the following sub-steps:
[0055] S41. Generate a disparity map from the human body images collected by the binocular camera through the stereo matching algorithm, as Figure 7 shown. After obtaining the disparity map, it is converted into a depth map, and the formula is as follows
[0056]
[0057] f is the focal length of the camera. baseline is the distance between the two cameras, also known as the baseline. U_L is the horizontal coordinate of the point on the left camera image plane. U_R is the horizontal coordinate of the point on the right camera image plane. The disparity is U L -(U R ), that is, the difference between the two coordinates. Specific points in the 3D space are obtained according to the depth map to generate the 3D point cloud of the human body, and finally the greedy projection triangulation algorithm is used to perform surface reconstruction on the 3D human body point cloud data.
[0058] So far, we have realized a method for 3D human body reconstruction based on binocular stereo vision and point cloud.
[0059] Compared with the prior art, the main improvements of the present invention are as follows:
[0060] 1. The present invention proposes cost calculation based on median and gradient. Our method reduces calculation errors caused by unstable factors of central pixels. By introducing gradient as one of the bases for measuring the cost value, it can make the method applicable to pictures with uneven illumination to a certain extent and improve the matching accuracy rate.
[0061] 2. The present invention optimizes outlier rejection. When the cost value of a noise point has a large difference from the cost value in the region, it will cause a large error in the result. Through the method of cost aggregation and outlier rejection optimization, the abnormal cost values in the cross domain are removed, and finally a more accurate average cost value can be obtained.
[0062] 3. The present invention is based on a hierarchical stereo matching optimization strategy, constructs a disparity map from coarse to fine at different resolutions, realizes better reconstruction results, reduces resource waste and improves the stereo matching speed at the same time. Description of the Drawings
[0063] Figure 1 is the experimental equipment for camera calibration of the present invention.
[0064] Figure 2 is the corner display diagram for calculating the disparity of the present invention.
[0065] Figure 3 is the corner mapping relationship diagram for calculating the disparity of the present invention.
[0066] Figure 4 are the convolution factors in the x and y directions of the Sobel operator.
[0067] Figure 5 is the cross domain of point P in the outlier rejection of the present invention.
[0068] Figure 6 is the pyramid image pair diagram based on the hierarchical stereo matching strategy of the present invention.
[0069] Figure 7 is the human body disparity map calculated based on the present invention.
[0070] Figure 8 is the three-dimensional reconstruction flow chart based on binocular vision of the present invention.
[0071] Term Explanation
[0072] SGM is the abbreviation of Semi - Global Matching, which is a semi - global stereo matching algorithm. Its implementation method refers to the reference [H. Hirschmuller, "Stereo Processing by Semiglobal Matching and Mutual Information," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 328 - 341, Feb. 2008, doi: 10.1109 / TPAMI.2007.1166.]
[0073] PatchMatch is a local stereo matching algorithm that can estimate the disparity at each pixel and its 3D disparity plane, and is more suitable for disparity estimation on inclined surfaces. Its implementation method refers to the reference [Bleyer, Michael et al. “PatchMatch Stereo - Stereo Matching with Slanted Support Windows.” British Machine Vision Conference (2011).]
[0074] AD - Census is a stereo matching algorithm that combines local and global algorithms. Its implementation method refers to the reference [Mei X, Sun X, Zhou M, et al. On building an accurate stereo matching system on graphics hardware[C] / / IEEE International Conference on Computer Vision Workshops. IEEE, 2012.] Specific implementation manners
[0075] The present invention will be further described in conjunction with the accompanying drawings.
[0076] Embodiment
[0077] The complete process of the 3D human body reconstruction method based on binocular stereo vision and point cloud is as shown in the attached Figure 8 figure.
[0078] The first step is to calibrate two cameras through step S1. Since in the process of converting a two-dimensional image to a three-dimensional object, the internal and external parameters of the cameras are required to calculate the depth information, and in order to satisfy the epipolar constraint under parallel views during stereo matching in steps S2 and S3, the stereo rectification projection matrices corresponding to the two cameras need to be obtained. Then, the parameters obtained in camera calibration are applied to stereo rectification, stereo matching, and three-dimensional reconstruction of the images. Therefore, next, the internal and external parameter matrices of the corresponding cameras and the stereo rectification projection matrices are obtained through the camera calibration experiment, and then these matrices are applied to the processes of stereo rectification and generating the three-dimensional point cloud of the human body in step S4 during the human body three-dimensional reconstruction experiment. Finally, the entire process of reconstructing the three-dimensional surface of the human body is realized.
[0079] Application Example
[0080] The method for three-dimensional reconstruction of the human body based on binocular stereo vision provided by the embodiment is used to experiment on different images in the MiddleBurry dataset. The source of the MiddleBurry dataset is shown in the reference [S. Baker, D. Scharstein, J. Lewis, S. Roth, M. J. Black, and R. Szeliski. A database and evaluation methodology for optical flow. IJCV, 2011.2, 3, 5, 6, 7]. In order to verify that the algorithm proposed in this paper can improve the correct rate of stereo matching to a certain extent when there is light interference, three groups of images of the Mobius pictures in the MiddleBurry dataset under different lighting conditions and different noise conditions are selected for experiments. The experimental results are shown in Tables 1 and 2.
[0081] Table 1 Matching Error Rate Table under Different Lighting Conditions (%)
[0082]
[0083] Table 2 Matching Error Rate Table under Different Noise Conditions (%)
[0084]
[0085]
[0086] As can be seen from the above table, the feasibility of the proposed improved method is verified through experiments. When performing stereo matching on images with noise interference and uneven lighting, the correct rate is increased by 5% to 8%.
[0087] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for 3D reconstruction of a human body based on binocular stereo vision and point cloud, which specifically comprises the following steps: S1, using a binocular camera to take pictures of the chessboard calibration plate at different viewing angles to obtain the camera's internal and external parameter matrix and the stereo correction projection matrix, and perform distortion correction on the input left and right human body images according to the matrix parameters. S2, stereo matching is performed on the input image that has completed distortion correction. Our invention is based on the cost calculation method of median and gradient, and proposes a cost aggregation outlier removal optimization method. S3, through the hierarchical stereo matching strategy, constructs the disparity map from coarse to fine method to optimize the operation speed of stereo matching. S4, obtains the three-dimensional point cloud data of the human body according to the disparity map obtained by stereo matching, and reconstructs the surface of the three-dimensional point cloud data of the human body through a greedy projection triangulation algorithm.
2. The method for three-dimensional reconstruction of a human body based on binocular stereo vision and point cloud according to claim 1, characterized in that: In step S1, for the binocular camera used for 3D reconstruction, first use the chessboard calibration method to obtain its internal and external parameter matrix. For the input left and right human body images, use the internal and external parameter matrix to correct the image distortion caused by the camera. Calculate the internal and external parameter matrix of the camera and the stereo correction projection matrix. First, obtain the internal and external parameter matrix of the camera by calibrating the monocular camera separately. A represents the internal parameter matrix, N represents the external parameter matrix, (X W ,Y W ,Z W ) represents the coordinate value in the world coordinate system, (u,v) represents the coordinate value in the image pixel coordinate system, and Z c Represents the value of the Z axis in the camera coordinate system, f represents the focal length of the camera, 1 / d x , 1 / d y They respectively represent the contraction ratio of the image physical coordinate system to the U axis and V axis of the image pixel coordinate system, (u0, v0) is the moving distance, and θ represents the angle of the final pixel coordinate system that deviates from the normal situation. We obtained the internal and external parameter matrices of the camera and the positional relationship of the binocular camera. Then, the human body image captured by the binocular camera is preprocessed to complete the distortion correction. First, we Gaussian blur the collected input image to remove noise points, and then use Unsharpen Mask to sharpen and enhance it. Through preprocessing, we get two pictures with slightly less noise and better quality. Then restore the image distortion through distortion correction and epipolar correction.
3. The method for three-dimensional reconstruction of a human body based on binocular stereo vision and point cloud according to claim 1, characterized in that: In step S2, after the left and right images captured by the binocular camera are processed in step S1, stereo matching is performed. The same-name points of the pixels in the left image are found in the right image, that is, all points in the same row in the right image are searched, and the similarity is compared. The points with the highest similarity are determined as the same-name points. The disparity is calculated by determining the same-name points, so the disparity map can be finally obtained through stereo matching. This step includes the following sub-steps: S21, measure the similarity of the pixel values of the left and right images through cost calculation. Inspired by the AD-Census algorithm, our invention combines the advantages of the AD (Absolute Differences) algorithm and the Census algorithm. The cost value under the AD algorithm is obtained by comparing the size of a single pixel. Then our invention improves the cost calculation method under the Census algorithm, based on the window median pixel as the core of the calculation and introducing gradient calculation as one of the bases for measuring the cost value. First is the first part of the cost calculation under the AD algorithm: Among them C AD (p,q) is the AD algorithm cost value of point p in the left figure and point q in the right figure. represents the gray value of channel i in the left image. Indicates the gray value of q in the i channel in the right figure, a total of three channels, and finally calculates the average value of the absolute value of the difference between the three channels p and q. The present invention is based on the cost calculation of the second part of the binary code calculation, and generates the corresponding binary code by taking the median value of the pixels in the window as the comparison value. The calculation formula is as follows: Formula (6) represents the cost calculation method of point p in the image based on the median binary code, t∈N p It means to calculate the gray value in the window of point P, and It is a bit connector, which is to calculate the corresponding bit value of each pixel in the window, and then connect it into a binary code through this symbol. The corresponding bit value of each pixel in the window is calculated through the function ζ, where I mid (p) represents the median grayscale value of all pixels in the window centered on p. I(t) represents the grayscale value of each pixel in the window. By comparing the two, we can get the bit value corresponding to the pixel in the window. By connecting all the bit values, we can get the binary code of the window corresponding to point p. S22, introduce gradient for cost calculation. By using the Sobel operator for gradient calculation, the Sobel convolution factor is divided into the convolution factor in the x direction and the y direction, corresponding to the gradients in the two directions in Figure 4. Finally, the gradients in the two directions are fused to obtain the final gradient value. The specific calculation of the gradients in the x direction and the y direction is shown in formula (8). The gradient in the x direction at (u, v) is G x (u,v),G y (u, v) represents the gradient in the y direction. The gradient values in these two directions are averaged to obtain the final gradient value, as shown in Figure (9). The cost value based on the AD algorithm, the cost value based on the median binary code, and the cost value based on the gradient calculation are obtained in the above manner. The three cost values are combined to obtain the final cost value. The cost value is the Hamming distance of the binary code. The formula for generating the cost value is shown in equations (10) and (11). C cg (p,q)=hamming(TP(p),TP(q)) (10) C cd (p,q)=hamming(GT(p),GT(q)) (11) C cg (p,q) represents the grayscale cost value based on the median, C cd (p, q) represents the cost value of the gradient. The hamming function represents the number of different corresponding positions of the binary code, that is, the Hamming distance. The AD algorithm has a different cost value measurement scale from the above two algorithms, and the cost value range is also different. The cost value needs to be normalized, the range is narrowed to [0, 1], and then the cost value is fused through different parameter coefficients. As shown in (12), the value is normalized, where λ is the control parameter and c is the cost value under different algorithms. The final fusion method of the three cost calculation methods is shown in formula (13), where λ cg represents the control parameter based on median calculation, λ cd represents the control parameter based on gradient calculation, λ AD It represents the control parameters of the AD algorithm. A more stable cost calculation algorithm is obtained through cost fusion. S23, after calculating the cost value of each pixel and other pixels, find the pixels around the pixel and other pixels with similar colors through cost aggregation. Construct a Gaussian distribution through the cost value data in the cross domain, first calculate the mean value and variance, and the calculation formula is shown in (14) and (15), where μ represents the mean value of the cost value in the cross domain, n nums Represents the number of elements, N p It represents the cross domain S. C(t) represents the cost value corresponding to each element in S, and σ represents the variance in the cross domain. Substituting the data in the cross domain into equations (14) and (15), we can get the corresponding average value and method value. After calculating the mean value and variance, the corresponding Gaussian distribution is obtained. The existing noise points are selectively removed through the Gaussian distribution. The specific principles of removal are as follows: 1) When the number of elements in the intersection domain satisfies n num <α, no outlier removal is performed. 2) When the number of elements in the intersection domain satisfies β>n num ≥α, the cost values greater than μ+3σ and less than μ-3σ are regarded as the cost values of noise points, and these points are eliminated, and the remaining ones are the correct cost points. 3) When the number of elements in the intersection domain satisfies n num ≥β, which judges the cost values greater than μ+σ and less than μ-σ as noise points and removes them. Among them, α and β are control thresholds, and these two values are set to determine whether to remove abnormal values and those values in the cross domain. The third step is to calculate the final result. After removing the outliers in the cross domain, a new cross domain S is obtained, and the cost values in the remaining cross domain are all correct cost values. The corresponding average value is calculated by the remaining cost values, which is the cost value of the final pixel. The specific formula is shown in (16), C agg (P) represents the final cost value of point P after the cross-domain cost aggregation, n' num represents the number of elements remaining after removing the noise in the cross domain, N' P This represents the new cross-domain S'. By introducing an optimization method for outlier removal in the process of cost aggregation in the cross-domain, the influence of noise in the image on the final stereo matching can be suppressed to a certain extent, and a more stable stereo matching algorithm can be obtained.
4. The method for three-dimensional reconstruction of a human body based on binocular stereo vision and point cloud according to claim 1, characterized in that: In step S3, a hierarchical stereo matching optimization strategy is proposed. This step specifically includes the following sub-steps: S31, establish a Gaussian pyramid for the left and right image pairs. The Gaussian pyramid established by the original left and right image pairs is Gaussian downsampling performed three times, each downsampling is one-fourth of the original size, the length and width are half of the original length and width, the smallest size is at the fourth layer, and the original image is at the first layer, as shown in FIG6 . S32, stereo matching is performed on the fourth layer left and right image pairs. For the fourth layer stereo matching, its algorithm is the same as the traditional stereo matching algorithm. First, the disparity search range of the first layer left and right image pairs must be set. At this time, the disparity search range for each pixel in the image is the same, and the search range is generally set to [0, width*20%], that is, the minimum disparity value is 0 and the maximum value is 20% of the image width. Finally, the disparity map D is obtained. 4 and confidence map C 4 The confidence map shows the credibility of each disparity value in the disparity map. A high confidence value can make the disparity range of the next layer smaller, otherwise it can be larger. S33, calculate the n-th layer left and right image pair I through the disparity map and confidence map obtained in the previous layer n The disparity search range of each pixel in the nth layer is It is composed of the disparity map D of its previous layer. n+1 and confidence map C n+1 The specific determination method is shown in equations (17), (18), (19) and (20). The disparity search range at the pixel (i, j) position is [2d b -t,2d b +t],d b It is the disparity value of the previous layer at position (i / 2, j / 2), and t is 16 or 32, which is based on the credibility of the disparity value c b To determine, if the credibility is greater than the threshold t c , then take 16, otherwise take 32. The disparity range of each pixel is calculated by formula (17), and then the left and right image pairs are stereo matched to finally obtain the disparity map of this layer. Since the disparity range of each pixel is different, and the maximum range is 64 pixels and the minimum range is 32 pixels, it can shorten the stereo matching time and reduce the memory usage. S34, determine the disparity search range and perform stereo matching for each layer (except the fourth layer) by the method of S33 until the last layer, and obtain the disparity map of the last layer, which is the final result. The confidence map in the layered stereo matching strategy is calculated as shown in formula (21): It represents the minimum cost value corresponding to the disparity search range at position (i, j) in the left figure. It represents the second smallest cost value. The confidence of the corresponding disparity value can be obtained through this formula, and the range of confidence is [0,1]. The greater the difference between the smallest cost value and the second smallest cost value, the more credible the corresponding disparity value is, and vice versa. The hierarchical stereo matching strategy can reduce the memory usage of the cost space and the time complexity of stereo matching. At the same time, the noise resistance of the algorithm can be improved by Gaussian filtering. Therefore, in the high-level low-resolution left and right image pairs of the Gaussian pyramid, the accuracy of stereo matching will be improved due to the reduction of noise, and finally a more reliable stereo matching result will be obtained.
5. The method for three-dimensional reconstruction of a human body based on binocular stereo vision and point cloud according to claim 1, characterized in that: In step S4, the median and gradient-based stereo matching algorithm proposed in the present invention is used, and the disparity map of the human body is obtained according to the layered stereo matching optimization strategy, and then the disparity map is converted into a three-dimensional point cloud map according to the internal and external parameter matrices of the camera, and finally the human body surface is reconstructed by the triangulation method. Through the stereo matching algorithm, the human body image collected by the binocular camera is used to generate a disparity map, as shown in Figure 7. After the disparity map is obtained, it is converted into a depth map. The formula is as follows: f is the focal length of the camera. baseline is the distance between the two cameras, also called the baseline. U_L is the horizontal coordinate of the point on the left camera image plane. U_R is the horizontal coordinate of the point on the right camera image plane. Disparity is U L -(U R ), which is the difference between the two coordinates. According to the depth map, the specific points in the three-dimensional space are obtained to generate the three-dimensional point cloud of the human body. Finally, the greedy projection triangulation algorithm is used to reconstruct the surface of the three-dimensional point cloud data of the human body. So far, we have realized a method for 3D reconstruction of the human body based on binocular stereo vision and point cloud.