Fast 3d face reconstruction method and system based on left and right face depth cameras
By simultaneously acquiring data from left and right face depth cameras and using a precise registration algorithm, combined with texture stitching processing, the speed, accuracy, and texture quality issues in existing 3D face reconstruction methods are resolved, achieving fast and high-quality 3D face reconstruction suitable for both consumer and industrial applications.
Patent Information
- Application Number
- CN202511487668.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-10-17
Smart Images

Figure CN120976441B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional face reconstruction, and in particular to a fast three-dimensional face reconstruction method and system based on left and right face depth cameras. Background Technology
[0002] A person's face contains a lot of important information. With the development of computer vision and depth camera technology, 3D face reconstruction technology is also constantly being updated and has important application prospects in fields such as medical aesthetics, public safety and virtual reality.
[0003] Traditional 3D face reconstruction methods typically rely on the analysis of face parameters in a single image or the stereo matching of face feature points in multi-view images. However, these methods have certain limitations in terms of reconstruction speed, reconstruction accuracy, 3D information integrity, and equipment structure feasibility, making them difficult to apply widely in ordinary environments.
[0004] With the advent of structured light depth cameras, the technology for obtaining 3D facial data using active 3D measurement has become relatively mature, capable of real-time output of high-resolution, high-precision 3D facial data. However, in applications involving complete 3D facial reconstruction, due to occlusion, single-view depth cameras often fail to capture 3D data of the preauricular region, postauricular region, or even the entire ear area. Existing 3D facial scanning devices based on handheld or mechanically controlled scanning, while capable of acquiring 3D data from multiple perspectives in space, suffer from excessively long data acquisition times. Facial movement and changes in facial expressions can cause deviations in the facial data acquired at different times, ultimately affecting the quality of the reconstructed 3D facial model. Besides accurate 3D facial data, rich facial color texture is also crucial for high-quality 3D facial models. Similar to 3D facial data, facial texture data acquired by RGB cameras also encounters the above issues, ultimately affecting the detail quality of the color texture and the accuracy of texture mapping in the resulting 3D facial model.
[0005] Currently, there is no fast 3D face reconstruction method with complete model texture details, which limits 3D face reconstruction. To address this, a fast 3D face reconstruction method based on left and right face depth cameras is proposed to achieve rapid acquisition of face data and real-time reconstruction of textured and detailed 3D face models. Summary of the Invention
[0006] In view of the above, the main objective of this invention is to propose a fast 3D face reconstruction method and system based on left and right face depth cameras to solve the above-mentioned technical problems.
[0007] This invention proposes a fast 3D face reconstruction method based on left and right face depth cameras, the method comprising the following steps:
[0008] Step 1: Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face;
[0009] Step 2: Based on the face feature point detection algorithm, identify the face regions in the left and right grayscale images respectively and obtain the face feature point region ROI; according to the face feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, segment the left face point cloud and the right face point cloud from the 3D point cloud data, and perform filtering and downsampling processing on the segmented left face point cloud and right face point cloud.
[0010] Step 3: Calculate the local feature descriptors of the left and right face point clouds, and based on the local feature descriptors, perform coarse registration of the left and right face point clouds using the sample consistency initial registration algorithm, and then perform fine registration using the iterative nearest point algorithm to obtain the fine registration transformation matrix; use the fine registration transformation matrix to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain the complete face point cloud; perform normal calculation and Poisson surface reconstruction on the complete face point cloud to generate the initial 3D face mesh model;
[0011] Step 4: Locate the corresponding texture regions in the left and right RGB images based on the ROI of the facial feature points, then perform color difference processing and brightness equalization to obtain the processed left and right texture regions; stitch the processed left and right texture regions together according to a preset ratio to generate a complete texture map.
[0012] Step 5: Using the nasal tip midline as the dividing line, divide the vertices in the initial 3D face mesh model into left-side vertices and right-side vertices; for the left-side edge vertices located in the stitching gap area, adjust their coordinates to the coordinates of the nearest right-side edge vertex to eliminate the gap; based on the intrinsic and extrinsic parameters of the left and right RGB cameras, calculate the texture coordinates of each vertex in the initial 3D face mesh model in the texture map; finally, output a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information.
[0013] This invention also proposes a fast 3D face reconstruction system based on left and right face depth cameras, wherein the system applies the fast 3D face reconstruction method based on left and right face depth cameras as described above, and the system includes:
[0014] The data acquisition module is used for:
[0015] Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face;
[0016] The point cloud segmentation module is used for:
[0017] Based on the facial feature point detection algorithm, the facial regions in the left and right grayscale images are identified respectively to obtain the facial feature point region ROI; according to the facial feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, the left face point cloud and the right face point cloud are segmented from the 3D point cloud data, and the segmented left face point cloud and right face point cloud are filtered and downsampled.
[0018] The mesh reconstruction module is used for:
[0019] Local feature descriptors of left and right face point clouds are calculated. Based on these descriptors, a sample consistency initial registration algorithm is used to coarsely register the left and right face point clouds, followed by a fine registration algorithm based on iterative nearest point to obtain a fine registration transformation matrix. The fine registration transformation matrix is then used to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain a complete face point cloud. Normals are calculated and Poisson surfaces are reconstructed from the complete face point cloud to generate an initial 3D face mesh model.
[0020] The texture generation module is used for:
[0021] Based on the ROI of the facial feature points, the corresponding texture regions in the left and right RGB images are located, and then color difference processing and brightness equalization are performed to obtain the processed left and right texture regions. The processed left and right texture regions are then stitched together according to a preset ratio to generate a complete texture map.
[0022] The model generation module is used for:
[0023] Using the nasal tip midline as the dividing line, the vertices in the initial 3D face mesh model are divided into left-side vertices and right-side vertices. For the left-side edge vertices located in the stitching gap region, their coordinates are adjusted to the coordinates of the nearest right-side edge vertex to eliminate the gap. Based on the intrinsic and extrinsic parameters of the left and right RGB cameras, the texture coordinates of each vertex in the initial 3D face mesh model in the texture map are calculated. Finally, a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information is output.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] 1. This invention uses simultaneous triggering of left and right depth cameras to ensure that the facial state is captured at the same instant, which can effectively avoid data misalignment caused by changes in facial micro-expressions and lay the foundation for accurate registration in the future.
[0026] 2. This invention employs a two-stage strategy of coarse registration followed by fine registration. SAC-IA quickly estimates a rough transformation matrix based on local feature descriptors, and ICP iteratively optimizes this matrix to minimize the distance between two point clouds, thereby obtaining a high-precision rigid transformation matrix and achieving accurate alignment of the point clouds.
[0027] 3. This invention proposes a partitioned projection and edge blending method. The model is divided into left and right sides, with the central plane of the nose tip as the boundary. The vertex textures on the left side come from the left depth camera, and the vertex textures on the right side come from the right depth camera. By "stitching" the vertex coordinates at the seams, the texture stitching gaps are eliminated, resulting in a natural and continuous texture mapping.
[0028] 4. This invention eliminates undesirable textured patches by directly modifying the vertex coordinates of the mesh model at the geometric level, pulling the vertices on one side of the seam to the other side, thus fundamentally solving the problem of undesirable textures being mapped onto the patches at the junction.
[0029] 5. The 3D face model generated by this invention possesses both high geometric accuracy and high visual fidelity. The model not only includes a complete facial structure (including often-obscured ears and profile), but also retains rich surface details, with seamless textures and realistic colors. Through an innovative dual-camera architecture and efficient algorithm flow, the time from data acquisition to model generation is significantly shortened, while hardware and computational costs are reduced. This enables high-quality 3D face reconstruction technology to move from expensive professional equipment to a wider range of consumer and industrial applications.
[0030] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description
[0031] Figure 1 This is a flowchart of the fast 3D face reconstruction method based on left and right face depth cameras proposed in this invention;
[0032] Figure 2 This is a schematic diagram of the structure of the fast 3D face reconstruction system based on left and right face depth cameras proposed in this invention;
[0033] Figure 3 This is a schematic diagram of data acquisition in this invention;
[0034] Figure 4 This is a flowchart of the left and right face point cloud fusion process of the present invention;
[0035] Figure 5 This is a flowchart of the left and right face texture fusion process of the present invention;
[0036] Figure 6 This is a flowchart illustrating the process of eliminating texture stitching gaps in the human face model according to the present invention.
[0037] Figure 7 This is a schematic diagram illustrating the texture coordinate calculation of the present invention.
[0038] In the diagram: 1 represents the left depth camera; 2 represents the right depth camera; 3 represents the casing; 4 represents the touchscreen; 5 represents the user; rgbx, rgby represent the pixel coordinate system in the left or right RGB image; 1_rgbx, 1_rgby represent the pixel coordinates of the left vertex in the left RGB image; r_rgbx, r_rgby represent the pixel coordinates of the right vertex in the right RGB image; 1x, 1y represent the pixel coordinates of the top-left corner vertex of the left texture area (rectangle) in the left RGB image; rx, ry represent the right texture... The top-left vertex of the region (rectangle) corresponds to the pixel coordinates in the right RGB image; 1_px, 1_py represent the pixel coordinates in the final texture map corresponding to the left vertex; r_px, r_py represent the pixel coordinates in the final texture map corresponding to the right vertex; px, py represent the pixel coordinate system in the final texture map; vtx represents the normalized value of px, and vty represents the normalized value of py; vtx, vty represent the relative coordinates of a pixel in the final texture map, i.e., the texture coordinates finally written to the OBJ file. Detailed Implementation
[0039] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0040] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0041] Please see Figure 1 This embodiment provides a fast 3D face reconstruction method based on left and right face depth cameras, the method including the following steps:
[0042] Step 1: Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face;
[0043] Step 2: Based on the face feature point detection algorithm, identify the face regions in the left and right grayscale images respectively and obtain the face feature point region ROI; according to the face feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, segment the left face point cloud and the right face point cloud from the 3D point cloud data, and perform filtering and downsampling processing on the segmented left face point cloud and right face point cloud.
[0044] Step 3: Calculate the local feature descriptors of the left and right face point clouds, and based on the local feature descriptors, perform coarse registration of the left and right face point clouds using the sample consistency initial registration algorithm, and then perform fine registration using the iterative nearest point algorithm to obtain the fine registration transformation matrix; use the fine registration transformation matrix to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain the complete face point cloud; perform normal calculation and Poisson surface reconstruction on the complete face point cloud to generate the initial 3D face mesh model;
[0045] Step 4: Locate the corresponding texture regions in the left and right RGB images based on the ROI of the facial feature points, then perform color difference processing and brightness equalization to obtain the processed left and right texture regions; stitch the processed left and right texture regions together according to a preset ratio to generate a complete texture map.
[0046] Step 5: Using the nasal tip midline as the dividing line, divide the vertices in the initial 3D face mesh model into left-side vertices and right-side vertices; for the left-side edge vertices located in the stitching gap area, adjust their coordinates to the coordinates of the nearest right-side edge vertex to eliminate the gap; based on the intrinsic and extrinsic parameters of the left and right RGB cameras, calculate the texture coordinates of each vertex in the initial 3D face mesh model in the texture map; finally, output a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information.
[0047] In a preferred embodiment of the present invention, in step 1, the depth camera is a structured light camera or a binocular vision depth camera, and each binocular vision depth camera integrates an RGB camera. While outputting grayscale images, RGB images and 3D point clouds, it also outputs the intrinsic and extrinsic parameter matrices of the RGB cameras on the left and right depth cameras.
[0048] In a preferred embodiment of the present invention, in step 2, the face feature point detection algorithm adopts the 68-point or 76-point detection model of the dlib library; the filtering process includes pass-through filtering and statistical filtering, which are used to remove outliers and noise.
[0049] In a preferred embodiment of the present invention, in step 3, the sampling consensus initial registration algorithm is SAC-IA (Sampling Consensus Initial Alignment), and the iterative closest point algorithm is ICP (Iterative ClosestPoint).
[0050] In a preferred embodiment of the present invention, step 3, after the Poisson surface is reconstructed, further includes a trimming step, which is detailed below:
[0051] Based on the spatial density distribution of the complete face point cloud, calculate the spatial occupancy grid;
[0052] Using the central plane of the nose tip as the boundary, calculate the density of the point cloud in the spatial occupancy grid on the left and right sides respectively;
[0053] The average density of the point cloud on the opposite side is calculated based on the density of the point cloud on the left and right sides in the spatial occupancy grid. The dynamic density threshold of the current side is set according to the average density of the point cloud on the opposite side, so as to identify and mark the sparse grids in the current side reconstructed grid whose density is lower than the dynamic density threshold.
[0054] The triangular patches corresponding to the sparse grid are identified as useless patches caused by point cloud registration errors or noise. The surface clipping tool is used to clip and remove them from the initial 3D face mesh model, thereby obtaining an optimized initial 3D face mesh model with fewer holes and a smoother surface, which is used for subsequent texture mapping.
[0055] This embodiment, based on the roughly left-right symmetry of a human face, sets a dynamic density threshold for the current side that is correlated with the average density of the point cloud on the opposite side. This allows the cropping criteria on both sides to be adaptive: strictly filtering the poor-quality side using the standard of the high-quality side to remove invalid data, and loosely filtering the high-quality side using the standard of the poor-quality side to retain high-quality data. This better adapts to differences in point cloud density among different faces and under different acquisition conditions, avoiding accidental deletion of valid or useless faces. Cropping is more intelligent and precise, automatically removing low-confidence mesh portions, significantly improving the accuracy and aesthetics of the final 3D face model, and reducing the workload of manual post-processing. Furthermore, by analyzing the "complete face point cloud" after registration and fusion to guide the cropping of the mesh model, it precisely solves the problem of redundant "beard" or thin-film-like faces in the reconstructed mesh caused by inconsistent point cloud density in overlapping areas or sparse point clouds in edge areas after left-right point cloud registration.
[0056] In a preferred embodiment of the present invention, in step 4, the color difference processing includes histogram matching and bilateral filtering, which are used to keep the color and brightness of the RGB images on the left and right sides consistent. In step 4, the preset ratio is 50%, that is, when splicing, the left half of the pixels of the left texture area is retained and the right half of the pixels of the right texture area is retained. Then, these two parts of pixels are spliced to generate a complete texture map.
[0057] In a preferred embodiment of the present invention, step 5, adjusting the coordinates of the left face edge vertex located in the splicing gap area to the coordinates of the nearest right face edge vertex, specifically involves: constructing a KD tree for the right face edge vertex; for the left face edge vertex, searching for its nearest right face edge vertex through the KD tree, and modifying the coordinates of the left face edge vertex to the coordinates of the nearest right face edge vertex.
[0058] In a preferred embodiment of the present invention, step 5, the process of calculating the texture coordinates of each vertex in the texture map of the initial three-dimensional face mesh model, is performed in the following manner depending on the face side to which the vertex belongs:
[0059] For vertices classified as left face side, the inverse of the fine registration transformation matrix is first used to transform the vertices of the left face side from the right depth camera coordinate system back to the left depth camera coordinate system, obtaining the 3D coordinates of the vertices in the left depth camera coordinate system. Then, based on the extrinsic parameter matrix of the left RGB camera relative to the left depth camera, the 3D coordinates of the vertices in the left depth camera coordinate system are transformed to the left RGB camera coordinate system, obtaining the 3D coordinates of the vertices in the left RGB camera coordinate system. Finally, using the intrinsic parameter matrix of the left RGB camera and the pinhole imaging model, the 3D coordinates of the vertices in the left RGB camera coordinate system are projected onto the left RGB image to obtain 2D pixel coordinates. Based on the position of the left texture region ROI in the complete texture map, the 2D pixel coordinates are mapped onto the complete texture map to obtain the texture coordinates of the vertices of the left face side in the complete texture map.
[0060] For vertices classified as the right side of the face, the vertices on the right side are directly transformed from the right depth camera coordinate system to the right RGB camera coordinate system based on the extrinsic parameter matrix of the right RGB camera relative to the right depth camera, obtaining the 3D coordinates of the vertices in the right RGB camera coordinate system. Then, using the intrinsic parameter matrix of the right RGB camera and the pinhole imaging model, the 3D coordinates of the vertices in the right RGB camera coordinate system are projected onto the right RGB image to obtain the 2D pixel coordinates. Based on the position of the right texture region ROI in the complete texture map, the 2D pixel coordinates are mapped onto the complete texture map to obtain the texture coordinates of the vertices on the right side of the face in the complete texture map.
[0061] The formula for calculating the conversion from two-dimensional pixel coordinates to texture coordinates is as follows:
[0062] ;
[0063] Where rgbx and rgby represent the pixel coordinates in the left or right RGB image; height and width represent the height and width of the left and right texture areas (rectangles) after unifying their size; rate represents the preset ratio; after stitching the left and right textures according to the preset ratio, the final texture map's height remains height, but the width becomes width*rate*2; 1_rgbx and 1_rgby represent the pixel coordinates of the left vertex corresponding to the pixel in the left RGB image; r_rgbx and r_rgby represent the pixel coordinates of the right vertex corresponding to the pixel in the right RGB image; 1x and 1y represent the pixel coordinates of the top left vertex of the left texture area (rectangle). The pixel coordinates in the left RGB image; rx,ry represents the pixel coordinates of the top-left vertex of the right texture area (rectangle) in the right RGB image; 1_px,1_py represents the pixel coordinates of the left vertex in the final texture map; r_px,r_py represents the pixel coordinates of the right vertex in the final texture map; px,py represents the pixel coordinate system in the final texture map; vtx represents the normalized value of px, and vty represents the normalized value of py; vtx,vty are the relative coordinates of a pixel in the final texture map, which are the texture coordinates finally written to the OBJ file.
[0064] To further improve the accuracy and naturalness of texture mapping, this embodiment will also determine the expansion width threshold by obtaining the registration error statistics of the final iteration of fine registration after obtaining the texture coordinates of the vertices on the left side of the face and the right side of the face in the complete texture map. Taking the central plane of the nose tip as the reference, the width threshold is expanded to both sides to form a three-dimensional gap processing area; the mesh vertices located in this area are marked as vertices to be processed.
[0065] For each vertex to be processed, perform the following operations:
[0066] a. Based on the intrinsic and extrinsic parameters of the left and right RGB cameras, calculate the first coordinate vector of the vertex in the coordinate system of the left RGB camera and the second coordinate vector in the coordinate system of the right RGB camera, respectively;
[0067] b. Calculate the first cosine value of the angle between the first coordinate vector and the normal vector of the vertex to be processed, Cosθ_L;
[0068] c. Calculate the second cosine value of the angle between the second coordinate vector and the normal vector of the vertex to be processed, Cosθ_R;
[0069] d. Compare the absolute values of the cosine values of the first and second included angles;
[0070] Based on the comparison between the cosine of the first included angle and the cosine of the second included angle, the final texture coordinates of the vertex to be processed are re-determined: if the absolute value of the second included angle cosine is greater than the absolute value of the first included angle cosine, then the right RGB camera is determined to be the optimal texture source for the vertex, and the texture coordinates calculated based on the right RGB image are used as the final texture coordinates of the vertex to be processed.
[0071] If the absolute value of the cosine of the first included angle is greater than the absolute value of the cosine of the second included angle, then the left RGB camera is determined to be the optimal texture source for the vertex to be processed, and the texture coordinates calculated based on the left RGB image are used as the final texture coordinates of the vertex to be processed.
[0072] This invention utilizes visual geometry principles to dynamically and intelligently select the optimal texture source for each vertex of the stitching region. This fundamentally avoids stretching, distortion, and blurring issues that may result from using textures from images viewed from a non-direct perspective, significantly improving the fidelity and visual realism of the final model texture. Furthermore, by automatically assigning appropriate texture sources to the vertices in the midline region, the transition between left and right textures becomes more natural and smooth, eliminating visible stitching gaps and color jumps, thus obtaining a visually highly unified and complete face texture map.
[0073] In a preferred embodiment of the present invention, in step 5, the output three-dimensional model file is in OBJ format, and is accompanied by MTL material library file and PNG texture map file.
[0074] In the above scheme, snapshot depth cameras with RGB cameras are used to collect RGB and 3D data of the face. Data collection of the entire face, including the ears, is performed using a minimum number of two depth cameras. After the two snapshot depth cameras, deployed on the left and right sides, have collected the face data, the 3D reconstruction task is completed by a computer C++ program. This includes left and right point cloud registration, point cloud surface reconstruction, texture mapping, and surface mesh model texture mapping, ultimately resulting in a complete and detailed 3D face model with texture.
[0075] Please refer to Figure 2 This embodiment also provides a fast 3D face reconstruction system based on left and right face depth cameras, wherein the system applies the fast 3D face reconstruction method based on left and right face depth cameras as described above, and the system includes:
[0076] The data acquisition module is used for:
[0077] Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face;
[0078] The point cloud segmentation module is used for:
[0079] Based on the facial feature point detection algorithm, the facial regions in the left and right grayscale images are identified respectively to obtain the facial feature point region ROI; according to the facial feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, the left face point cloud and the right face point cloud are segmented from the 3D point cloud data, and the segmented left face point cloud and right face point cloud are filtered and downsampled.
[0080] The mesh reconstruction module is used for:
[0081] Local feature descriptors of left and right face point clouds are calculated. Based on these descriptors, a sample consistency initial registration algorithm is used to coarsely register the left and right face point clouds, followed by a fine registration algorithm based on iterative nearest point to obtain a fine registration transformation matrix. The fine registration transformation matrix is then used to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain a complete face point cloud. Normals are calculated and Poisson surfaces are reconstructed from the complete face point cloud to generate an initial 3D face mesh model.
[0082] The texture generation module is used for:
[0083] Based on the ROI of the facial feature points, the corresponding texture regions in the left and right RGB images are located, and then color difference processing and brightness equalization are performed to obtain the processed left and right texture regions. The processed left and right texture regions are then stitched together according to a preset ratio to generate a complete texture map.
[0084] The model generation module is used for:
[0085] Using the nasal tip midline as the dividing line, the vertices in the initial 3D face mesh model are divided into left-side vertices and right-side vertices. For the left-side edge vertices located in the stitching gap region, their coordinates are adjusted to the coordinates of the nearest right-side edge vertex to eliminate the gap. Based on the intrinsic and extrinsic parameters of the left and right RGB cameras, the texture coordinates of each vertex in the initial 3D face mesh model in the texture map are calculated. Finally, a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information is output.
[0086] Please see Figures 3 to 7 To verify the effectiveness of this invention, the specific process of this invention is as follows:
[0087] S1: As Figure 3 As shown, connect the left and right depth cameras to the computer, run the program, and complete the camera initialization, including connecting the camera, setting the exposure and trigger mode, defining data, and opening the data stream channel;
[0088] After initialization, the user starts the left and right depth cameras to take pictures. Within 5 seconds, the grayscale images, RGB data and point cloud data of the left and right sides are collected. At the same time, the intrinsic parameter matrix and extrinsic parameter matrix of the RGB camera of the left and right depth cameras are obtained.
[0089] S2: As Figure 4 As shown, after completing the face data acquisition, the left and right grayscale images and the dlib face feature point detection model are loaded. After converting the left and right grayscale images to dlib format, the dlib face detector is used to detect face feature points on the left and right grayscale images.
[0090] After obtaining the facial feature points of the left and right grayscale images, their pixel coordinates are converted to Point2f format, and the Region of Interest (ROI) of the facial feature points in the left and right grayscale images are calculated respectively. The length and width of the left and right ROIs are then unified according to the smaller one.
[0091] After obtaining the face region ROI in the left and right grayscale images, load the left point cloud data. Utilize the one-to-one correspondence between the pixel number in the left grayscale image and the point number in the left point cloud data, for each pixel in the face region ROI of the left grayscale image, put the point in the point cloud data corresponding to the grayscale image pixel into the blank left face point cloud data variable to complete the segmentation of the left face point cloud.
[0092] For the right-side point cloud data, the right-side grayscale image and its face region ROI are used to similarly complete the segmentation of the right-side face point cloud; after completing the segmentation of the left and right face point clouds, the face point clouds are filtered and downsampled to obtain high-quality left and right face point clouds.
[0093] S3: As Figure 4 As shown, for the point clouds of the left and right faces, the normal vector and the local feature descriptor of the point cloud are calculated respectively. Then, using the two local feature descriptors, the sample consistency initial registration algorithm (SAC) is used to perform coarse registration to register the left face point cloud to the right face point cloud, and the coarse registration transformation matrix SAC_r is returned.
[0094] Next, the nearest point algorithm ICP is used to perform fine registration of the left and right face point clouds based on SAC_r, and the fine registration transformation matrix ICP_r is returned. Using ICP_r, the left face point cloud is transformed from the left depth camera world coordinate system to the right depth camera world coordinate system, and the transformed left face point cloud is fused with the right face point cloud, thus completing the registration of the left face point cloud to the right face point cloud. Finally, downsampling and point cloud smoothing are performed to obtain the complete face point cloud.
[0095] Based on the complete face point cloud, calculate its normal vector, fuse it with the face point cloud, and save it as a ply. First, use the SSDRecon tool to reconstruct the Poisson surface of the ply that has fused the point cloud and normal vector. Then, use the SurfaceTrimmer tool to remove low point density area patches from the reconstructed ply. Finally, obtain the initial 3D face mesh model ply that includes the ears.
[0096] S4: As Figure 5 As shown, by retrieving and mapping the face region ROI in the left and right grayscale images, the face region ROI in the left and right RGB data is calculated. Similarly, the length and width of the left and right ROIs are unified according to the smaller one, and the same redundancy processing is performed to obtain the texture mapping region rectangle ROI in the left and right RGB data.
[0097] Color difference processing is performed on the pixels in the left and right RGB data rectangles of the Region of Interest (ROI), including histogram matching, brightness equalization, and bilateral filtering, to obtain the ROI regions used for texture mapping of the 3D mesh model. The pixels in the left ROI region retain the left-side width portion (rate), and the pixels in the right ROI region retain the right-side width portion (rate is the width retention ratio). Then, the pixels of the two width portions are concatenated (e.g., ...). Figure 6 As shown in the image, this results in a complete texture map PNG that contains as many pixel areas as possible required for texture mapping.
[0098] The left vertex of the nose tip in the 3D face mesh model will be texture-mapped to the pixels of the left width portion, and the right vertex of the nose tip in the 3D face mesh model will be texture-mapped to the pixels of the right width portion; the pixels at the junction of the two width portions in the texture map will not be mapped, but will be accessed by the face formed by the left and right vertices of the nose tip as texture blocks, thus forming texture compression slits in a series of face blocks of the 3D face mesh model.
[0099] S5: Based on the obtained texture map png file, write a material library mtl file to access the texture map, load the 3D mesh model ply, extract the vertex v and face index f, and write it to the obj file;
[0100] First, write vertex v to the obj file: distinguish the left and right face vertices by the y-axis coordinates of the three-dimensional vertex corresponding to the nose tip pixel in the dlib face feature points of the left or right grayscale image. For the edge vertex on the right side of the boundary y-plane, construct its KD tree to move the edge vertex on the left side of the boundary y-plane to the edge neighbor vertex on the right side of the boundary y-plane.
[0101] like Figure 6As shown, for the 3D mesh model, each vertex is judged. If the vertex is not in the edge region of the left side of the boundary y-plane, the 3D coordinates of vertex v are directly written into the obj file. Otherwise, the nearest edge vertex on the right side of the boundary y-plane is searched and the 3D coordinates of the nearest edge vertex on the right side are written into the obj file to eliminate the texture compression gap at the splicing point.
[0102] Next, the texture coordinates VT are written to the OBJ file: using the extrinsic parameter matrix of the left and right RGB cameras, the three-dimensional coordinates in the coordinate system of the left and right face depth cameras can be converted into three-dimensional point coordinates in the coordinate system of the left and right RGB cameras. Then, using the pinhole imaging model and intrinsic parameter matrix, the three-dimensional points in the coordinate system of the left and right RGB cameras are converted into two-dimensional coordinates in the imaging plane coordinate system of the left and right RGB cameras, that is, the pixel coordinates in the RGB data.
[0103] For the vertex point on the left side of the face in the y-plane of the nose tip of the 3D mesh model, its coordinates are 3D coordinates in the coordinate system of the right face depth camera. It needs to be transformed back to the coordinate system of the left face depth camera through the fine registration matrix ICP_r, so as to obtain the 3D coordinates pointxyz of the vertex on the left side of the face in the coordinate system of the left face depth camera. Then, according to the extrinsic parameter matrices R and T of the left RGB camera relative to the left depth camera, the 3D coordinates of the vertex in the coordinate system of the left depth camera are transformed to the coordinate system of the left RGB camera. Finally, using the intrinsic parameter matrices Kc and K of the left RGB camera and the pinhole imaging model, the 2D pixel coordinates rgbxy of the pixel in the left RGB data corresponding to the 3D point point are calculated.
[0104] For the vertex point on the right side of the face in the y-plane of the nose tip point of the 3D mesh model, its coordinates are the 3D coordinates pointxyz in the right face depth camera coordinate system. Transform the vertex on the right side of the face from the right depth camera coordinate system to the right RGB camera coordinate system to obtain the 3D coordinates of the vertex in the right RGB camera coordinate system. Then, use the intrinsic parameter matrices Kc and K of the right RGB camera and the pinhole imaging model to calculate the 2D pixel coordinates rgbxy of the pixel in the right RGB data corresponding to the 3D point point.
[0105] like Figure 7 As shown, for a 3D mesh model, after calculating the pixel coordinates of the corresponding pixels in the initial RGB data for each vertex on the left and right sides, the pixel coordinates of the corresponding pixels in the texture map are calculated by using the position and size of the ROI of the left and right width parts of the texture map in the initial RGB data. This completes the mapping of all vertices of the 3D mesh model onto a single texture map. Then, based on the UV coordinate system, the texture coordinates (vt) of all vertices of the 3D mesh model can be calculated using the pixel coordinates in the texture map and the pixel size of the texture map, and written to the obj file.
[0106] For the set of vertices v written to obj, i.e., the vertex point cloud, calculate its normal vector vn and write it to the obj file; for the face 3D mesh model ply, directly write the face index to the obj file; after writing the MTL material library information, vertex v, texture coordinates vt, normal vector vn and face index f into the obj file, a face 3D model file that can access the texture map png through the material library MTL file is obtained.
[0107] The formula for calculating the conversion from two-dimensional pixel coordinates to texture coordinates is as follows:
[0108] ;
[0109] Where `height` and `width` represent the height and width of the left and right texture areas (rectangles) after unifying their size; `rate` represents the preset ratio; after stitching left and right according to the preset ratio, the height of the final texture map remains `height`, but the width becomes `width * rate * 2`; `1_rgbx` and `1_rgby` represent the pixel coordinates of the left vertex in the left RGB image; `r_rgbx` and `r_rgby` represent the pixel coordinates of the right vertex in the right RGB image; `1x` and `1y` represent the pixel coordinates of the top-left vertex of the left texture area (rectangle) in the left RGB image; `rx` and `ry` represent the pixel coordinates of the top-left vertex of the right texture area (rectangle) in the right RGB image; `1_px` and `1_py` represent the pixel coordinates of the left vertex in the final texture map; `r_px` and `r_py` represent the pixel coordinates of the right vertex in the final texture map; `vtx` represents the value after normalizing `px`, and `vty` represents the value after normalizing `py`; `vtx`, vty is the relative coordinate of a pixel on the final texture map, which is the texture coordinate written to the final OBJ file.
[0110] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0111] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0112] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0113] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A fast 3D face reconstruction method based on left and right face depth cameras, characterized in that, The method includes the following steps: Step 1: Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face; Step 2: Based on the face feature point detection algorithm, identify the face regions in the left and right grayscale images respectively and obtain the face feature point region ROI; according to the face feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, segment the left face point cloud and the right face point cloud from the 3D point cloud data, and perform filtering and downsampling processing on the segmented left face point cloud and right face point cloud. Step 3: Calculate the local feature descriptors of the left and right face point clouds. Based on these descriptors, perform coarse registration of the left and right face point clouds using the sample consistency initial registration algorithm, followed by fine registration using the iterative nearest point algorithm to obtain the fine registration transformation matrix. Use this matrix to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain the complete face point cloud. Perform normal calculation and Poisson surface reconstruction on the complete face point cloud to generate an initial 3D face mesh model. After Poisson surface reconstruction, a cropping step is included, as detailed below: Based on the spatial density distribution of the complete face point cloud, calculate the spatial occupancy grid; Using the central plane of the nose tip as the boundary, calculate the density of the point cloud in the spatial occupancy grid on the left and right sides respectively; The average density of the point cloud on the opposite side is calculated based on the density of the point cloud on the left and right sides in the spatial occupancy grid. The dynamic density threshold of the current side is set according to the average density of the point cloud on the opposite side, so as to identify and mark the sparse grids in the current side reconstructed grid whose density is lower than the dynamic density threshold. The triangular facets corresponding to the sparse grid are identified as useless facets caused by point cloud registration errors or noise, and the surface clipping tool is used to clip and remove them from the initial 3D face mesh model to obtain the initial 3D face mesh model. Step 4: Locate the corresponding texture regions in the left and right RGB images based on the ROI of the facial feature points, then perform color difference processing and brightness equalization to obtain the processed left and right texture regions; stitch the processed left and right texture regions together according to a preset ratio to generate a complete texture map. Step 5: Using the nasal tip midline as the dividing line, divide the vertices in the initial 3D face mesh model into left-side vertices and right-side vertices; for the left-side edge vertices located in the stitching gap area, adjust their coordinates to the coordinates of the nearest right-side edge vertex to eliminate the gap; based on the intrinsic and extrinsic parameters of the left and right RGB cameras, calculate the texture coordinates of each vertex in the initial 3D face mesh model in the texture map; finally, output a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information; specifically, adjusting the coordinates of the left-side edge vertices located in the stitching gap area to the coordinates of the nearest right-side edge vertex involves: constructing a KD tree for the right-side edge vertices; for the left-side edge vertices, search for their nearest right-side edge vertex using the KD tree, and modify the coordinates of the left-side vertex to the coordinates of the nearest right-side edge vertex.
2. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 1, characterized in that, In step 1, the depth camera is a structured light camera or a binocular vision depth camera, and each binocular vision depth camera integrates an RGB camera. While outputting grayscale images, RGB images and 3D point clouds, it also outputs the intrinsic and extrinsic parameter matrices of the RGB cameras on the left and right depth cameras.
3. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 2, characterized in that, In step 2, the face feature point detection algorithm uses the 68-point or 76-point detection model from the dlib library; the filtering process includes pass-through filtering and statistical filtering to remove outliers and noise.
4. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 3, characterized in that, In step 3, the initial registration algorithm for sample consistency is SAC-IA, and the iterative nearest point algorithm is ICP.
5. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 1, characterized in that, In step 4, color difference processing includes histogram matching and bilateral filtering to ensure that the color and brightness of the RGB images on both sides are consistent; in step 4, the preset ratio is 50%.
6. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 1, characterized in that, Step 5 involves calculating the texture coordinates of each vertex in the initial 3D face mesh model in the texture map. This is performed in the following ways depending on the face side to which the vertex belongs: For the vertices classified as left face side, the inverse of the fine registration transformation matrix is first used to transform the vertices on the left face side from the right depth camera coordinate system back to the left depth camera coordinate system, obtaining the 3D coordinates of the vertices in the left depth camera coordinate system; then, based on the extrinsic parameter matrix of the left RGB camera relative to the left depth camera, the 3D coordinates of the vertices in the left depth camera coordinate system are transformed to the left RGB camera coordinate system, obtaining the 3D coordinates of the vertices in the left RGB camera coordinate system. Finally, using the intrinsic parameter matrix of the left RGB camera and the pinhole imaging model, the three-dimensional coordinates of the vertex in the left RGB camera coordinate system are projected onto the left RGB image to obtain the two-dimensional pixel coordinates; based on the position of the left texture region ROI in the complete texture map, the two-dimensional pixel coordinates are mapped onto the complete texture map to obtain the texture coordinates of the vertex on the left side of the face in the complete texture map. For the vertices that are classified as the right side of the face, the vertices on the right side of the face are directly transformed from the right depth camera coordinate system to the right RGB camera coordinate system based on the extrinsic parameter matrix of the right RGB camera relative to the right depth camera, so as to obtain the three-dimensional coordinates of the vertices in the right RGB camera coordinate system. Then, using the intrinsic parameter matrix of the right RGB camera and the pinhole imaging model, the three-dimensional coordinates of the vertices in the right RGB camera coordinate system are projected onto the right RGB image to obtain two-dimensional pixel coordinates. Based on the position of the right texture region ROI in the complete texture map, the two-dimensional pixel coordinates are mapped onto the complete texture map to obtain the texture coordinates of the vertices on the right side of the face in the complete texture map.
7. The fast 3D face reconstruction method based on left and right face depth cameras according to claim 6, characterized in that, In step 5, the output 3D model file is in OBJ format, and includes an MTL material library file and a PNG texture map file.
8. A fast 3D face reconstruction system based on left and right face depth cameras, characterized in that, The system employs the fast 3D face reconstruction method based on left and right face depth cameras as described in any one of claims 1 to 7, and the system comprises: The data acquisition module is used for: Connect and initialize the left and right depth cameras, and simultaneously acquire the left and right grayscale images, left and right RGB images, and 3D point cloud data of the target face; The point cloud segmentation module is used for: Based on the facial feature point detection algorithm, the facial regions in the left and right grayscale images are identified respectively to obtain the facial feature point region ROI; according to the facial feature point region ROI and the correspondence between the pixels of the left and right grayscale images and the 3D point cloud data, the left face point cloud and the right face point cloud are segmented from the 3D point cloud data, and the segmented left face point cloud and right face point cloud are filtered and downsampled. The mesh reconstruction module is used for: Local feature descriptors of left and right face point clouds are calculated. Based on these descriptors, a sample consistency initial registration algorithm is used to coarsely register the left and right face point clouds, followed by a fine registration algorithm based on iterative nearest point to obtain a fine registration transformation matrix. The fine registration transformation matrix is then used to transform the left face point cloud to the coordinate system of the right face point cloud and fuse them to obtain a complete face point cloud. Normals are calculated and Poisson surfaces are reconstructed from the complete face point cloud to generate an initial 3D face mesh model. The texture generation module is used for: Based on the ROI of the facial feature points, the corresponding texture regions in the left and right RGB images are located, and then color difference processing and brightness equalization are performed to obtain the processed left and right texture regions. The processed left and right texture regions are then stitched together according to a preset ratio to generate a complete texture map. The model generation module is used for: Using the nasal tip midline as the dividing line, the vertices in the initial 3D face mesh model are divided into left-side vertices and right-side vertices. For the left-side edge vertices located in the stitching gap region, their coordinates are adjusted to the coordinates of the nearest right-side edge vertex to eliminate the gap. Based on the intrinsic and extrinsic parameters of the left and right RGB cameras, the texture coordinates of each vertex in the initial 3D face mesh model in the texture map are calculated. Finally, a textured 3D face model file containing vertex coordinates, texture coordinates, normal vectors, and face information is output.
Citation Information
Patent Citations
A real-time three-dimensional reconstruction method of a face
CN109242951A
Image processing method and device, storage medium and equipment
CN113129455A