Face three-dimensional modeling method based on RGBD image

By combining RGB images and depth images, multi-frame point cloud registration and ICP precision registration, combined with facial key points information for depth completion and refinement, the noise, hollowness and texture distortion problems in RGBD images are solved, and high-precision three-dimensional face reconstruction is achieved.

CN120279169APending Publication Date: 2025-07-08XI AN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510268446.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing three-dimensional modeling methods for RGBD images have limitations in real-time and hardware costs. It is difficult to obtain complete depth information in depth image noise and hollow areas. Traditional point cloud registration methods are prone to local optimal problems when texture information is lacking. Texture distortion is caused by lighting changes and viewing angle differences during texture mapping.

Method used

By combining RGB images and depth images, multi-frame point cloud registration and ICP accurate registration, combined with facial key points information for depth completion and refinement, optimize face depth maps, and generate high-precision three-dimensional models.

Benefits of technology

The depth map cavity filling rate is improved to 98%, the depth error of key facial areas is reduced to 0.5mm, the registration accuracy of facial feature areas is improved to 1.0mm, and the accuracy and robustness of three-dimensional modeling are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279169A_ABST
    Figure CN120279169A_ABST
Patent Text Reader

Abstract

The invention discloses a face three-dimensional modeling method based on an RGBD image. The face three-dimensional modeling method comprises the following steps: S100, acquiring an RGB image and depth video data by using an RGBD camera; s200, performing face detection on the RGB image, and identifying and positioning a face area; s300, face key point detection is carried out in the face area, and face key point coordinate information is extracted; s400, inputting a current depth frame in the depth video data into a face depth completion network to carry out face depth completion, inputting a completed depth image into a face depth refining network, and refining and enhancing by combining the extracted face key point coordinate information to obtain an optimized face depth image; and S500, performing multi-frame point cloud registration on the optimized face depth image in combination with the face key point coordinate information to generate a complete three-dimensional face model. According to the method, the precision and robustness of face three-dimensional modeling are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical fields of computer vision, 3D perception, and image processing, and particularly relates to a method for 3D face modeling based on RGBD images. Background Technique

[0002] In recent years, 3D face modeling technology has received high attention in the fields of biometrics, intelligent monitoring, virtual reality, and medical imaging. Traditional 3D modeling methods such as structured light and stereovision technology, although capable of achieving high-precision 3D reconstruction, have significant limitations in terms of real-time performance and hardware cost, and are difficult to meet the needs of large-scale data processing. With the popularization of depth sensors, RGBD cameras have become an important tool in the field of 3D face modeling because they can simultaneously collect color images and depth information.

[0003] However, there are still some technical bottlenecks in directly using RGBD data for 3D reconstruction. First, due to sensor limitations, depth images often have noise and hole regions, and it is difficult to obtain complete depth information, especially in the facial edge and facial feature regions. Second, traditional point cloud registration methods such as the Iterative Closest Point (ICP) algorithm are prone to local optimum problems when facing the lack of texture information or key point constraints, resulting in unstable point cloud fusion. In addition, during the texture mapping process, illumination changes and viewing angle differences often lead to texture distortion, affecting the realism of the model. Summary of the Invention

[0004] To solve the above problems, the present disclosure provides a method for 3D face modeling based on RGBD images, including the following steps:

[0005] S100: Use an RGBD camera to collect RGB images and depth video data;

[0006] S200: Perform face detection on the RGB image to identify and locate the face region;

[0007] S300: Perform face key point detection within the face region to extract the coordinate information of face key points;

[0008] S400: Input the current depth frame in the depth video data into a face depth completion network for face depth completion, input the completed depth map into a face depth refinement network and refine and enhance it in combination with the extracted coordinate information of face key points to obtain an optimized face depth map;

[0009] S500: Perform multi-frame point cloud registration on the optimized face depth map in combination with the coordinate information of face key points to generate a complete 3D face model.

[0010] In addition, the present invention also discloses a three-dimensional face modeling device based on RGBD images, including:

[0011] A device for collecting RGB images and depth video data using an RGBD camera;

[0012] A device for performing face detection on the RGB image, identifying and locating the face area;

[0013] A device for performing face key point detection within the face area and extracting the coordinate information of face key points;

[0014] A device for inputting the current depth frame in the depth video data into a face depth completion network for face depth completion, inputting the completed depth map into a face depth refinement network and refining and enhancing it in combination with the extracted coordinate information of face key points to obtain an optimized face depth map;

[0015] A device for performing multi-frame point cloud registration on the optimized face depth map in combination with the coordinate information of face key points to generate a complete three-dimensional face model.

[0016] In addition, the present invention also discloses a computer storage medium, wherein the storage medium includes computer instructions, which, when running on a computer, cause the computer to execute the above method.

[0017] In addition, the present invention also discloses an electronic device, wherein the electronic device includes:

[0018] A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein

[0019] when the processor executes the program, the above method is implemented.

[0020] Compared with the prior art, the beneficial effects of the present disclosure are as follows:

[0021] First, through the combination of the depth map and the RGB image, and by using multi-frame point cloud registration and ICP precise registration, the present disclosure can effectively fuse information from multiple perspectives to generate a high-precision three-dimensional face model. Through the face key point information, the details of key face parts are optimized, overcoming the problem of insufficient facial feature details in traditional methods.

[0022] Second, by combining RGB face key point information, the present disclosure completes and refines the face depth map, effectively improving the quality of the face depth map and solving the problem of insufficient facial details in the face depth map. The hole filling rate of the depth map hole filling based on the traditional interpolation method is 90%, and this method increases the depth map hole filling rate to 98%. The depth error in the facial key area of a single depth map refinement method is 0.9 mm, and this method reduces the error to 0.5 mm.

[0023] Third, by introducing face key point information into the ICP precise registration, the present disclosure is more targeted and accurate when aligning point clouds. When processing facial feature areas, it can better maintain the accuracy and consistency of facial details, overcoming the limitation that the traditional ICP registration algorithm cannot effectively process facial feature areas. The registration error of the traditional ICP registration algorithm is 2 mm, and the ICP registration error with key point constraints is reduced to 1.2 mm. By assigning higher weights to key points, the registration accuracy of the facial feature area is improved from 1.2 mm to 1.0 mm. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of a method for 3D face modeling based on RGBD images provided in an embodiment of the present disclosure;

[0025] Figure 2 is a structural diagram of a face depth completion network provided in an embodiment of the present disclosure;

[0026] Figure 3 is a structural diagram of a face depth refinement network provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In one embodiment, as Figure 1 shown, the present disclosure provides a method for 3D face modeling based on RGBD images, including the following steps:

[0028] S100: Use an RGBD camera to collect RGB images and depth video data;

[0029] S200: Perform face detection on the RGB image to identify and locate the face area;

[0030] S300: Perform face key point detection within the face area to extract the coordinate information of face key points;

[0031] S400: Input the current depth frame in the depth video data into a face depth completion network for face depth completion, and input the completed depth map into a face depth refinement network and refine and enhance it in combination with the extracted coordinate information of face key points to obtain an optimized face depth map;

[0032] S500: Perform multi-frame point cloud registration on the optimized facial depth map in combination with the facial key point coordinate information to generate a complete 3D facial model.

[0033] For this embodiment, the present method is a method that can integrate the advantages of RGB images and depth images. It uses deep learning for depth map completion and facial detail enhancement, introduces key point constraints in point cloud registration, and at the same time, through an optimized texture mapping strategy, realizes a more realistic and complete 3D facial reconstruction solution, further improving the accuracy and robustness of 3D facial modeling.

[0034] This embodiment uses the Orbbec P1Pro camera to synchronously collect depth image and RGB image sequences at a frame rate of 30fps. The resolution of both is 640×400, and time synchronization processing is performed to ensure data consistency and spatio-temporal alignment accuracy.

[0035] The collected depth map sequence and RGB image sequence are simultaneously input into the modeling system. First, face detection is performed on the RGB image to identify and locate the face region. Facial key point detection is performed within this region to extract facial key point coordinate information to accurately calibrate facial feature points. At the same time, to reduce the interference of irrelevant regions on depth information, the depth values of the depth image regions outside the face detection box are set to zero, thus retaining the depth features of the facial region.

[0036] Subsequently, the cropped facial depth map is input into the depth completion network for depth map hole completion to repair the missing depth information caused by the limitations of the acquisition device. The completed depth map is further input into the facial depth map refinement network. In combination with the extracted key point information, the depth features of the facial region are refined and enhanced, especially for detail regions such as the tip of the nose, the corners of the eyes, and the corners of the mouth, improving the accuracy and integrity of the depth map.

[0037] After the depth map and RGB image processing are completed, in combination with the facial key point information, multi-frame point cloud registration is performed to ensure the accurate fusion of depth data from different perspectives. Finally, the RGB image is used for texture mapping of the 3D model to seamlessly fit the color information onto the surface of the 3D mesh, generating a 3D facial model with high fidelity.

[0038] Therefore, this embodiment can achieve high-precision 3D facial modeling, especially having significant advantages in terms of facial depth details and realism.

[0039] In another embodiment, step S200 includes the following steps:

[0040] S201: Cache the RGB video data to obtain an N-frame RGB image sequence.

[0041] S202: Detect faces in N frames of RGB images to obtain N face image detection frames.

[0042] S203: Evaluate and screen the N face image detection frames to obtain N optimal face image detection frames corresponding to the N-frame RGB image sequence.

[0043] S204: Crop the N optimal face image detection frames to obtain N face image regions.

[0044] For this embodiment, in step S201, the cached face video data can be consecutive frames or image frames collected at intervals. N is an integer and N > 4, and the N-frame face RGB image sequence obtained by caching is updated in a first-in-first-out manner.

[0045] In step S202, in the N-frame RGB image sequence, if faces are detected in both the first frame and the last frame, frame the face image region and output the face image detection frame; otherwise, return to step S201 to re-cache the RGB video data to update the existing N-frame face RGB image sequence, re-read the updated N-frame face RGB image sequence, and re-perform the detection.

[0046] It should be noted that a face detection network can be used to detect face RGB images frame by frame. Here, an object detection network including YOLOV5, RetinaFace, etc. can be used as the face detection network.

[0047] In step S203, a decision is made by calculating the intersection over union of the areas of the face RGB image detection frames corresponding to two adjacent frames in the N-frame face RGB image sequence. The formula for the intersection over union is as follows:

[0048]

[0049] Where A and B are the areas of the detection frames corresponding to two adjacent frames of face RGB images before and after respectively. If the intersection over union IOU is greater than the threshold S, select the detection frame corresponding to the previous frame of RGB image as the common detection frame for these two adjacent frames of RGB images; otherwise, the actual detection frame of the latter frame of RGB image is used. By analogy, a decision is made on the detection frame corresponding to each frame of RGB image to achieve the effect of stable display of the detection frame.

[0050] In another embodiment, step S300 further includes the following steps:

[0051] S301: Extract the facial key points of each face RGB image region and obtain the coordinates of each key point;

[0052] S302: Optimize the facial key points in each facial RGB image region to obtain the optimized facial key points;

[0053] S303: According to the optimized facial key points, reconstruct and adjust the facial geometry to provide a geometric basis for subsequent depth completion refinement and 3D modeling.

[0054] For this embodiment, in step S301, through the facial key point detection algorithm, the cropped facial RGB image region is processed to identify the key feature points of the face, such as the coordinates of the eyes, nose, mouth and other parts.

[0055] In step S302, the optimization of the facial key points of N frames includes the following steps:

[0056] S3021: Correct the facial key points based on the context information. By analyzing the relative position relationship between facial features, correct the positions of the initially detected key points. Use geometric constraints and local region relationships to adjust inaccurate or abnormal key points to ensure that the key points are more reasonable within the facial feature region;

[0057] S3022: Use an adaptive optimization algorithm to optimize the key point positions. Match the extracted facial key points with the facial shape model to minimize the error between the key points and the model. Through the iterative optimization process, make the position of each key point more accurate and reduce the deviation caused by the error in the initial detection process;

[0058] S3023: According to the facial pose estimation algorithm, adjust the positions of the facial key points to adapt to different angles and ensure high accuracy in different viewing perspectives.

[0059] In another embodiment, step S400 includes the following steps:

[0060] S401: Denoise the original depth map through the Gaussian filtering algorithm to smooth the small errors in the depth map and eliminate the inaccurate depth information caused by sensor noise;

[0061] In this step, Gaussian filtering can effectively smooth the image while retaining most of the important structural information, providing high-quality input for subsequent completion and refinement processing.

[0062] S402: Input the denoised depth map into a pre-trained face depth completion network. This network learns the spatial and context relationships of the facial depth map through a deep learning model and automatically infers and completes the missing regions in the depth map. The completed depth map should fully restore the facial features and avoid unnatural depth value changes in the completed regions;

[0063] S403: Combine the coordinate information of the facial key points extracted in step S300, and input the complemented face depth map into the face depth map refinement network for further processing. Through deep learning algorithms, the refinement network finely adjusts and enhances the local details in the depth map for the key features in facial regions such as the nose, eyes, and mouth, thereby improving the depth performance in these regions and enhancing the accuracy and quality of the depth map;

[0064] S404: Further process the refined face depth map using a filtering algorithm to remove redundant noise and interference, ensuring that the overall quality of the finally output depth map is optimized.

[0065] In step S402, the depth completion process for the face depth map includes the following steps:

[0066] S4021: Take the depth map after denoising as input data and pass it into the face depth map completion network. This network consists of multiple convolutional neural network layers, designed to extract spatial features and context information in the face depth map; the network structure belongs to the encoder-decoder architecture, and a network similar to the UNet structure can be used as the face depth completion network;

[0067] S4022: The depth completion network extracts the global and local features of the depth map through convolutional neural networks, performs pixel-level processing on the image, and learns the context relationship of the facial structure through a deep learning model, especially the interdependence between different facial regions;

[0068] S4023: According to the depth information and facial structure model in the pre-trained data, the depth completion network infers the depth values of the missing parts through the information of the known regions. In this step, through the network model of the encoder-decoder model, the original input depth map is encoded and feature-extracted, and feature fusion is gradually performed, decoded and restored to the depth map, inferring and complementing the missing regions in the depth map, and generating depth values that conform to the pre-facial morphology;

[0069] S4024: Correct the image after depth completion to ensure a natural transition between the complemented region and the original region, and eliminate obvious depth jumps and discontinuities;

[0070] S4025: Output the face depth map after completion and correction, laying a foundation for further processing and optimization.

[0071] In step S403, the depth refinement process for the face depth map includes the following steps:

[0072] S4031: Align the completed depth map with the key point information in combination with the facial key point coordinate information obtained in step S300. The key point information provides important structural constraints for the refinement network, enabling the refinement process to focus on the facial feature regions;

[0073] S4032: Input the completed depth image combined with the key point information into a pre-trained face depth refinement network. The refinement network consists of multiple convolutional neural networks, which can learn to improve the quality of the depth map of the facial region at the microscopic level of the image; the network input is the completed depth map and the face key point heat map. The key point information provides important structural constraints for the refinement network, enabling the network refinement process to focus on the facial feature regions and effectively enhancing the depth details of the feature regions;

[0074] S4033: The refinement network refines the local regions of the facial depth map, such as the nose, mouth, and eye regions, to enhance the details of these regions, strengthen the depth gradient, and enhance the facial features;

[0075] S4034: Optimize the connection between the enhanced region and the original part to avoid abrupt depth jumps;

[0076] S4035: Output the refined face depth map.

[0077] In another embodiment, the face depth completion network includes three modules: an encoder, a bottleneck, and a decoder.

[0078] For this embodiment, as Figure 2 shown, the encoder generates feature maps of different scales from the original face depth map through a series of convolutions and downsamplings. After passing through the bottleneck and the decoder, through a series of convolutions and upsamplings, the depth map is gradually restored. The feature maps of different scales are fused in terms of features, and finally the completed face depth map is output.

[0079] The encoding module is composed of N cascaded feature extraction units. Each feature extraction unit includes a convolution calculation layer, a feature activation layer, and a downsampling layer. The convolution calculation layer extracts spatial features from the input feature map through a convolution kernel of a preset size; the downsampling layer reduces the resolution of the feature map through a strided convolution to generate low-dimensional feature expressions of different levels. The output feature maps of each feature extraction unit are parameter-transferred along the channel dimension to form a multi-scale feature map set.

[0080] The bottleneck module is composed of K cascaded splicing modules, which perform non-linear transformation and feature recombination on the deep features output by the encoding module to establish cross-channel feature correlation relationships.

[0081] The decoding module includes M feature reconstruction units (M = N), and each feature reconstruction unit is composed of a transposed convolutional layer, a feature concatenation layer, and a convolutional calculation layer. The transposed convolutional layer increases the resolution of the feature map through learnable parameters; the feature concatenation layer performs channel dimension fusion on the upsampled features of the current layer and the feature map of the corresponding level of the encoding module to form an enhanced feature representation; the convolutional calculation layer reconstructs the spatial information of the fused features. After the outputs of the respective feature reconstruction units are cascaded, a completed depth map is generated through the terminal convolutional layer.

[0082] The network realizes progressive depth recovery of the missing region through multi-level feature extraction of the encoding module, context information fusion of the bottleneck module, and multi-scale feature concatenation of the decoding module. In particular, the feature maps output by each level of the encoding module are channel-concatenated with the feature maps of the corresponding level of the decoding module through cross-layer skip connections to form a feature enhancement mechanism with spatial perception ability.

[0083] In another embodiment, the face depth refinement network includes a feature extraction module, a feature fusion module, and a depth recovery module.

[0084] For this embodiment, as Figure 3 shown, the two feature extraction modules extract the features of the completed depth map and the heat map of face key points respectively through convolution and downsampling, and generate feature maps of different scales. The feature fusion module fuses the feature maps of different scales and transmits them to the depth recovery module, and finally recovers the depth map and outputs the refined face depth map.

[0085] The feature extraction module includes a depth map feature branch and a key point heat map branch arranged in parallel. The depth map feature branch receives the completed depth map as input and extracts multi-scale spatial features step by step through cascaded convolutional layers and learnable downsampling layers. Each downsampling level outputs a geometric detail feature map with different resolutions; the key point heat map branch receives the Gaussian heat map generated from the face key point coordinates and extracts multi-level semantic feature maps through the same number of convolutional and downsampling operations. The heat map encodes the probability distribution of the key point positions to strengthen the spatial prior of the facial feature regions.

[0086] The feature fusion module receives the same-level multi-scale feature maps from two feature extraction branches and performs cross-modal feature alignment and fusion. For the same-level feature maps, after unifying the channel dimensions through 1×1 convolution, a spatial attention mechanism is used to calculate pixel-level fusion weights, and the depth features and semantic features are dynamically weighted and superimposed to generate a fusion feature map with spatial perception ability; for the cross-level feature maps, the high-level low-resolution feature maps are upsampled by transposed convolution and then concatenated with the low-level high-resolution feature maps in channels, and 3×3 convolution is used to achieve the collaborative expression of detail enhancement and semantic constraint; in the highest-level fusion process, a global context vector is introduced, and the overall facial structure features are captured through global average pooling and broadcast to each fusion level after being reconstructed by a fully connected layer to inject global geometric consistency constraints.

[0087] The depth recovery module performs progressive depth reconstruction based on the fused multi-scale feature maps. Among them, cascaded feature reconstruction units are used to upsample and fuse cross-level features step by step. Each unit realizes the resolution recovery through 2×2 transposed convolution, and then is concatenated with the fusion features of the corresponding level and the local depth value distribution is refined through 3×3 convolution; residual convolution is added to the upsampling path, and skip connections are constructed through 1×1 convolution and depthwise separable convolution to eliminate the checkerboard artifacts introduced by the deconvolution operation and improve the continuity of depth values; an edge perception constraint is superimposed on the end output layer, and the spatial alignment between the refined depth map and the facial key point coordinates is supervised through an auxiliary loss function to ensure the geometric accuracy of the facial contour edges.

[0088] The network realizes the depth refinement of the face depth map through the multi-level feature extraction of the feature extraction module, the cross-modal feature fusion of the feature fusion module, and the progressive depth reconstruction of the depth recovery module, effectively enhancing the features of the facial depth. The introduction of the facial key point information can effectively provide coordinate information for depth refinement and improve the accuracy of depth refinement of the face.

[0089] In another embodiment, step S500 further includes the following steps:

[0090] S501: According to the internal parameters of the RGBD camera, convert the optimized face depth map into three-dimensional point cloud data;

[0091] S502: Select the frontal depth point cloud data as the reference frame for the registration of other frame point clouds;

[0092] S503: Perform preliminary alignment on the point clouds from different frames. By using the feature matching algorithm, align the point cloud data at different angles with the reference frame;

[0093] S504: Use the iterative closest point ICP algorithm to perform fine registration on the preliminarily aligned point clouds in combination with the facial key point coordinate information;

[0094] S505: After precisely aligning multiple frames of point clouds, all the aligned point clouds are merged into a unified point cloud set, and the point clouds from different angles are integrated in one coordinate system to form a complete face point cloud;

[0095] S506: Denoise the merged point cloud set to remove the abnormal points and noise points during the registration process, obtaining optimized point cloud data;

[0096] S507: Based on the optimized point cloud data, use the triangular mesh generation algorithm to construct a three-dimensional face mesh model;

[0097] S508: On the basis of generating the mesh, perform mesh optimization to remove the redundant vertices and faces and perform smoothing processing;

[0098] S509: Map the color information of the RGB image onto the three-dimensional face mesh model to achieve texture mapping;

[0099] S510: Output the complete three-dimensional face model.

[0100] For this embodiment, in S501, the three-dimensional coordinates are calculated through the depth image pixel coordinates and depth values, and at the same time, the RGB pixel color information corresponding to the coordinates is associated with the point cloud coordinates, so that the point cloud data contains both geometric positions and color texture information. The three-dimensional point cloud data represents the position structure of the human face in the three-dimensional space.

[0101] In S502, the reference frame usually selects the point cloud from the front view of the face as the reference frame point cloud, and this reference frame will be used as the benchmark for registering other frames, and all other point cloud data and texture data will be registered according to this coordinate system.

[0102] In S507, by connecting the adjacent points in the point cloud, a complete face mesh is formed, representing the three-dimensional geometric shape of the face.

[0103] In another embodiment, step S504 further includes the following steps:

[0104] S5041: Select the point cloud of the current frame as the source point cloud, the point cloud of the reference frame as the target point cloud, initialize the maximum number of iterations, error threshold, and according to the preliminary calculation of the key points of the human face, initialize the rotation matrix and translation matrix of the source point cloud;

[0105] S5042: Extract the three-dimensional coordinates of the respective key points of the human face of the source point cloud and the target point cloud, and preliminarily match the key points in the source point cloud and the target point cloud according to the structural characteristics of the human face;

[0106] S5043: For each point in the source point cloud, find the point with the minimum distance in the target point cloud by calculating the Euclidean distance;

[0107] S5044: During each iteration of ICP, calculate the rotation matrix and translation vector between the source point cloud and the target point cloud, and assign a higher weight to the position information of the facial key points, so that the alignment between the facial key points is given priority when calculating the transformation matrix, ensuring the precise registration of the facial structure;

[0108] S5045: Apply the transformation matrix to each point in the source point cloud and transform it into a new coordinate system;

[0109] S5046: After each iteration of ICP, calculate the error of the current point cloud registration, and calculate the error of the positions of the facial key points in the source point cloud and the target point cloud, ensure that the three-dimensional position alignment degree of the facial key points meets the predetermined requirements, and judge whether the stop condition is satisfied;

[0110] S5047: If the registration error is lower than the threshold, stop the iteration; if the maximum number of iterations is reached but the threshold requirement is still not met, force the iteration to stop;

[0111] S5048: When the iteration stop condition is reached, output the source point cloud after fine registration.

[0112] For this embodiment, S5041 is generally initialized as the identity matrix.

[0113] In S5043, it can be optimized by the KD-tree acceleration method. In addition to the conventional nearest neighbor point matching, it is also necessary to ensure that the facial key points in the source point cloud and the target point cloud are aligned as much as possible. By calculating the distances between the key points in the source point cloud and the target point cloud, adjust the matching strategy to reduce the error of other point cloud matching.

[0114] In S5046, the registration error is the average distance between the matching points in the source point cloud and the target point cloud.

[0115] In another embodiment, step S509 further includes the following steps:

[0116] S5091: Extract the RGB image of the current frame, and according to the coordinate information of the facial key points, crop out the area containing the face and extract the texture information as the basis of the texture map;

[0117] S5092: According to the generated three-dimensional face model, map each three-dimensional point to the two-dimensional texture coordinate system, and use the UV texture coordinate generation technology to calculate the texture coordinates of each triangular mesh;

[0118] S5093: Map the pixels of the texture image to the surface of the three-dimensional model according to the calculated texture coordinates, ensuring that the facial features are presented in the correct positions and proportions on the model;

[0119] S5094: Use a texture smoothing algorithm to optimize the texture edges and reduce texture noise and discontinuities caused by image projection or lighting changes;

[0120] S5095: Output a three-dimensional face model with completed texture mapping.

[0121] In another embodiment, a three-dimensional face modeling device based on RGBD images includes:

[0122] A device for collecting RGB images and depth video data using an RGBD camera;

[0123] A device for performing face detection on the RGB image to identify and locate the face area;

[0124] A device for performing face key point detection within the face area to extract the coordinate information of face key points;

[0125] A device for inputting the current depth frame in the depth video data into a face depth completion network for face depth completion, inputting the completed depth map into a face depth refinement network and combining the extracted coordinate information of face key points for refinement and enhancement to obtain an optimized face depth map;

[0126] A device for performing multi-frame point cloud registration on the optimized face depth map in combination with the coordinate information of face key points to generate a complete three-dimensional face model.

[0127] In addition, the present invention also discloses a computer storage medium, wherein the storage medium includes computer instructions, and when it runs on a computer, it causes the computer to execute any of the methods described above.

[0128] In addition, the present invention also discloses an electronic device, wherein the electronic device includes:

[0129] A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein

[0130] When the processor executes the program, it implements any of the methods described above.

[0131] Although the embodiments of the present invention have been described above in combination with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and these all belong to the scope of protection of the present invention.

Claims

1. A 3D face modeling method based on RGBD images, comprising the following steps: S100: Use an RGBD camera to collect RGB images and depth video data; S200: Perform face detection on the RGB images to identify and locate the face region; S300: Perform face key point detection within the face region to extract the coordinate information of face key points; S400: Input the current depth frame in the depth video data into a face depth completion network for face depth completion, input the completed depth map into a face depth refinement network and combine the extracted coordinate information of face key points for refinement and enhancement to obtain a refined face depth map; S500: Combine the coordinate information of the face key points to perform multi-frame point cloud registration on the optimized face depth map to generate a complete 3D face model.

2. The method according to claim 1, preferably, step S300 further comprises the following steps: S301: Extract the face key points of each face RGB image region and obtain the coordinates of each key point; S302: Optimize the face key points of each face RGB image region to obtain optimized face key points; S303: Reconstruct and adjust the facial geometry according to the optimized face key points to provide a geometric basis for subsequent depth completion refinement and 3D modeling.

3. The method according to claim 1, wherein the face depth completion network comprises three modules: an encoder, a bottleneck, and a decoder.

4. The method according to claim 1, wherein the face depth refinement network comprises a feature extraction module, a feature fusion module, and a depth recovery module.

5. The method according to claim 1, step S500 further comprises the following steps: S501: Convert the optimized face depth map into 3D point cloud data according to the internal parameters of the RGBD camera; S502: Select the frontal depth point cloud data as a reference frame for use as a reference for point cloud registration of other frames; S503: Perform preliminary alignment on the point clouds from different frames. By applying a feature matching algorithm, preliminarily align the point cloud data at different angles with the reference frame; S504: Use the Iterative Closest Point (ICP) algorithm to perform fine registration on the preliminarily aligned point clouds in combination with the coordinate information of the face key points; S505: After precisely aligning the multi-frame point clouds, merge all the aligned point clouds into a unified point cloud set. The point clouds from different angles are integrated in a coordinate system to form a complete face point cloud; S506: Denoise the merged point cloud set to remove the abnormal points and noise points during the registration process to obtain optimized point cloud data; S507: Use a triangular mesh generation algorithm to construct a 3D face mesh model according to the optimized point cloud data; S508: On the basis of generating the mesh, perform mesh optimization to remove redundant vertices and faces and perform smoothing; S509: Map the color information of the RGB image onto the 3D face mesh model to achieve texture mapping; S510: Output a complete 3D face model.

6. The method according to claim 5, wherein step S504 further comprises the following steps: S5041: Select the point cloud of the current frame as the source point cloud, the point cloud of the reference frame as the target point cloud, initialize the maximum number of iterations, the error threshold, and according to the preliminary calculation of the facial key points, initialize the rotation matrix and translation matrix of the source point cloud; S5042: Extract the three-dimensional coordinates of the facial key points of the source point cloud and the target point cloud respectively, and preliminarily match the key points in the source point cloud and the target point cloud according to the structural characteristics of the human face; S5043: For each point in the source point cloud, find the point with the minimum distance in the target point cloud by calculating the Euclidean distance; S5044: In each iteration of ICP, calculate the rotation matrix and translation vector between the source point cloud and the target point cloud, and assign a higher weight to the position information of the facial key points, so that the alignment between the facial key points is given priority when calculating the transformation matrix, and ensure the accurate registration of the facial structure; S5045: Apply the transformation matrix to each point in the source point cloud and transform it into a new coordinate system; S5046: After each iteration of ICP, calculate the error of the current point cloud registration, and calculate the error of the positions of the facial key points in the source point cloud and the target point cloud, ensure that the three-dimensional position alignment degree of the facial key points meets the predetermined requirements, and determine whether the stop condition is satisfied; S5047: If the registration error is lower than the threshold, stop the iteration; if the maximum number of iterations is reached but the threshold requirement is still not met, force the iteration to stop; S5048: When the iteration stop condition is reached, output the source point cloud after fine registration.

7. The method according to claim 5, wherein step S509 further comprises the following steps: S5091: Extract the RGB image of the current frame, and according to the facial key point coordinate information, crop the area containing the human face, extract the texture information, as the basis of the texture map; S5092: According to the generated three-dimensional human face model, map each three-dimensional point to a two-dimensional texture coordinate system, and use the UV texture coordinate generation technology to calculate the texture coordinates of each triangular mesh; S5093: Map the pixels of the texture image to the surface of the three-dimensional model according to the calculated texture coordinates, and ensure that the facial features are presented in the correct positions and proportions on the model; S5094: Use the texture smoothing algorithm to optimize the texture edges and reduce the texture noise and discontinuity caused by image projection or lighting changes; S5095: Output the three-dimensional human face model with texture mapping completed.

8. A three-dimensional human face modeling device based on RGBD images, comprising: A device for collecting RGB images and depth video data using an RGBD camera; A device for performing face detection on the RGB image, identifying and locating the face area; A device for performing face key point detection within the face area and extracting the coordinate information of the facial key points; A device for performing face depth completion on the current depth frame input in the depth video data through a face depth completion network, inputting the completed depth map into a face depth refinement network and performing refinement and enhancement in combination with the extracted face key point coordinate information to obtain an optimized face depth map; A device for performing multi-frame point cloud registration on the optimized face depth map in combination with the face key point coordinate information to generate a complete three-dimensional face model.

9. A computer storage medium, wherein, The storage medium includes computer instructions which, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

10. An electronic device, wherein, The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • 4D face reconstruction method and device

    CN120876784A