A system for 3D face reconstruction based on deep learning

Through a three-dimensional face reconstruction system based on deep learning, using multi-angle images to acquire and train artificial neural network models, the problems of low accuracy of three-dimensional face reconstruction and poor expression distinction in the existing technology are solved, and high-precision three-dimensional face reconstruction and expression tracking are achieved.

CN113971715BActive Publication Date: 2025-08-26ARCSOFT CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010709985.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-22
Publication Date
2025-08-26
Estimated Expiration
2041-02-20

AI Technical Summary

Technical Problem

The existing three-dimensional face reconstruction technology and expression tracking technology have problems such as low reconstruction accuracy and low expression distinction. Further processing of the reconstructed image is required to obtain accurate three-dimensional reconstruction images.

Method used

A three-dimensional face reconstruction system based on deep learning is adopted, and images are obtained from different angles using the main color depth camera and multiple auxiliary color cameras. Combined with the processor and memory, the front and side viewing three-dimensional images of the benchmark truth three-dimensional model are generated by training the artificial neural network model, achieving high-precision three-dimensional face reconstruction and expression tracking.

Benefits of technology

The accuracy and expression distinction ability of three-dimensional face reconstruction are improved, the generated three-dimensional images are more accurate, and the ability to track large-view expressions is better, which avoids the shortcomings of inaccurate fitting of large-angle three-dimensional models in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971715B_ABST
    Figure CN113971715B_ABST
Patent Text Reader

Abstract

The present invention discloses a system for three-dimensional face reconstruction, comprising a primary color depth camera, multiple auxiliary color cameras, a processor, and memory. The primary color depth camera captures a primary color image and a primary depth image of a reference user from a frontal perspective. The multiple auxiliary color cameras capture multiple auxiliary color images of the reference user from multiple side perspectives. The memory stores multiple instructions. The processor executes the multiple instructions to establish a frontal perspective three-dimensional image of a reference truth three-dimensional model based on the primary color image and the primary depth image, generate multiple side perspective three-dimensional images of the reference truth three-dimensional model based on the frontal perspective three-dimensional image and the multiple auxiliary color images, and train an artificial neural network model based on the training images and the reference truth three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to face reconstruction, and in particular to a system for three-dimensional face reconstruction based on deep learning. Background Art

[0002] 3D face reconstruction and expression tracking technologies in computer vision are both programs that capture and establish facial shape and appearance, and are used in fields such as face recognition and expression detection. Existing 3D face reconstruction and expression tracking technologies generally suffer from low reconstruction accuracy and limited expression differentiation, requiring further processing of the reconstructed images to obtain accurate 3D reconstructions. Summary of the Invention

[0003] An embodiment of the present invention provides a system for three-dimensional face reconstruction based on deep learning, including a main color depth camera, multiple auxiliary color cameras, a processor and a memory. The main color depth camera is set at a front view position to obtain a main color image and a main depth image of a reference user from the front view position. The multiple auxiliary color cameras are set at multiple side view positions to obtain multiple auxiliary color images of the reference user from the multiple side view positions. The processor is coupled to the main color depth camera and the multiple auxiliary color cameras. The memory is coupled to the processor and stores multiple instructions. The processor executes the multiple instructions to establish a front view three-dimensional image of a reference truth three-dimensional model based on the main color image and the main depth image; generate multiple side view three-dimensional images of the reference truth three-dimensional model based on the front view three-dimensional image and the multiple auxiliary color images; and train an artificial neural network model based on the training image and the reference truth three-dimensional model. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 4 is a block diagram of a three-dimensional face reconstruction system according to an embodiment of the present invention.

[0005] Figure 2 yes Figure 1 Schematic diagram of the setup of the main color depth camera and auxiliary color camera in the system.

[0006] Figure 3 yes Figure 1 Flowchart of the training method of the artificial neural network model in the system.

[0007] Figure 4 yes Figure 3 Flowchart of step S302.

[0008] Figure 5 yes Figure 3 Flowchart of step S304.

[0009] Figure 6 yes Figure 3Flowchart of step S306.

[0010] Figure 7 yes Figure 4 Flowchart of step S404.

[0011] Figure 8 is a schematic diagram of the cropped training image in step S602.

[0012] Figure 9 yes Figure 1 Schematic diagram of the artificial neural network model in the system.

[0013] Figure 10 is used Figure 1 Flowchart of the 3D image reconstruction method based on the trained artificial neural network model.

[0014] The description of the accompanying drawings is as follows:

[0015] 1: 3D facial reconstruction system

[0016] 10: Processor

[0017] 100: Artificial Neural Network Model

[0018] 102: 3D deformation model

[0019] 104: Benchmark Truth 3D Model

[0020] 12: Memory

[0021] 14: Main color depth camera

[0022] 16(1) to 16(N): Auxiliary color camera

[0023] 18: Display

[0024] 19: Image sensor

[0025] 300: Training Methods

[0026] S302 to S306, S402 to S406, S502 to S508, S602 to S608, S702 to S722, S1002 to S1008: Steps

[0027] 900: Phase 1

[0028] 902: Phase 2

[0029] 904: Phase 3

[0030] 906: Stage 4

[0031] 908: Full connection stage

[0032] 910: Convolution stage

[0033] 912: Fully connected stage

[0034] 1000: Three-dimensional image reconstruction method

[0035] IP: Main color image

[0036] Is(1) to Is(N): Auxiliary color images

[0037] Dp: Main depth image

[0038] R: Reference user

[0039] Tex: Expression coefficient

[0040] Tid: face shape coefficient DETAILED DESCRIPTION

[0041] Figure 1 This is a block diagram of a 3D face reconstruction system 1 according to an embodiment of the present invention. The 3D face reconstruction system 1 can receive a 2D face image and perform 3D face reconstruction and expression tracking based on the 2D face image. The 3D face reconstruction system 1 can also be used for 3D face reconstruction, expression-driven, and avatar animation. The reconstructed face can be used to obtain 3D feature points, a face mask map, and face segmentation. It can also use face shape coefficients for face recognition and driver facial attribute analysis. The 3D face reconstruction system 1 can fit a 3D morphable model (3DMM) to the 2D face image to reconstruct the 3D face model. The 3D morphable model can be based on principal component analysis (PCA) and use multiple model coefficients to generate facial features of the 3D face model. For example, face shape coefficients can be used to control the face shape of the 3D face model, and expression coefficients can be used to control the expression of the 3D face model. Furthermore, the 3D face reconstruction system 1 can use an artificial neural network model to generate the required model coefficients. The artificial neural network model can be trained using a ground truth (GT) 3D model as a training target. The ground-truth 3D model can be generated based on actual measurements and can contain multiple accurate 3D images corresponding to multiple viewpoints, including a yaw angle range of -90° to 90° and a pitch angle range of -45° to 45°. Because the ground-truth 3D model covers 3D images from a wide range of angles, an artificial neural network model trained using the ground-truth 3D model can accurately predict the model coefficients for a wide-angle 3D face model. The artificial neural network model can also train facial shape coefficients and expression coefficients separately, increasing the accuracy of expression initialization and individual expressions.

[0042] The three-dimensional face reconstruction system 1 may include a processor 10, a memory 12, a primary color depth camera 14, auxiliary color cameras 16(1) to 16(N), a display 18, and an image sensor 19, where N is a positive integer, for example, N=18. The processor 10 may be coupled to the memory 12, the primary color depth camera 14, the auxiliary color cameras 16(1) to 16(N), the display 18, and the image sensor 19. The processor 10, the memory 12, the display 18, and the image sensor 19 may be integrated into a common device, which may be, for example, a mobile phone, a computer, or an embedded device. The processor 10 may include an artificial neural network model 100, a three-dimensional deformation model 102, and a GT three-dimensional model 104. The artificial neural network model 100 may be a convolutional neural network model. In some embodiments, the artificial neural network model 100 may be a visual geometry group (VGG) model, an AlexNet model, a GoogleNet Inception model, a ResNet model, a DenseNet model, a SENet model, a feature pyramid network (FPN) model, or a MobileNet model.

[0043] The three-dimensional face reconstruction system 1 can operate in a training phase and a face reconstruction phase. In the training phase, the three-dimensional face reconstruction system 1 can generate a GT three-dimensional model 104 and train an artificial neural network model 100 based on the training image and the GT three-dimensional model 104. In the face reconstruction phase, the three-dimensional face reconstruction system 1 can input the user's two-dimensional image into the trained artificial neural network model 100 to generate a three-dimensional model of the user, and display the three-dimensional model of the user on the display 18. The processor 10 can control the operation of the memory 12, the main color depth camera 14, the auxiliary color cameras 16 (1) to 16 (N), the display 18 and the image sensor 19 to perform the training phase and the face reconstruction phase. After the GT three-dimensional model 104 is generated, the connection between the main color depth camera 14 and the auxiliary color cameras 16 (1) to 16 (N) and the processor 10 can be cut off.

[0044] Figure 21 is a schematic diagram of the arrangement of the main color depth camera 14 and the auxiliary color cameras 16(1) to 16(18) in the three-dimensional face reconstruction system 1. The main color depth camera 14 can be arranged at a front view position of the reference user R, and the auxiliary color cameras 16(1) to 16(18) can be arranged at 18 side view positions of the reference user R, respectively. The front view position and the side view position can be defined by the yaw angle and pitch angle of the reference user R. The yaw angle is the rotation angle of the head of the reference user R around the z-axis, and the pitch angle is the rotation angle of the head of the reference user R around the y-axis. The main color depth camera 14 can be arranged at a position where the yaw angle is 0° and the pitch angle is 0°. The 18 side view positions can be evenly distributed in the range of -90° to 90° yaw angle and -45° to 45° pitch angle. For example, the auxiliary color camera 16(6) can be arranged at a position where the yaw angle is -90° and the pitch angle is 0°. The arrangement of the auxiliary color cameras 16(1) to 16(18) is not limited to Figure 2 The embodiments may also be arranged in other distribution manners, for example, within other yaw angle or pitch angle ranges.

[0045] The primary color depth camera 14 and the auxiliary color cameras 16(1) to 16(18) can substantially simultaneously capture images of the reference user R from different angles, obtaining 19 color camera images and 1 depth image of the face of the reference user R in one shot. The primary color depth camera 14 can capture a primary color image Ip and a primary depth map Dp of the reference user R from a frontal viewing position. The auxiliary color cameras 16(1) to 16(18) can respectively capture multiple auxiliary color images Is(1) to Is(18) of the reference user R from multiple side viewing positions.

[0046] The memory 12 may store a plurality of instructions. The processor 10 may execute the plurality of instructions stored in the memory 12 to perform the training method 300 in the training phase and the 3D image reconstruction method 1000 in the face reconstruction phase.

[0047] Figure 3 This is a flow chart of a training method 300 for the artificial neural network model 100 in the 3D face reconstruction system 1. The training method 300 includes steps S302 through S306. Steps S302 and S304 are used to prepare the GT 3D model 104. Step S306 is used to train the artificial neural network model 100. Any reasonable technical changes or adjustments to the steps fall within the scope of the present invention. Steps S302 through S306 are explained below:

[0048] Step S302: The processor 10 generates a GT based on the main color image Ip and the main depth image Dp

[0049] a three-dimensional image of the three-dimensional model 104 from an orthographic perspective;

[0050] Step S304: The processor 10 generates a plurality of side-view 3D images of the GT 3D model 104 according to the front-view 3D image and the plurality of auxiliary color images Is(1) to Is(N);

[0051] Step S306 : The processor 10 trains the artificial neural network model 100 according to the training image, the front-view 3D image, and the plurality of side-view 3D images.

[0052] In step S302, the processor 10 uses the main color image Ip and the main depth image Dp obtained from the front view position to perform high-precision expression fitting to generate an accurate front view three-dimensional image. Then, in step S304, the processor 10 uses the accurate front view three-dimensional image and the calibration parameters of the auxiliary color cameras 16(1) to 16(N) to perform reference truth migration for the perspective of the auxiliary color cameras 16(1) to 16(N) to generate other accurate multiple side view three-dimensional images. Finally, in step S306, the processor 10 uses the accurate front view three-dimensional image and the accurate multiple side view three-dimensional images to train an accurate artificial neural network model 100.

[0053] In some embodiments, the generation of the side-view 3D image in the GT 3D model 104 in step S304 can be replaced by using an existing pre-trained model to pre-process the large-pose image, and then manually adjust it, or other various methods of using consistency constraints between the front view and other views to perform baseline truth migration.

[0054] The training method 300 utilizes the primary depth image Dp to perform high-precision expression fitting to generate an accurate frontal three-dimensional image, and then performs reference truth migration to migrate the frontal three-dimensional image to other camera perspectives to generate accurate side-view three-dimensional images, thereby training an accurate artificial neural network model 100, avoiding the shortcomings of inaccurate large-angle three-dimensional model fitting in traditional facial reconstruction methods.

[0055] Figure 4 yes Figure 3 The flowchart of step S302 includes steps S402 to S406. Steps S402 to S406 are used to generate an orthographic 3D image. Any reasonable technical changes or step adjustments fall within the scope of the present invention. The following explains steps S402 to S406:

[0056] Step S402: The primary color and depth camera 14 acquires a primary color image Ip and a primary depth image Dp of a reference user R from a normal viewing angle;

[0057] Step S404: The processor 10 optimizes and fits the main color image Ip and the main depth image Dp to generate a posture, a frontal view facial coefficient group, and a frontal view expression coefficient group; Step S406: The processor 10 uses the three-dimensional deformation model 102 to generate a frontal view three-dimensional image based on the posture, the frontal view facial coefficient group, and the frontal view expression coefficient group.

[0058] In step S402, the primary color and depth camera 14 captures the frontal face of a reference user R to obtain a color image Ip and a primary depth image Dp. In step S404, the processor 10 performs landmark detection on the color image Ip and then optimizes and fits it with the primary depth image Dp to obtain a pose, a frontal face shape coefficient set, and a frontal expression coefficient set. The pose can be the head pose of the 3D model, indicating the head's orientation and position relative to the primary color and depth camera 14. The frontal face shape coefficient set can include a plurality of face shape coefficients, for example, 100 face shape coefficients, each representing a facial shape feature, such as a fat or thin face. The frontal expression coefficient set can include a plurality of expression coefficients, for example, 48 expression coefficients, each representing a facial expression feature, such as squinting and a slightly raised corner of the mouth. Finally, in step S406, the processor 10 generates an accurate frontal 3D image based on the pose, the frontal face shape coefficient set, and the frontal expression coefficient set.

[0059] Compared with the method of using only the main color image Ip, the orthographic face shape coefficient set and the orthographic expression coefficient set obtained by using the main depth image Dp and the main color image Ip are more accurate, thereby generating a more accurate orthographic 3D image.

[0060] Figure 5 yes Figure 3 The flowchart of step S304 includes steps S502 to S508. Step S502 is used to generate corresponding calibration parameters for the auxiliary color camera 16(n). Steps S504 to S508 are used to perform a reference truth migration based on the corresponding calibration parameters of the auxiliary color camera 16(n) to the viewing angle of the auxiliary color camera 16(n), thereby generating an accurate side-view 3D image. Any reasonable technical changes or step adjustments fall within the scope of the present invention. The following explains steps S502 to S506:

[0061] Step S502: Calibrate the auxiliary color camera 16 ( n ) according to the main color depth camera 14 to generate corresponding calibration parameters of the auxiliary color camera 16 ( n );

[0062] Step S504: the auxiliary color camera 16(n) acquires the auxiliary color image Is(n) of the reference user R;

[0063] Step S506: The processor 10 migrates the frontal view 3D image according to the corresponding calibration parameters of the auxiliary color camera 16(n) to generate a corresponding side view facial shape coefficient set and a corresponding side view expression coefficient set;

[0064] Step S508 : The processor 10 generates a corresponding side-view 3D image using the 3D deformation model 102 according to the auxiliary color image Is(n), the corresponding side-view face shape coefficient set, and the corresponding side-view expression coefficient set.

[0065] The auxiliary color camera 16(n) is one of the auxiliary color cameras 16(1) to 16(N), where n is a positive integer from 1 to N. In step S502, the auxiliary color camera 16(n) can be calibrated with the main color depth camera 14 as a reference camera to generate calibration parameters. The calibration parameters may include external parameters of the auxiliary color camera 16(n), and the external parameters may include rotation parameters, translation parameters, scaling parameters, affine translation parameters, and other external camera parameters. The main color depth camera 14 and the auxiliary color camera 16(n) may each have intrinsic parameters, and the intrinsic parameters may include lens deformation parameters, focal length parameters, and other internal camera parameters. In step S506, the processor 10 generates an orthographic three-dimensional image based on the posture, the orthographic face coefficient group, and the orthographic expression coefficient group, and migrates the orthographic three-dimensional image to the angle of the auxiliary color camera 16(n) based on the corresponding calibration parameters of the auxiliary color camera 16(n) to generate a corresponding side view face coefficient group and a corresponding side view expression coefficient group. In step S508, the processor 10 generates an accurate corresponding side-view 3D image based on the corresponding side-view facial shape coefficient set and the corresponding side-view expression coefficient set. Steps S502 to S508 may be performed alternately for the auxiliary color cameras 16(1) to 16(N) to generate the corresponding side-view 3D images for the auxiliary color cameras 16(1) to 16(N).

[0066] In step S304 , the perspective of the auxiliary color camera 16 ( n ) is transferred to the reference truth according to the calibration parameters, so as to generate an accurate corresponding side perspective 3D image.

[0067] Figure 6 yes Figure 3 The flowchart of step S306 includes steps S602 to S608. Step S602 is used to crop the training image to obtain a stable cropped training image. Steps S604 to S608 are used to train the artificial neural network model 100. Any reasonable technical changes or step adjustments fall within the scope of the present invention. The following explains steps S602 to S608:

[0068] Step S602: The processor 10 crops the training image to generate a cropped training image;

[0069] Step S604: The processor 10 inputs the cropped training image into the artificial neural network model 100 to generate a face shape coefficient set and an expression coefficient set;

[0070] Step S606: The processor 10 generates a 3D predicted image using the 3D deformation model 102 according to the facial shape coefficient set and the expression coefficient set;

[0071] Step S608 : The processor 10 adjusts the parameters of the artificial neural network model 100 to reduce the difference between the 3D predicted image and the GT 3D model 104 .

[0072] In step S602, the processor 10 performs face detection on the training image, then detects two-dimensional feature points of the face, selects a minimum bounding rectangle based on the two-dimensional feature points, appropriately enlarges the minimum bounding rectangle, and crops the training image based on the enlarged minimum bounding rectangle. The training image may be a two-dimensional image and may be captured by one of the image sensor 19, the primary color depth camera 14, and the auxiliary color cameras 16(1) to 16(N). Figure 8 is a schematic diagram of the cropped training image in step S602, where the circles are two-dimensional feature points, including two-dimensional contour points 80 and multiple internal points, and 8 represents the enlarged minimum enclosing rectangle. Two-dimensional contour points 80 may include jaw contour points, while the other multiple internal points may include eye contour points, eyebrow contour points, nose contour points, and mouth contour points. Processor 10 may select a minimum enclosing rectangle based on two-dimensional contour points 80. In some embodiments, to stabilize the image input to artificial neural network model 100, processor 10 may normalize the cropped training image for roll angles, then input the normalized image into the trained artificial neural network model 100 to generate a three-dimensional image of the user. In other embodiments, processor 10 may normalize the cropped training image to generate a cropped training image of a predetermined size, such as a 128-bit × 128-bit × 3-bit two-dimensional three-primary-color (red, green, blue, RGB) image, and then input the normalized image into the trained artificial neural network model 100 to generate a three-dimensional image of the user. In other embodiments, the processor 10 may perturb the minimum bounding rectangle of the training image to improve algorithm robustness. The perturbation may be performed by affine transformations such as translation, rotation, and scaling.

[0073] In step S604, the processor 10 inputs the cropped training image into the artificial neural network model 100, performing forward propagation to generate a face shape coefficient set and an expression coefficient set. Next, in step S606, the processor 10 uses the face shape coefficient set and the expression coefficient set in the 3D deformable model 102 based on principal component analysis to obtain a 3D model point cloud as a 3D predicted image. In step S608, the 3D predicted image, supervised by the GT 3D model 104, undergoes backpropagation within the artificial neural network model 100, ultimately bringing the 3D predicted image closer and closer to one of the frontal 3D image and multiple side 3D images of the GT 3D model 104. The artificial neural network model 100 uses a loss function to adjust its parameters to reduce the difference between the 3D predicted image and one of the frontal 3D image and multiple side 3D images of the GT 3D model 104. The parameters of the artificial neural network model 100 can be represented by a regression face shape coefficient matrix and a regression expression coefficient matrix. Processor 10 calculates face shape loss based on the regressed face shape coefficient matrix and expression loss based on the regressed expression coefficient matrix. By adjusting parameters in the regressed face shape coefficient matrix and the regressed expression coefficient matrix, the face shape loss and expression loss are reduced, respectively, thereby reducing the total face shape loss and expression loss. The trained artificial neural network model 100 can separately generate face shape coefficients and expression coefficients for the 3D deformable model 102, enabling more refined expression capabilities, eliminating 3D image initialization issues, and providing enhanced expression tracking capabilities across a wide viewing angle.

[0074] Figure 7 yes Figure 4 The flowchart of step S404 includes steps S702 to S722 for executing the optimization fitting procedure to generate a posture, a frontal view face coefficient group, and a frontal view expression coefficient group. Steps S702 to S708 are used to generate a depth point cloud corresponding to the feature points of the main color image Ip. Steps S710 to S710 are used to generate a posture, a frontal view face coefficient group, and a frontal view expression coefficient group based on the depth point cloud. Any reasonable technical changes or step adjustments fall within the scope disclosed by the present invention. The following explains steps S702 to S722:

[0075] Step S702: The processor 10 receives the main color image Ip;

[0076] Step S704: The processor 10 detects feature points of the main color image Ip;

[0077] Step S706: The processor 10 receives the primary depth image Dp;

[0078] Step S708 : The processor 10 generates a depth point cloud in the coordinate system of the primary color and depth camera 14 based on the feature points of the primary color image Ip and the primary depth image Dp;

[0079] Step S710 : The processor 10 generates a pose based on the depth point cloud and the interior points of the average 3D model using an iterative closest point (ICP) algorithm;

[0080] Step S712: The processor 10 generates 3D contour points of the orthographic 3D image according to the posture;

[0081] Step S714: The processor 10 updates the pose according to the 3D contour points and interior points of the orthographic 3D image;

[0082] Step S716: The processor 10 generates 3D contour points of the orthographic 3D image according to the updated posture;

[0083] Step S718: The processor 10 searches for corresponding points of the depth point cloud corresponding to the three-dimensional contour points;

[0084] Step S720: The processor 10 updates the frontal view face shape coefficient group according to the corresponding points of the depth point cloud. Step S722: The processor 10 updates the frontal view expression coefficient group according to the corresponding points of the depth point cloud. The process continues with step S714.

[0085] The processor 102 performs feature point detection on the primary color image Ip (step S704) and aligns the primary depth image Dp with the primary color image Ip. The inliers of the primary color depth camera 14 are then converted into a depth point cloud in the color depth camera 14 coordinate system based on the primary depth image Dp using the intrinsic parameters of the primary color depth camera 14 (step S708). The inliers are then compared with the 3D inliers in the face database using an iterative closest point (ICP) algorithm to initialize the pose (step S710). Based on the initialized pose, the processor 102 locates the extreme points of parallel lines on the orthographic 3D image of the GT 3D model 104 as the corresponding 3D contour points of the 2D feature points (step S712). The processor 102 then updates the pose using the 3D inliers and 3D contour points (step S714), and then updates the 3D contour points of the orthographic 3D image based on the updated pose (step S716). Next, processor 102 updates the face shape coefficient set (step S720) and expression coefficient set (step S722) based on the current pose by finding corresponding point pairs in the depth point cloud and corresponding vertices in the orthographic 3D image. Processor 102 then updates the orthographic 3D image using the updated face shape coefficient set and expression coefficient set, and re-updates the pose using the new 3D contour points and new 3D interior points in the updated orthographic 3D image (step S714). Repeating steps S714 through S722 several times in this manner yields a more accurate orthographic 3D image. In some embodiments, the face shape coefficient set and expression coefficient set may be manually adjusted to obtain a more accurate orthographic 3D image.

[0086] In step S404, the depth information is used to generate a more accurate orthographic 3D image through an optimization fitting process.

[0087] Figure 9 1 is a schematic diagram of an artificial neural network model 100. Artificial neural network model 100 includes a first stage 900, a second stage 902, a third stage 904, a fourth stage 906, a fully connected stage 908, a convolutional stage 910, and a fully connected stage 912. The first stage 900, the second stage 902, and the third stage 904 are executed sequentially. Following the third stage 904 are the fourth stage 906 and the convolutional stage 910. Following the fourth stage 906 is the fully connected stage 908. Following the convolutional stage 910 is the fully connected stage 912.

[0088] Artificial neural network model 100 can be a ShuffleNet V2 lightweight network. The training image can be a 128-bit × 128-bit × 3-bit two-dimensional RGB image. The training image is input into the first stage 900. After the first stage 900, the second stage 902, and the third stage 904, the artificial neural network model 100 is divided into two paths. The first path includes the fourth stage 906 and the fully connected stage 908, and the second path includes the convolution stage 910 and the fully connected stage 912. Through the first path, 48 expression coefficients Tex are obtained. Through the convolution stage 910 of the second path, two 3×3 convolution kernels are used for processing. Then, after the fully connected stage 908, 100 facial shape coefficients Tid are obtained. Since the facial shape coefficient Tid regresses facial features such as fatness and thinness, it can be separated after the third stage 904. The expression coefficient Tex requires more refined features for regression. For example, squinting eyes, slightly raised corners of the mouth, etc. require more abstract features, so it goes through the fourth stage 906 and then fully connected output.

[0089] The artificial neural network model 100 can utilize the features of networks of different depths to achieve optimal network performance.

[0090] Figure 10 3D image reconstruction method 1000 using a trained artificial neural network model 1 is shown as a flow chart. 3D image reconstruction method 1000 includes steps S1002 through S1008, which are used to generate a 3D image of the user based on the user image. Any reasonable technical changes or adjustments to these steps fall within the scope of the present invention. Steps S1002 through S1008 are explained below:

[0091] Step S1002: the image sensor 109 acquires a user image of the user;

[0092] Step S1004: the processor 10 detects a plurality of feature points in the user image;

[0093] Step S1006: The processor 10 crops the user image according to the plurality of feature points to generate a cropped image of the user;

[0094] Step S1008: The processor 10 inputs the cropped image into the trained artificial neural network model 100 to generate a three-dimensional model of the user.

[0095] In step S1008, the cropped user image is input into the trained artificial neural network model 100 to obtain a corresponding facial shape coefficient set and expression coefficient set. Finally, the 3D deformable model 102 is used to generate a 3D image of the user's 3D model based on the corresponding facial shape coefficient set and expression coefficient set. Because the artificial neural network model 100 is trained using the GT 3D model 104 and the loss function, the trained artificial neural network model 100 can separately generate the facial shape coefficients and expression coefficients of the 3D deformable model 102. The resulting 3D image of the user has more refined expressions, no 3D image initialization issues, and better expression tracking performance over a wide viewing angle.

[0096] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A system for 3D face reconstruction based on deep learning, characterized in that: include: a primary color and depth camera, disposed at a front viewing angle of a reference user, for acquiring a primary color image and a primary depth image of the reference user from the front viewing angle; a plurality of auxiliary color cameras, disposed at a plurality of side viewing positions of the reference user, for acquiring a plurality of auxiliary color images of the reference user from the plurality of side viewing positions; a processor coupled to the primary color depth camera and the plurality of auxiliary color cameras; and a memory, coupled to the processor, storing a plurality of instructions; wherein the processor executes the plurality of instructions to: generating an orthographic 3D image of a ground truth 3D model based on the primary color image and the primary depth image; generating a plurality of side-view 3D images of the reference truth 3D model based on the front-view 3D image and the plurality of auxiliary color images; and An artificial neural network model is trained according to the training image, the front-view 3D image, and the plurality of side-view 3D images.

2. The system according to claim 1, wherein: The processor optimizes and fits the main color image and the main depth image to generate a posture, a frontal view face shape coefficient, and a frontal view expression coefficient; and The processor generates the orthographic three-dimensional image using a three-dimensional deformable model according to the posture, the orthographic face shape coefficient, and the orthographic expression coefficient.

3. The system according to claim 2, characterized in that The processor: detecting a plurality of feature points in the primary color image; generating a depth point cloud in a coordinate system of the primary color and depth camera according to the plurality of feature points in the primary color image and the primary depth image; Generate a pose based on the depth point cloud and the interior points of the average 3D model using an iterative closest point algorithm; Generating three-dimensional contour points of the orthographic three-dimensional image according to the posture; Finding corresponding points of the depth point cloud corresponding to the three-dimensional contour points; and The frontal view face shape coefficient and the frontal view expression coefficient are updated according to the corresponding point of the depth point cloud.

4. The system according to claim 1, wherein: Calibrate the plurality of auxiliary color cameras according to the primary color depth camera to generate corresponding calibration parameters of the plurality of auxiliary color cameras; The processor migrates the front-view 3D image according to corresponding calibration parameters of one of the plurality of auxiliary color cameras to generate corresponding side-view facial shape coefficients and corresponding side-view expression coefficients; and The processor generates one of the plurality of side-view 3D images using a 3D deformation model according to the corresponding side-view facial shape coefficient, the corresponding side-view expression coefficient, and the corresponding auxiliary color image.

5. The system according to claim 1, wherein: Also includes: an image sensor, coupled to the processor, for acquiring an image of the user; wherein the processor further: Detecting a plurality of landmarks in the user image; cropping the user image according to the plurality of feature points to generate a cropped image of the user; and The cropped image is input into the trained artificial neural network model to generate a three-dimensional image of the user.

6. The system according to claim 5, characterized in that The processor: normalizing the cropped image to generate a normalized image; and The normalized image is input into the trained artificial neural network model to generate the three-dimensional image of the user.

7. The system according to claim 5, characterized in that Also includes: A display is coupled to the processor and is used to display the three-dimensional image of the user.

8. The system according to claim 1, wherein: The processor: Inputting the training image into the artificial neural network model to generate face shape coefficients and expression coefficients; generating a three-dimensional predicted image using a three-dimensional deformation model according to the face shape coefficient and the expression coefficient; and Parameters of the artificial neural network model are adjusted to reduce a difference between the 3D predicted image and one of the front-view 3D image and the plurality of side-view 3D images.

9. The system according to claim 1, wherein: The artificial neural network model is a convolutional neural network model.

10. The system according to claim 1, wherein: The processor defines the plurality of side viewing positions using a yaw angle and a pitch angle of the reference user.

Citation Information

Patent Citations

  • A method and apparatus for image processing

    CN109697688A

  • Method and Apparatus for image-based photorealistic 3Dface modeling

    KR1020050022306A