Face Three-Dimensional Reconstruction Method, Computer Device and Medium Based on Deep Learning

Through deep learning methods, deep reconstruction is performed using the trained network model and camera parameters to segment the image blocks, which solves the problem of insufficient three-dimensional reconstruction accuracy of faces in the prior art, and achieves high-precision face reconstruction.

CN116109778BActive Publication Date: 2025-07-22NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310191074.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2025-07-22
Estimated Expiration
2043-03-02

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision reconstruction in three-dimensional reconstruction of human faces, and the amount of network parameters increases rapidly with the image resolution, making it difficult to effectively improve the reconstruction accuracy.

Method used

Through a deep learning-based method, the trained rough matching network model is used to generate predicted optical flow and virtual camera parameters, the face image block is segmented and the initial depth map is generated. Combined with the surface reconstruction network and the decoder, the depth information of each image block is gradually reconstructed to achieve high-precision face reconstruction.

Benefits of technology

The accuracy of face reconstruction is improved, and the problem of rapid increase in network parameters with image resolution is avoided, thereby achieving high-precision face reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109778B_ABST
    Figure CN116109778B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for three-dimensional face reconstruction based on deep learning, a computer device and a medium, relating to the technical field of three-dimensional face reconstruction. The method includes predicting, by a trained coarse matching network model, multiple different perspective images of a target face to obtain a predicted optical flow from each perspective image to the target perspective image, generating a rough face according to real camera parameters, dividing the rough face into several image blocks, generating an initial depth map corresponding to each image block by virtual camera parameters, then obtaining a surface prediction code of each image block through a trained surface reconstruction network, and decoding through a trained surface decoder to obtain the reconstructed depth value of each pixel point on the initial depth map corresponding to each image block, and obtaining a reconstructed face according to all the above-mentioned reconstructed depth values. The present invention can perform depth reconstruction on each image block respectively through a deep learning model, and can achieve high-precision face reconstruction with a small number of network parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional face reconstruction, and particularly to a three-dimensional face reconstruction method, a computer device and a medium based on deep learning. Background Art

[0002] The indoor three-dimensional face reconstruction technology mainly restores the three-dimensional shape of the face from information such as images and radars, and has a wide range of applications in many fields such as virtual reality, human-computer interaction, and game graphics. Three-dimensional face reconstruction is a very important issue in the field of computer vision, and how to perform high-precision reconstruction is one of the more challenging tasks in the academic and industrial circles at present. Summary of the Invention

[0003] The purpose of the present invention is to provide a three-dimensional face reconstruction method, a computer device and a medium based on deep learning, which can improve the accuracy of face reconstruction.

[0004] To achieve the above purpose, the present invention provides the following solutions:

[0005] A three-dimensional face reconstruction method based on deep learning, the method includes:

[0006] S1: Obtain multiple images of the target face from different perspectives;

[0007] S2: For each of the perspective images, taking the perspective image as the source perspective image, inputting the source perspective image and the target perspective image into a trained rough matching network model to obtain the predicted optical flow from the source perspective image to the target perspective image; the target perspective image is the perspective image other than the source perspective image among all the perspective images; the trained rough matching network model is a model trained with the sample source perspective image and the sample target perspective image as inputs and the sample optical flow from the sample source perspective image to the sample target perspective image as the label;

[0008] S3: Fuse all the perspective images according to the predicted optical flow and the true camera parameters corresponding to the perspective images to generate a rough face of the target face;

[0009] S4: Segment the rough face into several image blocks, generate the virtual camera parameters corresponding to each image block, and generate the initial depth map corresponding to the image block through the virtual camera parameters corresponding to the image block;

[0010] S5: For each of the said image patches, input all the said perspective images and the initial depth map corresponding to the image patch into the trained surface reconstruction network to obtain the surface prediction encoding of the image patch; input the surface prediction encoding and the coordinates of each pixel point on the initial depth map corresponding to the image patch into the trained surface decoder to obtain the reconstructed depth value of each pixel point on the initial depth map corresponding to the image patch; the trained surface reconstruction network is a model trained with all the sample perspective images and sample initial depth maps of the sample face as the input and the sample surface encoding as the label; the trained surface decoder is a model trained with the sample surface prediction encoding and sample point coordinates as the input and the true depth value corresponding to the sample point as the label;

[0011] S6: Determine the reconstructed face of the target face based on the reconstructed depth values of each pixel point on the initial depth maps corresponding to all the said image patches.

[0012] Optionally, S3 specifically includes:

[0013] Generate the true depth map corresponding to each of the said perspective images according to the predicted optical flow and the true camera parameters corresponding to the perspective image;

[0014] Fuse the true depth maps corresponding to all the said perspective images to generate the rough face of the target face.

[0015] Optionally, the trained coarse matching network model includes an RGB feature extraction module and an optical flow prediction module connected in sequence;

[0016] The RGB feature extraction module includes a number of convolutional layers connected in sequence, and is used for feature extraction of the source perspective image and the target perspective image;

[0017] The optical flow prediction module uses a U-Net network and is used to obtain the predicted optical flow from the source perspective image to the target perspective image according to the extracted features.

[0018] Optionally, generating the virtual camera parameters corresponding to each of the said image patches specifically includes:

[0019] For each of the said image patches, perform the following steps:

[0020] Process the image patch using the principal component analysis method to obtain three eigenvectors;

[0021] Sort the three eigenvectors in descending order of eigenvalues, denote the eigenvector ranked first as the first eigenvector, the eigenvector ranked second as the second eigenvector, and the eigenvector ranked third as the third eigenvector;

[0022] Take the first eigenvector and the second eigenvector as the x-axis and y-axis of the virtual camera respectively, and take the reverse direction of the third eigenvector as the z-axis of the virtual camera to generate the virtual camera coordinate system of the virtual camera corresponding to the image patch;

[0023] Determine the true coordinates of the first eigenvector, the second eigenvector, and the third eigenvector in the world coordinate system respectively;

[0024] Determine the external parameter rotation matrix R according to the true coordinates;

[0025] Determine the external parameter translation matrix T according to the external parameter rotation matrix R;

[0026] Determine the virtual camera coordinates of each image point according to the coordinates of the image points on the image patch, the external parameter rotation matrix R, and the external parameter translation matrix T;

[0027] Determine the scaling factor s according to the maximum value in the x-axis direction and the maximum value in the y-axis direction among the virtual camera coordinates of all image points;

[0028] Generate the external parameters of the virtual camera according to the external parameter rotation matrix R, the external parameter translation matrix T, and the scaling factor s;

[0029] Determine the internal parameters of the virtual camera according to the resolution of the initial depth map corresponding to the image patch; the external parameters of the virtual camera and the internal parameters of the virtual camera constitute the virtual camera parameters of the virtual camera.

[0030] Optionally, the trained surface reconstruction network includes a feature pyramid network, a feature cross-correlation module, and a surface coding regression module connected in sequence;

[0031] The feature pyramid network is used to extract features from each of the perspective images to obtain the features of each of the perspective images;

[0032] The feature cross-correlation module is used to select several search points in the initial depth map corresponding to the image patch; for each perspective image, project the coordinates of each search point into the image coordinate system corresponding to the perspective image based on the true camera parameters corresponding to the perspective image to obtain the projected coordinates of each search point in the image coordinate system corresponding to the perspective image, and calculate the perspective features corresponding to each search point in the perspective image based on the features of the perspective image and the projected coordinates; for each search point, perform pairwise cross-correlation calculations on the perspective features corresponding to the search point in all the perspective images to obtain the cross-correlation calculation results of each search point; fuse the cross-correlation calculation results of all the search points to obtain the depth direction cost volume;

[0033] The surface coding regression module is used to encode the features of each of the perspective images, the depth direction cost volume, and the initial depth map corresponding to the image block, so as to obtain the surface prediction coding of the image block.

[0034] Optionally, before S5, it further includes: training the surface decoder, and the training process is as follows:

[0035] Obtain a first sample set; the first sample set includes the sample initial depth map of the sample face, the sample point coordinates, and the true depth value corresponding to the sample point;

[0036] Use the first sample set to train the surface coding and decoding network to obtain a trained surface coding and decoding network; the trained surface coding and decoding network includes a trained surface encoder and a trained surface decoder connected in sequence.

[0037] Optionally, the loss function adopted during the training process of the surface coding and decoding network includes a depth loss function and a normal vector loss function;

[0038] The expression of the depth loss function is:

[0039]

[0040] where loss d represents the depth loss function value, n represents the number of pixel points on the sample initial depth map, represents the true depth value of the i-th pixel point on the sample initial depth map; represents the reconstructed depth value of the i-th pixel point on the sample initial depth map;

[0041] The expression of the normal vector loss function is:

[0042]

[0043] where loss n represents the normal vector loss function value, represents the true normal vector of the i-th pixel point on the sample initial depth map, represents the predicted normal vector of the i-th pixel point on the sample initial depth map.

[0044] The determination process of the predicted normal vector is as follows:

[0045] Select adjacent pixel points of the pixel points on the sample initial depth map in the x-axis and y-axis directions respectively to obtain x adjacent pixel points and y adjacent pixel points;

[0046] Connect the x adjacent pixel points, the y adjacent pixel points and the pixel points to obtain a triangular patch;

[0047] Determine the virtual camera coordinates of the x-adjacent pixel points, the y-adjacent pixel points, and the pixel point according to the reconstruction depth value of the x-adjacent pixel points, the y-adjacent pixel points, and the pixel point and the virtual camera parameters corresponding to the sample initial depth map;

[0048] Determine the direction vector of each side of the triangular patch according to the virtual camera coordinates of the x-adjacent pixel points, the y-adjacent pixel points, and the pixel point;

[0049] Select any two sides of the triangular patch, perform a cross product on the direction vectors of the two sides to obtain the predicted normal vector of the pixel point on the sample initial depth map.

[0050] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above-mentioned face three-dimensional reconstruction method based on deep learning.

[0051] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to perform the above-mentioned face three-dimensional reconstruction method based on deep learning.

[0052] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The face three-dimensional reconstruction method, computer device, and medium based on deep learning provided by the present invention predict multiple different perspective images of a target face through a trained coarse matching network model to obtain the predicted optical flow from each perspective image to the target perspective image. Then, a rough face is generated according to the predicted optical flow and the real camera parameters. The rough face is segmented into several image blocks, and an initial depth map corresponding to each image block is generated through the virtual camera parameters corresponding to the image blocks. Then, the trained surface reconstruction network encodes all the perspective images and the initial depth map of the target face to obtain the surface prediction encoding of each image block. Then, the trained surface decoder decodes the surface prediction encoding to obtain the reconstruction depth value of each pixel point on the corresponding initial depth map of each image block. Finally, a reconstructed face is obtained based on the reconstruction depth value. The present invention restores depth information based on the matching information between multi-perspective images through a deep learning model, improving the accuracy of face reconstruction. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 Schematic flow chart of the 3D face reconstruction method based on deep learning provided by the present invention;

[0055] Figure 2 Collection device diagram of multi-view face images provided by the present invention;

[0056] Figure 3 Schematic structural diagram of the surface encoding and decoding network provided by the present invention;

[0057] Figure 4 Schematic principle diagram of the 3D face reconstruction method based on deep learning provided by the present invention;

[0058] Figure 5 Schematic structural diagram of a computer device provided by the present invention.

[0059] Symbol description:

[0060] 1000 - Computer device; 1001 - Processor; 1002 - Communication bus; 1003 - User interface; 1004 - Network interface; 1005 - Memory. Detailed implementation manners

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0062] In recent years, with the continuous improvement of computer performance, deep learning algorithms have been widely used in the field of vision, especially in the field of 3D face reconstruction. For example, the multi-view matching reconstruction method. The multi-view matching reconstruction method draws on the idea of stereo matching and searches for matching points between views to restore depth information. Although this method can obtain the accurate depth of the face, the reconstruction accuracy is limited by the resolution of the multi-view images themselves, and the number of network parameters increases exponentially with the resolution, making it difficult to achieve high-precision face reconstruction.

[0063] Based on the deficiencies of the above-mentioned prior art, the present invention provides a method for three-dimensional face reconstruction based on deep learning, a computer device and a medium. By using a deep learning model to restore depth information based on the matching information between multi-view images, the accuracy of face reconstruction is improved. Moreover, in the present invention, depth reconstruction is performed on each image block respectively, that is, by reconstructing only the local depth of a fixed size each time, the number of network parameters is decoupled from the image resolution, avoiding the problem that the number of network parameters increases rapidly with the image resolution, and enabling high-precision face reconstruction with a relatively small number of network parameters.

[0064] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] As Figure 1 shown, the present invention provides a method for three-dimensional face reconstruction based on deep learning, and the method includes:

[0066] S1: Obtain multiple different-view images of the target face.

[0067] S2: For each of the view images, using the view image as the source view image, input the source view image and the target view image into the trained coarse matching network model to obtain the predicted optical flow from the source view image to the target view image; the target view image is the view image other than the source view image among all the view images; the trained coarse matching network model is a model trained with the sample source view image and the sample target view image as inputs and the sample optical flow from the sample source view image to the sample target view image as the label.

[0068] S3: Fuse all the view images according to the predicted optical flow and the true camera parameters corresponding to the view images to generate a rough face of the target face.

[0069] S4: Segment the rough face into several image blocks, generate the virtual camera parameters corresponding to each image block, and generate the initial depth map corresponding to the image block through the virtual camera parameters corresponding to the image block.

[0070] S5: For each of the image patches, input all the perspective images and the initial depth map corresponding to the image patch into the trained surface reconstruction network to obtain the surface prediction encoding of the image patch; input the surface prediction encoding and the coordinates of each pixel point on the initial depth map corresponding to the image patch into the trained surface decoder to obtain the reconstructed depth value of each pixel point on the initial depth map corresponding to the image patch; the trained surface reconstruction network is a model trained with all the sample perspective images and sample initial depth maps of the sample face as inputs and the sample surface encoding as the label; the trained surface decoder is a model trained with the sample surface prediction encoding and sample point coordinates as inputs and the true depth value corresponding to the sample point as the label.

[0071] S6: Determine the reconstructed face of the target face based on the reconstructed depth values of each pixel point on the initial depth maps corresponding to all the image patches.

[0072] First, collect the perspective images of the target face. The collection device is as Figure 2 shown. The collection device consists of twelve single-lens reflex cameras, one surface light source, and thirteen polarizing films. The cameras are hardware-synchronized through signal lines and are controlled by the same shutter for shooting. Since the delay of hardware synchronization is at the microsecond level, it can be considered that the faces collected by each camera are at the same moment. Polarizing films are used to block in front of the camera lenses and the light source to ensure that the direction of the polarizing film in front of each camera is orthogonal to the polarizing film in front of the light source. The way to ensure that the polarizing films in front of the camera and the light source are orthogonal is: photograph a metal ball and adjust the position of the polarizing film until there is no highlight on the metal ball photographed by the camera. Then, perform preprocessing to remove anisotropy on the photos collected by the camera to obtain the perspective images of the target face. Among them, the preprocessing to remove anisotropy includes white balance removal, inverse gamma transformation, calibration of the response curves of each camera, and camera response inverse transformation operations.

[0073] Then, take each perspective image as the source perspective image one by one, and take the other perspective images except the source perspective image as the target perspective images. Input the source perspective image and each target perspective image into the trained coarse matching network model to output the predicted optical flow from the source perspective image to each target perspective image.

[0074] Among them, the trained coarse matching network model includes an RGB feature extraction module and an optical flow prediction module connected in sequence;

[0075] The RGB feature extraction module includes several convolutional layers connected in sequence, and is used to extract features from the source perspective image and the target perspective image. In this embodiment, the number of convolutional layers is 5.

[0076] The optical flow prediction module uses a U-Net network to obtain the predicted optical flow from the source perspective image to the target perspective image based on the extracted features. After the predicted optical flow from the source perspective image to the target perspective image is predicted by the coarse matching network model, the binocular vision depth calculation method is used to obtain the depth according to the predicted optical flow and the real camera parameters, and this depth is used for the supervised training of the coarse matching network model. That is, in this embodiment, fine tune is performed on the dataset used in the training process of the coarse matching network model, and the depth loss function is used for supervised training. In this embodiment, the PWC Net model is selected as the coarse matching network model.

[0077] In this embodiment, S3 specifically includes:

[0078] Generate the real depth map corresponding to each perspective image according to the predicted optical flow and the real camera parameters corresponding to the perspective image.

[0079] Fuse the real depth maps corresponding to all perspective images to generate the rough face of the target face.

[0080] Specifically, the predicted optical flow obtained in S2 is combined with the real camera parameters corresponding to the current source perspective image, and the binocular vision depth calculation method is used to calculate the real depth map corresponding to the current source perspective image. Then, after multiple loop calculations, the real depth maps corresponding to all perspective images are obtained, and the TSDF algorithm is used to fuse all real depth maps, and finally a three-dimensional rough face (rough mesh) with low precision and lack of details is generated.

[0081] After obtaining the rough mesh, the rough mesh is segmented into several image blocks (local patches), and the segmentation process is as follows:

[0082] Randomly select a rough mesh point from the rough mesh. According to the connection relationship provided by the triangular patches of the rough mesh, an adjacency matrix π is established, where π[i,j] represents whether the i-th point is connected to the j-th point (each rough mesh point in the rough mesh is connected to several points). If it is connected, it is 1, otherwise it is 0. Subsequently, according to the Markov chain, π is continuously multiplied by itself ten times to obtain the neighborhood points within ten orders of the rough mesh point (the points connected to the rough mesh point are first-order neighborhood points, and the points connected to the first-order neighborhood points are second-order neighborhood points). The points corresponding to the non-zero elements in the matrix obtained by the first multiplication are first-order neighborhood points, and so on. The points corresponding to the non-zero elements in the matrix obtained by the tenth multiplication are tenth-order neighborhood points. The rough mesh point and all neighborhood points within the first-order to tenth-order neighborhood points form a local patch. Loop the above process until the rough mesh is completely segmented.

[0083] After obtaining the above image patches (local patches), virtual cameras corresponding to each of the image patches are generated to find the perspective that is most suitable for unfolding the patch (image patch), facilitating the reconstruction of patch details. In S4, the virtual camera parameters corresponding to each of the image patches are generated, specifically including:

[0084] For each of the image patches, the following steps are performed:

[0085] The principal component analysis method is used to process the image patch to obtain three eigenvectors.

[0086] The three eigenvectors are sorted in descending order of eigenvalues. The eigenvector ranked first is denoted as the first eigenvector, the eigenvector ranked second is denoted as the second eigenvector, and the eigenvector ranked third is denoted as the third eigenvector.

[0087] The first eigenvector and the second eigenvector are respectively used as the x-axis and y-axis of the virtual camera, and the opposite direction of the third eigenvector is used as the z-axis of the virtual camera to generate the virtual camera coordinate system of the virtual camera corresponding to the image patch.

[0088] The true coordinates of the first eigenvector, the second eigenvector, and the third eigenvector in the world coordinate system are respectively determined.

[0089] The external parameter rotation matrix R is determined according to the true coordinates.

[0090] The external parameter translation matrix T is determined according to the external parameter rotation matrix R.

[0091] The virtual camera coordinates of each image point are determined according to the coordinates of the image points on the image patch, the external parameter rotation matrix R, and the external parameter translation matrix T.

[0092] The scaling coefficient s is determined according to the maximum value in the x-axis direction and the maximum value in the y-axis direction among the virtual camera coordinates of all image points.

[0093] The external parameters of the virtual camera are generated according to the external parameter rotation matrix R, the external parameter translation matrix T, and the scaling coefficient s.

[0094] The internal parameters of the virtual camera are determined according to the resolution of the initial depth map corresponding to the image patch; the external parameters of the virtual camera and the internal parameters of the virtual camera constitute the virtual camera parameters of the virtual camera.

[0095] The specific process is as follows: Perform principal component analysis on the points in the generated patch. The obtained eigenvectors are all 1×3 vectors. Specifically, if there are n points in the patch, a 3×n matrix is obtained according to the coordinates of the n points in the world coordinate system. The 3×n matrix is processed using the principal component analysis method to obtain three 1×3 eigenvectors. Since the eigenvectors are orthogonal to each other, the two eigenvectors with the largest eigenvalues (the first eigenvector and the second eigenvector) are respectively used as the x-axis direction vector and y-axis direction vector of the virtual camera, and the reverse of the third eigenvector is used as the z-axis direction vector of the virtual camera. Determine the true coordinates of the first eigenvector, the second eigenvector, and the third eigenvector in the world coordinate system according to the lengths of the three eigenvectors on the x, y, and z axes of the virtual camera, and thus determine the external parameter rotation matrix R, that is, invert the 3×3 matrix composed of the true coordinates of these three eigenvectors to obtain the rotation matrix R. Multiply the 3×n matrix by the rotation matrix R to obtain the rotated coordinates, and take the average value of the x, y, and z values of the rotated coordinates to obtain a point, which is the center point of the patch. Then move the center point of the patch in the reverse direction of the z-axis by the focal length to determine the optical center position. Multiply the coordinates of the optical center position by the rotation matrix R to obtain the coordinates in the world coordinate system, and then take the opposite number to obtain the translation matrix T. Calculate the scaling factor s through s = 1 / max(rang x ,range y ), where rang x ,range y represent the maximum ranges of the x value and y value of the patch in the virtual camera coordinate system respectively. Furthermore, obtain the final external parameters s×R and s×T of the virtual camera. The internal parameters of the virtual camera are based on the resolution of the initial depth map corresponding to the patch. For example, if the resolution of the initial depth map is required to be 32*32, then the parameters f x ,f y ,c x ,c y in the virtual camera internal parameter K are all 16.

[0096] Through the above, the virtual camera parameters corresponding to each image patch can be obtained, and the initial depth map corresponding to each image patch is calculated based on the virtual camera parameters. That is, K×(rotation matrix R×coordinates of points on the patch + translation matrix T), and thus the initial depth corresponding to the points on the patch is obtained.

[0097] For each image patch, input all the above perspective images and the initial depth map corresponding to the image patch into the trained surface reconstruction network, and the surface prediction coding of the image patch can be obtained.

[0098] In this embodiment, the above trained surface reconstruction network includes a feature pyramid network, a feature cross-correlation module, and a surface coding regression module connected in sequence.

[0099] The feature pyramid network is used to extract features from each of the perspective images, obtaining the features of each of the perspective images. As Figure 4 shown, the feature pyramid network FPN consists of six layers of convolution and four layers of deconvolution, and performs cross-layer connection on the features obtained by convolution of the same size and the features obtained by deconvolution.

[0100] The feature cross-correlation module is used to select a number of search points in the initial depth map corresponding to the image patch; for each of the perspective images, based on the true camera parameters corresponding to the perspective image, project the coordinates of each of the search points into the image coordinate system corresponding to the perspective image, obtaining the projected coordinates of each of the search points in the image coordinate system corresponding to the perspective image, and calculate the perspective features corresponding to each of the search points in the perspective image based on the features of the perspective image and the projected coordinates; for each of the search points, perform pairwise cross-correlation calculation on the perspective features corresponding to the search point in all the perspective images, obtaining the cross-correlation calculation result of each of the search points; fuse the cross-correlation calculation results of all the search points, obtaining the depth direction cost volume. Specifically:

[0101] In the surface reconstruction network, taking the virtual camera perspective corresponding to the current patch as the main perspective (source perspective), and the true perspectives of all perspective images as the target perspectives. Project the position coordinates of the source perspective image pixels at different depths in the virtual camera coordinate system into the target perspective, extract the features at the corresponding positions, and perform pairwise cross-correlation on the features between perspectives. The cross-correlation result is Corr(f i , f j ) = <f i , f j >, where f i , f j are the features of the i-th and j-th target perspectives respectively, and <> represents the inner product. Finally, obtain the cost volume in the depth search direction (depth direction cost volume), specifically:

[0102] Sample 2k + 1 search points at a fixed interval in the initial depth map corresponding to the patch. Assume that the pixel coordinates of the current search point position are (u, v), the depth is d, and the fixed interval is r. Then the depths of the 2k + 1 searched points are d - kr, d - (k - 1)r,..., d, d + r, d + 2r... d + kr respectively. For each search point, extract the feature vectors at the projected positions of the search point in each perspective according to the depth of the search point, and perform pairwise cross-correlation. The internal and external parameters of the virtual camera corresponding to the initial depth map are K0, P0, then its corresponding three-dimensional coordinate p = Prj -1(u, v, d, K0, P0) is projected onto the target view of a certain view image I1, and the coordinates in the corresponding true camera coordinate system of this view are obtained as (u1, v1) = Prj(p, K1, P1), where K1 and P1 are the true internal and external camera parameters of view image I2. The feature of the corresponding point is f1 = BIL(FPN(I1), u1, v1), where Prj represents the projection process, BIL represents bilinear interpolation, and FPN represents the Feature Pyramid Network. Similarly, the current search point is projected onto the view corresponding to another view image I2 to obtain the feature f2 = BIL(FPN(I2), u2, v2), and the calculation result of cross-correlation is Corr(f1, f2) = <f1, f2>, where <> represents the inner product. The cross-correlation results of all search points are stacked along the channel dimension to obtain the cost volume V, with a size of where n is the number of views. Among them, H and W are the length and width of the initial depth map of the patch respectively.

[0103] The surface encoding regression module is used to encode the features of each view image, the depth direction cost volume, and the initial depth map corresponding to the image patch to obtain the surface prediction encoding of the image patch. The features of multiple views, the depth direction cost volume, and the initial depth of the patch are input into the encoding regression module together, and the input shape size is n is the number of views, c is the feature channel dimension output by FPN, and then the surface prediction encoding (implicit encoding code) and the decoding operator multiplier are output. In this embodiment, the surface encoding regression module is composed of 4 layers of ResNet Block and two layers of fully connected networks.

[0104] After obtaining the above surface prediction encoding, pixel points are uniformly sampled on the initial depth map corresponding to the image patch. The pixel point coordinates (u, v) and the surface prediction encoding are input into the trained surface decoder together to obtain the reconstructed depth value of each pixel point. Finally, a high-precision reconstructed depth map can be obtained based on the reconstructed depth values of the pixel points corresponding to all patches.

[0105] Before S5, it also includes: training the surface decoder, and the training process is as follows:

[0106] Obtain the first sample set; the first sample set includes the sample initial depth map, sample point coordinates, and the true depth value corresponding to the sample points of the sample face.

[0107] Use the first sample set to train the surface encoding and decoding network to obtain the trained surface encoding and decoding network; the trained surface encoding and decoding network includes a trained surface encoder and a trained surface decoder connected in sequence.

[0108] In this embodiment, the structure of the above-mentioned surface encoding and decoding network is as follows Figure 3 As shown, the surface encoder takes a depth map with a resolution of 32*32 as input. The network consists of 4 ResNet blocks, as well as two heads, namely the multiplier head and the code head. The 4 ResNet blocks are responsible for extracting the features of the depth map and providing them as input to the subsequent head modules. The multiplier head is composed of two layers of 3*3 convolutional networks, with an output dimension of (B, 32, 2, 2), which is used as the operator encoding of the patch and is denoted as multiplier; the code head consists of one convolutional network and one fully connected network, with an output of (B, 64), which is used as the shape encoding (surface prediction encoding) of the depth map and is denoted as code. Then, an arbitrary point (pixel point) is queried on the pixel plane of the virtual camera corresponding to the patch. The decoder takes the query point coordinates (u, v) and the multiplier and code obtained by the encoder as input. After the matrix dot product of the (u, v) coordinates and the multiplier, it is concatenated with the code along the channel dimension and input to the subsequent MLP network (surface decoder), and finally the depth is output. In this embodiment, the surface decoder adopts an MLP structure containing 12 fully connected layers.

[0109] During the training process of the above-mentioned surface encoding and decoding network, the loss function used includes a depth loss function and a normal vector loss function; the depth loss function is the two-norm error between the true depth value and the reconstructed depth value of the pixel points on the sample initial depth map, and the normal vector loss function is the cosine value of the angle between the true normal direction and the predicted normal direction. Among them, the expression of the depth loss function is:

[0110]

[0111] Among them, loss d represents the value of the depth loss function, n represents the number of pixel points on the sample initial depth map, represents the true depth value of the i-th pixel point on the sample initial depth map; represents the reconstructed depth value of the i-th pixel point on the sample initial depth map;

[0112] The expression of the normal vector loss function is:

[0113]

[0114] Among them, loss n represents the value of the normal vector loss function, represents the true normal vector of the i-th pixel point on the sample initial depth map, represents the predicted normal vector of the i-th pixel point on the sample initial depth map.

[0115] Among them, the process of determining the predicted normal vector is as follows:

[0116] Adjacent pixels of the pixel on the sample initial depth map are respectively selected in the x-axis and y-axis directions to obtain x-adjacent pixels and y-adjacent pixels.

[0117] Connect the x-adjacent pixels, the y-adjacent pixels and the pixel to obtain a triangular patch.

[0118] Determine the virtual camera coordinates of the x-adjacent pixels, the y-adjacent pixels and the pixel according to the reconstructed depth values of the x-adjacent pixels, the y-adjacent pixels and the pixel and the virtual camera parameters corresponding to the sample initial depth map.

[0119] Determine the direction vector of each side of the triangular patch according to the virtual camera coordinates of the x-adjacent pixels, the y-adjacent pixels and the pixel.

[0120] Select any two sides of the triangular patch, and perform a cross product of the direction vectors of the two sides to obtain the predicted normal vector of the pixel on the sample initial depth map.

[0121] The specific process of determining the above predicted normal vector is as follows:

[0122] Select two new pixels (u-1, v) and (u, v-1) around the pixel to form a triangular patch. According to the reconstructed depth values of the three pixels and the corresponding virtual camera internal parameter K, calculate the position coordinates of the three pixels (x-adjacent pixels, y-adjacent pixels and the pixel) in the virtual camera coordinate system, and then obtain the direction vector of each side of the triangular patch. The patch normal vector obtained by cross multiplying the direction vectors of any two sides of the patch is used as the normal vector of the current pixel (u, v).

[0123] Based on the above process, a trained surface encoding and decoding network can be obtained. The trained surface encoding and decoding network includes a trained surface encoder and a trained surface decoder. This trained surface decoder is used for the reconstruction of the S6 face.

[0124] Then, the surface reconstruction network and the trained surface decoder are trained. It should be noted that in this training process, the parameters in the trained surface decoder remain fixed. The present invention uses a depth estimation error function and an encoding integration error function to train the surface reconstruction network and the trained surface decoder. Among them, the expression of the depth estimation error function is:

[0125]

[0126] In the formula, loss dRepresents the value of the depth estimation error function F(c, i) represents the true depth value and the reconstructed depth value of the i-th pixel point respectively. F represents the decoding algorithm, c represents the sample surface prediction coding, and n is the number of pixel points on the initial depth map of the sample. The reconstructed depth value is predicted by the trained surface decoder mentioned above.

[0127] The coding integration error is In the formula, loss i Represents the value of the coding integration error. F represents the decoding algorithm, c1 and c2 represent the sample surface prediction codings of two sample patches respectively. When performing mesh segmentation, there is a certain sample mesh point located in two different sample patches after segmentation. Therefore, i1 and i2 represent the numbers of the current point under different sample patches.

[0128] After obtaining the reconstructed depth values of each pixel point on the initial depth map corresponding to the image block, high-precision sampling is performed on each patch and fused together to obtain the reconstruction result of the three-dimensional face.

[0129] The three-dimensional face reconstruction method based on deep learning provided by the present invention restores depth information based on the matching information between multi-view images through a deep learning model, improving the accuracy of face reconstruction. And the present invention performs depth reconstruction on each image block separately, that is, by reconstructing only the local depth of a fixed size each time, decoupling the number of network parameters from the image resolution, avoiding the problem that the number of network parameters increases rapidly with the image resolution, and being able to achieve high-precision face reconstruction with a small number of network parameters.

[0130] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above-mentioned three-dimensional face reconstruction method based on deep learning.

[0131] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by this application. As Figure 5As shown in the figure, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen and a keyboard. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As Figure 5 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0132] In Figure 5 the computer device 1000 shown in the figure, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the face three-dimensional reconstruction method based on deep learning described in the above embodiments, which will not be elaborated here.

[0133] The present invention also provides a computer-readable storage medium storing a computer program, which is suitable for being loaded and executed by a processor to implement the face three-dimensional reconstruction method based on deep learning described in the above embodiments, which will not be elaborated here.

[0134] The above program can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain network.

[0135] The above computer-readable storage medium may be an internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0136] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference may be made to the description in the method section.

[0137] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A 3D face reconstruction method based on deep learning, characterized in that, The method includes: S1: Obtain multiple images of the target face from different perspectives; S2: For each of the perspective images, using the perspective image as the source perspective image, input the source perspective image and the target perspective image into the trained coarse matching network model to obtain the predicted optical flow from the source perspective image to the target perspective image; the target perspective image is the perspective image other than the source perspective image among all the perspective images; the trained coarse matching network model is a model trained with sample source perspective images and sample target perspective images as inputs and the sample optical flow from the sample source perspective image to the sample target perspective image as labels; S3: Fuse all the perspective images according to the predicted optical flow and the true camera parameters corresponding to the perspective images to generate a rough face of the target face; S4: Segment the rough face into several image patches, generate virtual camera parameters corresponding to each image patch, and generate an initial depth map corresponding to the image patch through the virtual camera parameters corresponding to the image patch; S5: For each image patch, input all the perspective images and the initial depth map corresponding to the image patch into the trained surface reconstruction network to obtain the surface prediction encoding of the image patch; input the surface prediction encoding and the coordinates of each pixel point on the initial depth map corresponding to the image patch into the trained surface decoder to obtain the reconstructed depth value of each pixel point on the initial depth map corresponding to the image patch; the trained surface reconstruction network is a model trained with all the sample perspective images and sample initial depth maps of the sample face as inputs and the sample surface encoding as labels; the trained surface decoder is a model trained with the sample surface prediction encoding and sample point coordinates as inputs and the true depth value corresponding to the sample point as labels; S6: Determine the reconstructed face of the target face based on the reconstructed depth values of each pixel point on the initial depth maps corresponding to all the image patches.

2. The 3D face reconstruction method based on deep learning according to claim 1, wherein S3 specifically includes: Generate a true depth map corresponding to each perspective image according to the predicted optical flow and the true camera parameters corresponding to the perspective image; Fuse the true depth maps corresponding to all the perspective images to generate a rough face of the target face.

3. The 3D face reconstruction method based on deep learning according to claim 1, wherein The trained coarse matching network model includes an RGB feature extraction module and an optical flow prediction module connected in sequence; The RGB feature extraction module includes several convolutional layers connected in sequence and is used to extract features from the source perspective image and the target perspective image; The optical flow prediction module uses a U-Net network and is used to obtain the predicted optical flow from the source perspective image to the target perspective image according to the extracted features.

4. The 3D face reconstruction method based on deep learning according to claim 1, wherein The generation of the virtual camera parameters corresponding to each image patch specifically includes: For each image patch, perform the following steps: Process the image patch using the principal component analysis method to obtain three feature vectors; Sort the three feature vectors in descending order of eigenvalues, denote the feature vector ranked first as the first feature vector, the feature vector ranked second as the second feature vector, and the feature vector ranked third as the third feature vector; Use the first feature vector and the second feature vector as the x-axis and y-axis of the virtual camera respectively, and use the opposite direction of the third feature vector as the z-axis of the virtual camera to generate the virtual camera coordinate system of the virtual camera corresponding to the image block; Determine the true coordinates of the first feature vector, the second feature vector, and the third feature vector in the world coordinate system respectively; Determine the external parameter rotation matrix R according to the true coordinates; Determine the external parameter translation matrix T according to the external parameter rotation matrix R; Determine the virtual camera coordinates of each image point according to the coordinates of the image points on the image block, the external parameter rotation matrix R, and the external parameter translation matrix T; Determine the scaling coefficient s according to the maximum value in the x-axis direction and the maximum value in the y-axis direction among the virtual camera coordinates of all image points; Generate the external parameters of the virtual camera according to the external parameter rotation matrix R, the external parameter translation matrix T, and the scaling coefficient s; Determine the internal parameters of the virtual camera according to the resolution of the initial depth map corresponding to the image block; the external parameters of the virtual camera and the internal parameters of the virtual camera constitute the virtual camera parameters of the virtual camera.

5. The method for three-dimensional face reconstruction based on deep learning according to claim 1, wherein The trained surface reconstruction network includes a feature pyramid network, a feature cross-correlation module, and a surface encoding regression module connected in sequence; The feature pyramid network is used to extract features from each of the perspective images to obtain the features of each of the perspective images; The feature cross-correlation module is used to select several search points in the initial depth map corresponding to the image block; for each perspective image, project the coordinates of each search point into the image coordinate system corresponding to the perspective image based on the true camera parameters corresponding to the perspective image to obtain the projected coordinates of each search point in the image coordinate system corresponding to the perspective image, and calculate the perspective features corresponding to each search point in the perspective image based on the features of the perspective image and the projected coordinates; for each search point, perform pairwise cross-correlation calculations on the perspective features corresponding to the search point in all perspective images to obtain the cross-correlation calculation results of each search point; fuse the cross-correlation calculation results of all search points to obtain the depth direction cost volume; The surface encoding regression module is used to encode the features of each perspective image, the depth direction cost volume, and the initial depth map corresponding to the image block to obtain the surface prediction encoding of the image block.

6. The method for three-dimensional face reconstruction based on deep learning according to claim 1, wherein Before S5, it further includes: training the surface decoder, and the training process is as follows: Obtain a first sample set; the first sample set includes the sample initial depth map of the sample human face, the sample point coordinates, and the true depth value corresponding to the sample point; The surface encoding and decoding network is trained using the first sample set to obtain a trained surface encoding and decoding network; the trained surface encoding and decoding network includes a trained surface encoder and a trained surface decoder connected in sequence.

7. The method for three-dimensional face reconstruction based on deep learning according to claim 6, wherein The loss function used in the training process of the surface encoding and decoding network includes a depth loss function and a normal vector loss function; The expression of the depth loss function is: Among them, loss d represents the value of the depth loss function, n represents the number of pixel points on the initial depth map of the sample, represents the true depth value of the i-th pixel point on the initial depth map of the sample; represents the reconstructed depth value of the i-th pixel point on the initial depth map of the sample; The expression of the normal vector loss function is: Among them, loss n represents the value of the normal vector loss function, represents the true normal vector of the i-th pixel on the initial depth map of the sample, represents the predicted normal vector of the i-th pixel on the initial depth map of the sample.

8. The 3D face reconstruction method based on deep learning according to claim 7, characterized in that, The determination process of the predicted normal vector is as follows: Adjacent pixels of the pixel points on the sample initial depth map are respectively selected in the x-axis and y-axis directions to obtain x adjacent pixels and y adjacent pixels; The x adjacent pixels, the y adjacent pixels and the pixel points are connected to obtain a triangular patch; The virtual camera coordinates of the x adjacent pixels, the y adjacent pixels and the pixel points are determined according to the reconstructed depth values of the x adjacent pixels, the y adjacent pixels and the pixel points and the virtual camera parameters corresponding to the sample initial depth map; The direction vectors of each side of the triangular patch are determined according to the virtual camera coordinates of the x adjacent pixels, the y adjacent pixels and the pixel points; Any two sides of the triangular patch are selected, and the direction vectors of the two sides are cross-multiplied to obtain the predicted normal vector of the pixel points on the sample initial depth map.

9. A computer device, characterized in that, It includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by the processor to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Three-dimensional face model reconstruction method and system based on self-supervised learning

    CN112950775A

  • Scene-adaptive fine three-dimensional face reconstruction method and system, and electronic equipment

    CN113269862A