A three-dimensional face reconstruction method, device, electronic equipment and storage medium
By acquiring images from different perspectives and recognizing facial features in 3D face reconstruction, and then filtering and adjusting the images, the problems of unattractive 3D face meshes and insufficient detail were solved, achieving fast and high-quality 3D face reconstruction.
Patent Information
- Application Number
- CN202211165258.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-09-23
AI Technical Summary
Existing 3D face reconstruction technologies suffer from problems such as inconsistent facial expressions and angles in multiple images, resulting in aesthetically unappealing 3D face meshes and issues like strabismus. Furthermore, they lack sufficient reconstruction speed and detail.
By acquiring depth and texture images from different perspectives, facial feature information such as head pose, hair region, and facial key points is identified. Based on the recognition results, the images are filtered and adjusted, and then meshed and textured to generate a 3D face.
It achieves fast 3D face reconstruction, generates 3D face meshes with high integrity and rich details, and has good visualization effects, meeting the needs of high-efficiency and high-quality reconstruction.
Smart Images

Figure CN116091687B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, and particularly relates to a three-dimensional face reconstruction method and device, electronic equipment and storage medium. BACKGROUND
[0002] Generally, in the field of professional orthodontics and medical cosmetology, accurate and beautiful three-dimensional facial data needs to be obtained through a facial scanner and provided to doctors as a basis for diagnosis and treatment.
[0003] The three-dimensional face reconstruction technology in the related art is usually based on the fusion of multiple face data collected by a camera, and the details are rich, but due to the differences in facial expressions and facial gaze angles in multiple images, the three-dimensional face grid generated finally has problems such as not being beautiful enough and eyes being slanted. SUMMARY
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a three-dimensional face reconstruction method, device, electronic equipment and storage medium.
[0005] The present disclosure provides a three-dimensional face reconstruction method, which comprises the following steps:
[0006] Obtaining a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles;
[0007] Identifying facial feature information of the to-be-processed texture image to obtain an identification result;
[0008] Processing the to-be-processed depth image based on the identification result to obtain a target depth image set;
[0009] Performing grid processing based on the target depth image set to obtain a face grid model;
[0010] Performing texture mapping processing on the face grid model based on the identification result and a target texture image set determined from the to-be-processed texture image to obtain a three-dimensional face.
[0011] The present disclosure also provides a three-dimensional face reconstruction device, which comprises the following modules:
[0012] An acquisition module is configured to acquire a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles;
[0013] An identification module is configured to identify facial feature information of the to-be-processed texture image to obtain an identification result;
[0014] A first processing module is configured to process the to-be-processed depth image based on the identification result to obtain a target depth image set;
[0015] a second processing module, configured to perform grid processing based on the target depth image set to obtain a face grid model;
[0016] a third processing module, configured to perform texture mapping processing on the face grid model based on the target texture image set determined from the to-be-processed texture image according to the recognition result to obtain a three-dimensional face.
[0017] The embodiments of the present disclosure further provide an electronic device, which comprises a processor, a memory for storing executable instructions of the processor, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the three-dimensional face reconstruction method provided by the embodiments of the present disclosure.
[0018] The embodiments of the present disclosure further provide a computer-readable storage medium, which stores a computer program for executing the three-dimensional face reconstruction method provided by the embodiments of the present disclosure.
[0019] The technical scheme provided by the embodiments of the present disclosure has the following advantages compared with the prior art: the three-dimensional face reconstruction scheme provided by the embodiments of the present disclosure obtains to-be-processed depth images and to-be-processed texture images of a target face under different perspectives, performs face feature information recognition on the to-be-processed texture images to obtain a recognition result, processes the to-be-processed depth images based on the recognition result to obtain a target depth image set, performs grid processing based on the target depth image set to obtain a face grid model, and performs texture mapping processing on the face grid model based on the target texture image set determined from the to-be-processed texture images according to the recognition result to obtain a three-dimensional face. By using the above technical scheme, the three-dimensional face reconstruction speed is fast, the reconstructed three-dimensional face has high grid integrity, rich details, and good visualization effect, and the further demand of users for efficient and high-quality three-dimensional face reconstruction is met. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0021] Figure 1 A flowchart of a three-dimensional face reconstruction method provided by the embodiments of the present disclosure is shown in the figure;
[0022] Figure 2 A flowchart of another three-dimensional face reconstruction method provided by the embodiments of the present disclosure is shown in the figure;
[0023] Figure 3 An example diagram of a hair region provided for an embodiment of the present disclosure;
[0024] Figure 4 An example diagram of a target map region provided for an embodiment of the present disclosure;
[0025] Figure 5 A structural schematic diagram of a three-dimensional face reconstruction device provided for an embodiment of the present disclosure;
[0026] Figure 6 A structural schematic diagram of an electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term “comprising” and variations thereof as used in the present disclosure are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions are given below in the description of the embodiments.
[0030] It should be noted that the terms “first”, “second”, and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the adjectives “one”, “more than one” mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as “one or more”.
[0032] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0033] In practical applications, in the way of using camera to collect multiple face data and then fusing to reconstruct the face, some background areas will have reconstructed garbage data, which will affect the splicing speed and effect. In the scanning process, the person cannot keep the rigid body state, and there will be situations of closing eyes and staring at other areas, resulting in the final generated face mesh having closed eyes and oblique eyes. The directly scanned face is affected by light and other factors, the skin is dark and rough, the visualization effect of the final three-dimensional face mesh is poor, and when extracting the mesh, the point cloud generated in the hair area is sparse, and it is difficult to form a mesh with good quality. Or the three-dimensional face reconstruction method based on deep learning, which is usually based on a general template of a variable face model and is reconstructed by parameterization. The reconstruction speed is fast, but it lacks real detailed information, that is, the geometric details of the face are lost.
[0034] To solve the above problems, the embodiment of the present disclosure proposes a three-dimensional face reconstruction method, which assists the reconstruction of a three-dimensional face mesh by recognizing multiple face feature information. The reconstruction speed is fast, the reconstructed three-dimensional face mesh has high completeness, rich details, and good visualization effect.
[0035] Figure 1 A flowchart of a three-dimensional face reconstruction method provided by the embodiment of the present disclosure is shown. The method can be executed by a three-dimensional face reconstruction device, which can be implemented by software and / or hardware and can be integrated in an electronic device. As shown in Figure 1 The method comprises the following steps.
[0036] Step 101, acquiring a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles.
[0037] The target face is a face to be three-dimensionally reconstructed, which can be selected and set according to the needs of the application scenario. Different viewing angles refer to different angles of the target face. In the embodiment of the present disclosure, the target face is photographed under different viewing angles, and the to-be-processed depth image and the to-be-processed texture image are acquired synchronously, that is, one to-be-processed depth image and one to-be-processed texture image of the target face can be acquired under each viewing angle. The to-be-processed depth image refers to a to-be-processed image in which the distance (depth) value of each point in the scene collected by the camera is taken as the pixel value; the to-be-processed texture image refers to a to-be-processed image reflecting the visual features of homogeneity in the image.
[0038] In the embodiments of the present disclosure, there are many ways to obtain the to-be-processed depth images and the to-be-processed texture images of the target face under different perspectives. In some embodiments, the camera is controlled to remain stationary, the target face is controlled to rotate according to a preset angle and direction, and the camera is controlled to capture the target face to obtain the to-be-processed depth images and the to-be-processed texture images of the target face under different perspectives. In other embodiments, the target face is controlled to remain stationary, the camera is controlled to rotate according to a preset angle and direction, and the target face is captured to obtain the to-be-processed depth images and the to-be-processed texture images of the target face under different perspectives.
[0039] The above two ways are only examples of obtaining the to-be-processed depth images and the to-be-processed texture images of the target face under different perspectives, and the embodiments of the present disclosure do not limit the specific implementation of obtaining the to-be-processed depth images and the to-be-processed texture images of the target face under different perspectives.
[0040] Step 102, performing face feature information recognition on the to-be-processed texture image to obtain a recognition result.
[0041] The face feature information includes hair region, whether to blink, gaze direction of the eyes, head pose, and the like.
[0042] In the embodiments of the present disclosure, there are many ways to perform face feature information recognition on the to-be-processed texture image to obtain a recognition result. In some embodiments, head pose recognition is performed on the to-be-processed texture image to obtain a head pose, hair region recognition is performed on the to-be-processed texture image to obtain a hair region, and face key point recognition is performed on the to-be-processed texture image to obtain a face key point.
[0043] In other embodiments, head pose recognition is performed on the to-be-processed texture image to obtain a head pose. The above two ways are only examples, and the embodiments of the present disclosure do not specifically limit the way of performing face feature information recognition on the to-be-processed texture image to obtain a recognition result.
[0044] Step 103, performing processing on the to-be-processed depth image based on the recognition result to obtain a target depth image set.
[0045] In the embodiments of the present disclosure, the target depth image refers to a depth image selected from a plurality of to-be-processed depth images based on the recognition result.
[0046] Specifically, there are many ways to process the to-be-processed depth image based on the recognition result to obtain the target depth image. In some embodiments, a first weight value of the to-be-processed texture image is determined based on the head pose, a second weight value of the to-be-processed texture image is determined based on the face key point, a first total weight value of the to-be-processed texture image is determined based on the first weight value and the second weight value, the corresponding to-be-processed depth image is filtered based on the first total weight value of the to-be-processed texture image, and a target depth image set is obtained.
[0047] In other embodiments, a target region of the to-be-processed depth image is determined based on the hair region, and a depth map reconstruction parameter of the target region is adjusted to obtain a target depth image set. The above two ways are only examples of processing the to-be-processed depth image based on the recognition result to obtain the target depth image, and the embodiments of the present disclosure do not specifically limit the way of processing the to-be-processed depth image based on the recognition result to obtain the target depth image.
[0048] Step 104: performing meshing processing based on the target depth image set to obtain a face mesh model.
[0049] Specifically, after the target depth image set is determined, the target depth image set is processed by triangular meshing or the like to obtain the face mesh model.
[0050] Step 105: performing texture mapping processing on the face mesh model based on the recognition result from the target texture image set determined from the to-be-processed texture image to obtain a three-dimensional face.
[0051] In the embodiments of the present disclosure, the target texture image refers to a texture image selected from a plurality of to-be-processed texture images based on the recognition result.
[0052] Specifically, there are many ways to perform texture mapping processing on the face mesh model based on the recognition result from the target texture image set determined from the to-be-processed texture image to obtain the three-dimensional face. In some embodiments, a third weight value of the to-be-processed texture image is determined based on the head pose, a fourth weight value of the to-be-processed texture image is determined based on the face key point, a second total weight value of the to-be-processed texture image is determined based on the third weight value and the fourth weight value, and the texture mapping processing is performed on the face mesh model based on the second total weight value from the target texture image set determined from the to-be-processed texture image to obtain the three-dimensional face.
[0053] In some embodiments, the target texture image set is determined based on the head region to perform texture mapping on the face mesh model, and a three-dimensional face is obtained. The above two manners are only examples of determining the target texture image set from the to-be-processed texture image based on the recognition result to perform texture mapping on the face mesh model, and obtaining a three-dimensional face. The present embodiment does not specifically limit the manner of determining the target texture image set from the to-be-processed texture image based on the recognition result to perform texture mapping on the face mesh model, and obtaining a three-dimensional face.
[0054] The three-dimensional face reconstruction scheme provided by the present embodiment comprises the following steps: obtaining a to-be-processed depth image and a to-be-processed texture image of a target face under different perspectives; performing face feature information recognition on the to-be-processed texture image to obtain a recognition result; processing the to-be-processed depth image based on the recognition result to obtain a target depth image set; performing meshing processing based on the target depth image set to obtain a face mesh model; and determining a target texture image set from the to-be-processed texture image based on the recognition result to perform texture mapping on the face mesh model, and obtaining a three-dimensional face. The face feature information recognition result based on the texture image is used to assist the meshing processing and the texture mapping, so that the three-dimensional face reconstruction is fast, the three-dimensional face mesh has high completeness, rich details, and good visualization effect, and the user's demand for efficient and high-quality three-dimensional face reconstruction is further met.
[0055] Figure 2 For another flowchart of the three-dimensional face reconstruction method provided by the present embodiment, the present embodiment further optimizes the above three-dimensional face reconstruction method on the basis of the above embodiment. As shown in Figure 2 the method comprises the following steps:
[0056] Step 201: obtaining a to-be-processed depth image and a to-be-processed texture image of a target face under different perspectives.
[0057] It should be noted that step 201 is the same as step 101, which will not be described in detail here. For details, please refer to the detailed description of step 101.
[0058] Step 202: performing head posture recognition on the to-be-processed texture image to determine the head posture in the to-be-processed texture image, performing hair region recognition on the to-be-processed texture image to determine the hair region in the to-be-processed texture image, and performing face key point recognition on the to-be-processed texture image to determine the face key point in the to-be-processed texture image.
[0059] Specifically, in the embodiments of the present disclosure, the deep learning algorithm or deep learning model can be used to identify the texture image to be processed to determine the head posture, hair region, and face key points, etc. in the texture image to be processed. The face key points refer to the eyes, mouth, etc. The state information corresponding to the face key points can be obtained, such as whether the eyes are open or closed, the eye gaze direction, etc. The head posture refers to the angle of the head, such as the front face, back face, and tilt angle, etc.
[0060] Specifically, the recognition model can be pre-trained to identify the texture image to be processed, that is, the recognition result is directly outputted by inputting the texture image to be processed, thereby further improving the processing efficiency.
[0061] Specifically, the identification of the texture image to be processed can be understood as the identification of each pixel point in the texture image to be processed, such as determining the hair region based on the color information of the pixel point.
[0062] Step 203, determining a first weight value of the texture image to be processed based on the head posture, determining a second weight value of the texture image to be processed based on the face key points, and determining a first total weight value of the texture image to be processed based on the first weight value and the second weight value.
[0063] Step 204, filtering the corresponding depth image to be processed based on the first total weight value of the texture image to be processed to obtain a target depth image set.
[0064] Specifically, when the grid is extracted, the face feature information of the collected texture image to be processed is identified, and the weight value is calculated based on whether the eyes are closed, the head posture, etc. The redundant pictures are removed to speed up the grid extraction. The first weight value and the second weight value are selected and set according to the application scenario, and the first total weight value can be obtained by adding or weighted averaging the first weight value and the second weight value, etc.
[0065] Specifically, different weight values are given based on different identification results, thereby obtaining the first total weight value of each texture image to be processed. The first total weight values are sorted, and the target depth image corresponding to the texture image to be processed with a higher first total weight value is obtained based on the sorting result to perform grid processing, thereby improving the grid extraction speed.
[0066] It should be noted that in the embodiments of the present disclosure, filtering the corresponding depth image to be processed to obtain a target depth image set can be understood as marking part of the whole depth image to be processed as not participating in subsequent grid processing, or adding a mask to part of the region in the depth image to be processed so that the part of the region does not participate in subsequent grid processing, such as closing the eyes in the eye region and adding a mask to the eye region to not participate in subsequent grid processing, thereby further meeting the diversified needs of three-dimensional reconstruction.
[0067] Step 205, determining a target region of the to-be-processed depth image based on the hair region, adjusting the depth map reconstruction parameter of the target region, and obtaining a target depth image set.
[0068] Specifically, the hair region is identified through the collected to-be-processed texture image, and a mask of the hair region is generated (as shown in Figure 3 According to the mask of the hair region and the to-be-processed depth image, the hair region of the to-be-processed depth image can be correspondingly determined as the target region.
[0069] By adjusting the depth map reconstruction parameter of the hair region as the target region, the reconstruction of the hair region is assisted, and the integrity of the face mesh reconstruction is improved. The point cloud of the hair region is relatively sparse, so the parameter (point distance or constraint) can be adjusted to increase the integrity of the face mesh reconstruction.
[0070] At this time, the target depth image is the to-be-processed depth image including the target region; and the depth map reconstruction parameter of the target region has been adjusted.
[0071] Specifically, the hair region is identified, and the parameter of the hair region is adjusted to assist the subsequent reconstruction of the hair region, that is, to assist the reconstruction of the sparse region (such as hair), and to further improve the reconstruction integrity of the hair region.
[0072] After step 202, steps 203-204 and / or step 205 can be performed, and the specific execution order Figure 2 is only an example.
[0073] Step 206, determining a third weight value of the to-be-processed texture image based on the head pose, determining a fourth weight value of the to-be-processed texture image based on the face key point, and determining a second total weight value of the to-be-processed texture image based on the third weight value and the fourth weight value.
[0074] Step 207, determining a target texture image set from the to-be-processed texture image based on the second total weight value, determining a beauty region of the target texture image set, performing beauty processing on the beauty region, and obtaining a final texture image.
[0075] Specifically, when the texture is mapped, the face feature information of the collected to-be-processed texture image is identified, and the weight value is calculated according to whether the eyes are closed or the head pose, so as to remove redundant pictures and speed up the texture mapping. The third weight value and the fourth weight value are selected and set according to the application scene, and the second total weight value can be obtained by adding or weighted averaging the third weight value and the fourth weight value.
[0076] Specifically, different weight values are assigned based on different recognition results, thereby obtaining a second total weight value for each texture image to be processed. The sorting results based on the second total weight values are used to select the texture images with higher second total weight values as the target texture image set for texture mapping, thereby improving the texture mapping speed.
[0077] The beautification area can be selected and set as needed, and can be the entire target texture image or a portion of the image such as teeth.
[0078] Specifically, by performing beautification, whitening, or teeth whitening operations on the target texture image set, a 3D face mesh with good visualization and effect can be reconstructed. That is, by beautifying the texture image, a 3D face mesh with good visualization and effect can be reconstructed, thereby improving the visualization effect of the 3D face mesh.
[0079] Step 208: Identify the target texture region and other texture regions in the face mesh model, determine the texture image from the final texture image based on the correlation of face key points, perform texture processing on the target texture region based on the texture image, and perform texture processing on other texture regions based on the remaining target texture image to obtain a three-dimensional face.
[0080] Specifically, after obtaining the final texture image, the target texture region and other texture regions in the face mesh model can be identified. Based on the correlation of facial key points, the texture image is determined from the final texture image. The target texture region is processed by the texture image, and other texture regions are processed by the remaining target texture image to obtain a three-dimensional face.
[0081] Among them, the correlation of facial key points refers to the relative relationship between eyes, the relative relationship between eyes and nose, mouth, and nose and mouth; the relative relationship includes distance and direction.
[0082] For example, such as Figure 4 As shown, the region including the eyes, nose, and mouth is designated as the region of interest (ROI) A. A texture image is used to apply texture mapping to this target region. Then, other regions are mapped using the remaining target texture image to obtain a 3D face. This involves applying texture mapping to the eyes, nose, and mouth using the same texture image. The texturing process primarily involves adding color to the face model.
[0083] Specifically, by identifying the facial features and other regions of interest (target texture regions) of the target texture image that is a frontal face, the texture image is made to ensure that the regions of interest come from the same frontal face texture image (i.e., the texture image), thus avoiding texture problems such as closed eyes or inconsistent gaze between the left and right eyes.
[0084] The three-dimensional face reconstruction scheme provided by the embodiments of the present disclosure includes obtaining a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles, performing head posture recognition on the to-be-processed texture image, determining the head posture in the to-be-processed texture image, performing hair region recognition on the to-be-processed texture image, determining the hair region in the to-be-processed texture image, performing face key point recognition on the to-be-processed texture image, determining the face key point in the to-be-processed texture image, determining a first weight value of the to-be-processed texture image based on the head posture, determining a second weight value of the to-be-processed texture image based on the face key point, determining a first total weight value of the to-be-processed texture image based on the first weight value and the second weight value, filtering the corresponding to-be-processed depth image based on the first total weight value of the to-be-processed texture image to obtain a target depth image set, determining a target region of the to-be-processed depth image based on the hair region, adjusting the depth image reconstruction parameter of the target region to obtain a target depth image set, determining a third weight value of the to-be-processed texture image based on the head posture, determining a fourth weight value of the to-be-processed texture image based on the face key point, determining a second total weight value of the to-be-processed texture image based on the third weight value and the fourth weight value, determining a target texture image set from the to-be-processed texture image based on the second total weight value, determining a beautifying region of the target texture image set, performing beautifying processing on the beautifying region to obtain a final texture image, recognizing a target mapping region in a face mesh model, determining a mapping texture image from the final texture image based on the face key point correlation, performing mapping processing on the target mapping region based on the mapping texture image, and obtaining a three-dimensional face. By using the technical scheme, the face hair region is recognized, the related parameters are configured, the reconstruction completeness of the hair region is improved, the face feature information is used to obtain the closed eye and head posture estimation, the mesh extraction efficiency is improved, the face feature information is used to obtain the closed eye and head posture estimation, the texture mapping efficiency is improved, the texture mapping effect is improved through beautifying, and the texture mapping effect is improved by recognizing the front face region of interest, so that the reconstructed mesh is relatively complete (the hair region), the details are rich, the visual effect of the reconstructed three-dimensional face mesh is good, and there is no closed eye, squinting and the like.
[0085] In some embodiments, it is determined whether the head posture in the corresponding to-be-processed texture image meets the posture requirement for the to-be-processed depth image under different viewing angles, if yes, it is added to the target depth image set, and if not, it is not added. It is determined whether the eyes in the corresponding to-be-processed texture image are closed and whether the eye gaze angle meets the preset angle condition for the to-be-processed depth image under different viewing angles, if the eyes are closed and / or the eye gaze angle does not meet the preset angle condition, it is not added or partially added to the target depth image set (which means that it is added after the eye region is processed by adding a mask), and if not, it is added to the target depth image set. In this way, the efficiency and accuracy of subsequent mesh reconstruction and texture mapping are further improved.
[0086] Figure 5 A structural schematic diagram of a three-dimensional face reconstruction device is provided for an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in the figure, the device comprises: Figure 5
[0087] An acquisition module 301 is configured to acquire a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles;
[0088] An identification module 302 is configured to perform face feature information identification on the to-be-processed texture image to obtain an identification result;
[0089] A first processing module 303 is configured to perform processing on the to-be-processed depth image based on the identification result to obtain a target depth image set;
[0090] A second processing module 304 is configured to perform meshing processing based on the target depth image set to obtain a face mesh model;
[0091] A third processing module 305 is configured to perform texture mapping processing on the face mesh model based on the identification result and a target texture image set determined from the to-be-processed texture image to obtain a three-dimensional face.
[0092] Optionally, the identification module 302 is specifically configured to:
[0093] perform head posture identification on the to-be-processed texture image to determine a head posture in the to-be-processed texture image; and / or,
[0094] perform hair region identification on the to-be-processed texture image to determine a hair region in the to-be-processed texture image; and / or,
[0095] perform face key point identification on the to-be-processed texture image to determine a face key point in the to-be-processed texture image.
[0096] Optionally, the first processing module 303 is specifically configured to:
[0097] determine a first weight value of the to-be-processed texture image based on the head posture;
[0098] determine a second weight value of the to-be-processed texture image based on the face key point;
[0099] determine a first total weight value of the to-be-processed texture image based on the first weight value and the second weight value;
[0100] filter the corresponding to-be-processed depth image based on the first total weight value of the to-be-processed texture image to obtain the target depth image set.
[0101] Optionally, the first processing module 303 is specifically configured to:
[0102] determining whether the head pose in the corresponding texture image meets the pose requirement, if yes, adding to the target depth image set, if not, not adding;
[0103] determining whether the eyes in the corresponding texture image are closed and whether the eye gaze angle meets the preset angle condition, if the eyes are closed and / or the eye gaze angle does not meet the preset angle condition, not adding or adding to the target depth image set after adding a mask to the eye region, if not, adding to the target depth image set;
[0104] determining whether the mouth opening angle in the corresponding texture image is greater than a preset angle threshold, if yes, not adding or adding to the target depth image set after adding a mask to the mouth region, if not, adding to the target depth image set.
[0105] Optionally, the first processing module 303 is specifically configured to:
[0106] determining a target region of the depth image based on the hair region;
[0107] adjusting the depth image reconstruction parameter of the target region to obtain the target depth image set.
[0108] Optionally, the third processing module 305 is specifically configured to:
[0109] determining a third weight value of the texture image based on the head pose;
[0110] determining a fourth weight value of the texture image based on the face key point;
[0111] determining a second total weight value of the texture image based on the third weight value and the fourth weight value;
[0112] determining a target texture image set from the texture image based on the second total weight value, and performing texture mapping on the face mesh model to obtain a three-dimensional face.
[0113] Optionally, the device further comprises:
[0114] a first determination module configured to determine a beauty region of the target texture image;
[0115] a beauty processing module configured to perform beauty processing on the beauty region to obtain a final texture image.
[0116] Optionally, the apparatus further comprises:
[0117] The identification module is configured to identify a target mapping region and other mapping regions in the face mesh model.
[0118] The second determination module is configured to determine a mapping texture image from the target texture image based on a face key point correlation.
[0119] The mapping processing module is configured to perform mapping processing on the target mapping region based on the mapping texture image, and perform mapping processing on the other mapping regions based on the remaining target texture image.
[0120] The three-dimensional face reconstruction apparatus provided by the embodiments of the present disclosure can perform the three-dimensional face reconstruction method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.
[0121] The embodiments of the present disclosure further provide a computer program product, which comprises computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the three-dimensional face reconstruction method provided by any of the embodiments of the present disclosure.
[0122] Figure 6 A structural schematic diagram of an electronic device provided by the embodiments of the present disclosure is provided. The following specifically refers to Figure 6 which shows a structural schematic diagram of an electronic device 400 suitable for being used to implement the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure can include but is not limited to a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (for example, a vehicle navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 6 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0123] As shown in Figure 6 , the electronic device 400 can include a processing apparatus (for example, a central processing unit, a graphic processing unit, and the like) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage apparatus 408 to a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing apparatus 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0124] In general, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 408 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data. Although Figure 6 The electronic device 400 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.
[0125] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing devices 401, the above-mentioned functions defined in the three-dimensional face reconstruction method of embodiments of the present disclosure are performed.
[0126] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF, etc., or any suitable combination of the above.
[0127] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0128] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.
[0129] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a to-be-processed depth image and a to-be-processed texture image of a target face under different viewing angles, perform face feature information recognition on the to-be-processed texture image to obtain a recognition result, perform processing on the to-be-processed depth image based on the recognition result to obtain a target depth image set, perform grid processing based on the target depth image set to obtain a face grid model, perform mapping processing on the face grid model based on the recognition result from the to-be-processed texture image to obtain a three-dimensional face.
[0130] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0131] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0132] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0133] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- transitory machine-readable media can include RAM, ROM, programmable ROM (EPROM, EEPROM or flash memory), or any other storage device(s) through which program instructions can be stored and executed by a processing unit. The above described functions can be implemented as software modules or software functions using object-oriented design methodology, or using any other suitable programming technique.
[0134] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0135] The above description is only preferred embodiments of the present disclosure and the explanation of the technical principles of the application. It should be understood by those skilled in the art that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0136] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0137] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A three-dimensional face reconstruction method, characterized in that, include: Obtain the depth image and texture image of the target face from different viewpoints; The texture image to be processed is subjected to facial feature information recognition to obtain the recognition result; wherein, the facial feature information includes hair region, facial key points and head posture; The depth image to be processed is processed based on the recognition result to obtain a target depth image set; wherein, the process of processing the depth image to be processed based on the recognition result to obtain the target depth image set includes: determining a first weight value of the texture image to be processed based on the head pose; determining a second weight value of the texture image to be processed based on the facial key points; determining a first total weight value of the texture image to be processed based on the first weight value and the second weight value; and filtering the corresponding depth images to be processed based on the first total weight value of the texture image to be processed to obtain the target depth image set. A face mesh model is obtained by performing meshing processing on the target depth image set; Based on the recognition results, a set of target texture images is determined from the texture image to be processed, and the face mesh model is processed by mapping to obtain a three-dimensional face.
2. The three-dimensional face reconstruction method according to claim 1, characterized in that, The process of recognizing facial features in the texture image to be processed to obtain the recognition result includes: Perform head pose recognition on the texture image to be processed to determine the head pose in the texture image; and / or, Perform hair region recognition on the texture image to be processed to determine the hair regions in the texture image to be processed; and / or, Facial key points are identified in the texture image to be processed.
3. The three-dimensional face reconstruction method according to claim 1, characterized in that, The step of processing the depth image to be processed based on the recognition result to obtain a target depth image set further includes: For depth images to be processed from different perspectives, determine whether the head pose in the corresponding texture image to be processed meets the pose requirements. If yes, add it to the target depth image set; otherwise, do not add it. For the depth image to be processed from different perspectives, determine whether the eyes in the corresponding texture image to be processed are closed and whether the eye gaze angle meets the preset angle conditions. If the eyes are closed and / or the eye gaze angle does not meet the preset angle conditions, do not add or add the eye area to the target depth image set after adding a mask. Otherwise, add the image to the target depth image set.
4. The three-dimensional face reconstruction method according to claim 1, characterized in that, The step of processing the depth image to be processed based on the recognition result to obtain a target depth image set further includes: The target region of the depth image to be processed is determined based on the hair region. Adjust the depth map reconstruction parameters of the target region to obtain the target depth image set.
5. The three-dimensional face reconstruction method according to claim 1, characterized in that, The step of determining a target texture image set from the texture image to be processed based on the recognition result and performing texture mapping on the face mesh model to obtain a three-dimensional face includes: The third weight value of the texture image to be processed is determined based on the head pose; The fourth weight value of the texture image to be processed is determined based on the facial key points; The second total weight value of the texture image to be processed is determined based on the third weight value and the fourth weight value; Based on the second total weight value, a set of target texture images is determined from the texture image to be processed, and the face mesh model is processed by mapping to obtain a three-dimensional face.
6. The three-dimensional face reconstruction method according to any one of claims 1-5, characterized in that, Before performing texture mapping on the face mesh model to obtain a 3D face by determining the target texture image set from the texture image to be processed based on the recognition result, the method further includes: Determine the beautification area of the target texture image; The beautification area is then processed to obtain the final texture image.
7. The three-dimensional face reconstruction method according to any one of claims 1-5, characterized in that, The process of determining the target texture image set from the texture image to be processed based on the recognition result and performing texture mapping on the face mesh model further includes: Identify the target texture region and other texture regions in the face mesh model; The texture image is determined from the target texture image based on the correlation of facial key points; The target texture region is processed by applying the texture image, and the other texture regions are processed by applying the remaining target texture image.
8. A three-dimensional face reconstruction device, characterized in that, include: The acquisition module is used to acquire the depth image and texture image of the target face from different perspectives. The recognition module is used to recognize facial feature information in the texture image to be processed and obtain the recognition result; wherein, the facial feature information includes hair area, facial key points and head posture; A first processing module is configured to process the depth image to be processed based on the recognition result to obtain a target depth image set; wherein, the process of processing the depth image to be processed based on the recognition result to obtain the target depth image set includes: determining a first weight value of the texture image to be processed based on the head pose; determining a second weight value of the texture image to be processed based on the facial key points; determining a first total weight value of the texture image to be processed based on the first weight value and the second weight value; and filtering the corresponding depth images to be processed based on the first total weight value of the texture image to be processed to obtain the target depth image set. The second processing module is used to perform meshing processing on the target depth image set to obtain a face mesh model; The third processing module is used to determine the target texture image set from the texture image to be processed based on the recognition result, and to perform texture mapping on the face mesh model to obtain a three-dimensional face.
9. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the three-dimensional face reconstruction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the three-dimensional face reconstruction method according to any one of claims 1-7.
Citation Information
Patent Citations
Eye texture image generation method and device, texture mapping method and device and electronic equipment
CN111862287A
Hair reconstruction method and equipment
CN113570701A
Three-dimensional face model reconstruction method and device, electronic equipment and storage medium
CN113902849A