Three-dimensional face reconstruction method and system based on occlusion segmentation

CN115619933BActive Publication Date: 2026-09-15BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211286327.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2026-09-15
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种基于遮挡分割的三维人脸重建方法及系统,能够减少三维人脸重建应用场景中模型部署对计算资源的占用,压缩模型计算量,解决三维人脸重建应用场景中计算资源占用过多的技术问题

Benefits of technology

本申请实施例通过将目标人脸图像输入预构建的参数预测模型,该参数预测模型包括图像特征提取器和图像分割解码器,参数预测模型基于多个人脸训练图片、人脸训练图片的人脸关键点信息和人脸遮挡分割区域进行训练,直至图像特征提取器和图像分割解码器的关联损失函数达到设定状态;进而基于参数预测模型输出目标人脸图像的目标人脸重建参数和目标人脸遮挡区域,并基于人脸重建参数和人脸遮挡分割区域进行三维人脸重建后处理。采用上述技术手段,通过训练包含图像特征提取器和图像分割解码器的参数预测模型,直至图像特征提取器和图像分割解码器的关联损失函数达到设定状态,以此可以使参数预测模型融合三维人脸重建和人脸遮挡分割功能,减少模型部署对计算资源的占用,减少模型的冗余度,压缩模型计算量,提升三维人脸重建效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115619933B_ABST
    Figure CN115619933B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a three-dimensional face reconstruction method and system based on occlusion segmentation. The technical scheme provided by the embodiment of the application inputs a target face image into a pre-constructed parameter prediction model, the parameter prediction model includes an image feature extractor and an image segmentation decoder, the parameter prediction model is trained based on multiple face training pictures, face key point information of the face training pictures and face occlusion segmentation areas, until the correlation loss function of the image feature extractor and the image segmentation decoder reaches a set state; then the target face reconstruction parameters and the target face occlusion area of the target face image are output based on the parameter prediction model, and three-dimensional face reconstruction post-processing is performed based on the face reconstruction parameters and the face occlusion segmentation area. By using the above technical means, the occupation of the calculation resources by the model deployment can be reduced, the redundancy of the model can be reduced, the model calculation amount can be compressed, and the three-dimensional face reconstruction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method and system for three-dimensional face reconstruction based on occlusion segmentation. Background Technology

[0002] Currently, 3D face reconstruction technology is widely used in film, gaming, healthcare, and live-streaming social media. For example, in live-streaming social media, by acquiring a user's 2D face image, 3D face reconstruction technology can reconstruct the user's 3D (expression, texture) information, thereby enabling functions such as 3D face reshaping and 3D makeup. However, in real-world applications, 2D face images do not always contain the complete face as expected; there may be occlusions from limbs or objects. Therefore, face occlusion segmentation is necessary to locate the occluded areas. Post-processing is then performed based on the 3D face reconstruction results and the face occlusion segmentation results to ensure the effectiveness of 3D face reshaping and 3D makeup functions.

[0003] However, in existing 3D face reconstruction applications, the 3D face reconstruction model and the face occlusion segmentation model are deployed independently, treating 3D face reconstruction and occlusion region segmentation as two separate tasks. Due to the limited computing power of the deployment platform, deploying multiple models simultaneously consumes excessive computing resources, increases the platform's computational pressure, and impacts the operation of the platform's computing services. Summary of the Invention

[0004] This application provides a method and system for 3D face reconstruction based on occlusion segmentation, which can reduce the computational resource consumption of model deployment in 3D face reconstruction application scenarios, compress the model computation, and solve the technical problem of excessive computational resource consumption in 3D face reconstruction application scenarios.

[0005] In a first aspect, embodiments of this application provide a three-dimensional face reconstruction method based on occlusion segmentation, comprising: The target face image is input into a pre-built parameter prediction model, which includes an image feature extractor and an image segmentation decoder. The parameter prediction model is trained based on multiple face training images, facial key point information of the face training images, and face occlusion segmentation regions until the association loss function of the image feature extractor and the image segmentation decoder reaches a set state. The target face reconstruction parameters and target face occlusion region of the target face image are output based on the parameter prediction model, and 3D face reconstruction post-processing is performed based on the face reconstruction parameters and the face occlusion segmentation region.

[0006] In a second aspect, embodiments of this application provide a three-dimensional face reconstruction system based on occlusion segmentation, comprising: The input module is configured to input the target face image into a pre-built parameter prediction model. The parameter prediction model includes an image feature extractor and an image segmentation decoder. The parameter prediction model is trained based on multiple face training images, the facial key point information of the face training images, and the face occlusion segmentation region until the association loss function of the image feature extractor and the image segmentation decoder reaches a set state. The output module is configured to output the target face reconstruction parameters and the target face occlusion region of the target face image based on the parameter prediction model, and perform 3D face reconstruction post-processing based on the face reconstruction parameters and the face occlusion segmentation region.

[0007] In a third aspect, embodiments of this application provide a three-dimensional face reconstruction device based on occlusion segmentation, comprising: Memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the occlusion segmentation-based three-dimensional face reconstruction method as described in the first aspect.

[0008] In a fourth aspect, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions that, when executed by a computer processor, are configured to perform the occlusion-based three-dimensional face reconstruction method as described in the first aspect.

[0009] In a fifth aspect, embodiments of this application provide a computer program product containing instructions that, when executed on a computer or processor, cause the computer or processor to perform the occlusion-segmentation-based 3D face reconstruction method as described in the first aspect. This application embodiment inputs a target face image into a pre-constructed parameter prediction model. This model includes an image feature extractor and an image segmentation decoder. The model is trained based on multiple face training images, facial keypoint information from these training images, and face occlusion segmentation regions until the association loss function of the image feature extractor and image segmentation decoder reaches a predetermined state. Then, based on the parameter prediction model, it outputs the target face reconstruction parameters and the target face occlusion region of the target face image, and performs 3D face reconstruction post-processing based on the face reconstruction parameters and the face occlusion segmentation region. By employing the above technical means, and training a parameter prediction model containing an image feature extractor and an image segmentation decoder until the association loss function of these two components reaches a predetermined state, the parameter prediction model can integrate 3D face reconstruction and face occlusion segmentation functions. This reduces the computational resource consumption of the model deployment, reduces model redundancy, compresses model computation, and improves the efficiency of 3D face reconstruction.

[0010] Furthermore, by customizing the parameter dimensions of the target face reconstruction parameters and the number of channels of the image segmentation decoder, the computational load of the parameter prediction model can be adaptively configured to adapt the parameter prediction model to deployment environments supported by different computing power. Attached Figure Description

[0011] Figure 1 This is a flowchart of a three-dimensional face reconstruction method based on occlusion segmentation provided in an embodiment of this application; Figure 2 This is a flowchart of the parameter prediction model training process in the embodiments of this application; Figure 3 This is a schematic diagram of the sample input and output of the parameter prediction model in the embodiments of this application; Figure 4 This is a flowchart of the image preprocessing process in an embodiment of this application; Figure 5 This is a flowchart of the prediction process of the parameter prediction model in the embodiments of this application; Figure 6 This is a flowchart of the target face image processing in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a three-dimensional face reconstruction system based on occlusion segmentation provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a three-dimensional face reconstruction device based on occlusion segmentation provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0013] This application provides a 3D face reconstruction method based on occlusion segmentation, aiming to train a parameter prediction model that integrates an image feature extractor and an image segmentation decoder. This model couples 3D face reconstruction and occlusion region segmentation, thereby reducing the computational resource consumption of the model deployment and compressing the computational load. In traditional 3D face reconstruction scenarios, to avoid face occlusion affecting post-processing, a separate face occlusion segmentation model is deployed to predict occlusion regions, and then post-processing is performed based on the located occlusion regions. Since the 3D face reconstruction model and the face occlusion segmentation model are deployed independently, and there are redundant image processing steps between them, treating 3D face reconstruction and occlusion region segmentation as two separate tasks increases the platform's computational pressure and affects the operation of other platform services. Therefore, this application provides an embodiment that treats 3D face reconstruction and occlusion region segmentation as two separate tasks to solve the technical problem of excessive computational resource consumption in existing 3D face reconstruction applications.

[0014] Example: Figure 1 A flowchart of a 3D face reconstruction method based on occlusion segmentation provided in this application embodiment is given. This 3D face reconstruction method based on occlusion segmentation can be executed by a 3D face reconstruction device based on occlusion segmentation. This device can be implemented through software and / or hardware. The device can consist of two or more physical entities, or it can consist of a single physical entity. Generally, this 3D face reconstruction device based on occlusion segmentation can be a computer, mobile phone, tablet, image processing server, or other processing device.

[0015] The following description uses the occlusion-segmentation-based 3D face reconstruction device as an example to illustrate the occlusion-segmentation-based 3D face reconstruction method. (Refer to...) Figure 1 The 3D face reconstruction method based on occlusion segmentation specifically includes: S110. Input the target face image into the pre-constructed parameter prediction model. The parameter prediction model includes an image feature extractor and an image segmentation decoder. The parameter prediction model is trained based on multiple face training images, the facial key point information of the face training images, and the face occlusion segmentation region until the association loss function of the image feature extractor and the image segmentation decoder reaches the set state.

[0016] In this embodiment of the application, during 3D face reconstruction, a pre-constructed parametric prediction model is used to input a 2D face image intended for 3D face reconstruction into the target face image. The 2D face image is defined as the target face image. The parametric prediction model then predicts the 3D face reconstruction parameters and the occlusion region of the target face image, which are defined as the target face reconstruction parameters and the target face occlusion region. The parametric prediction model integrates an image feature extractor and an image segmentation decoder, and is trained based on the correlation loss function between the two. This allows the 3D face reconstruction and occlusion region segmentation tasks to be processed in parallel, improving model processing efficiency and promoting mutual improvement.

[0017] Prior to this, the parametric prediction model was pre-trained to enable it to perform 3D face reconstruction and occluded region segmentation tasks. (See reference...) Figure 2 The training process for the parameter prediction model includes: S1001. Multiple face training images, facial landmark information of the face training images, and face occlusion segmentation regions are used as training samples. S1002. Based on the training samples, a parameter prediction model is trained, and the corresponding face prediction key point information is output through the image feature extractor. The face prediction occlusion segmentation region is output through the image segmentation decoder. Based on the face prediction key point information, three-dimensional face reconstruction is performed to generate a three-dimensional face prediction image. S1003. Using the 3D face prediction image, face prediction key point information, and face prediction occlusion segmentation region as prediction samples, calculate the correlation loss function of the image feature extractor and the image segmentation decoder based on the training samples and prediction samples. When the correlation loss function reaches the set state, complete the training process of the parameter prediction model.

[0018] In this embodiment of the application, when training the parameter prediction model, multiple face images are input as face training images. Facial landmark information from each face training image is obtained, and a pre-trained face region segmentation model is used to determine the face occlusion segmentation region of the face training images. Using the aforementioned multiple face training images, facial landmark information from the face training images, and face occlusion segmentation regions as training samples, the image feature extractor and image segmentation decoder of the parameter prediction model are trained.

[0019] The process involves several steps. First, a face training image is input into an image feature extractor, which outputs corresponding face prediction keypoint information. This keypoint information is then input into a 3D face reconstruction model to reconstruct the 3D face. The 3D face model is determined using the keypoint information, and a differentiable renderer projects this model onto a 2D plane to render a 2D image, resulting in the predicted rendered image, i.e., the 3D face prediction image. Second, an image feature extractor extracts image features from the face training image, which are then input into an image segmentation decoder for face image segmentation, outputting the predicted occlusion segmentation region. The parametric prediction model uses the aforementioned 3D face prediction image, face prediction keypoint information, and predicted occlusion segmentation region as prediction samples. It then calculates the association loss function between the image feature extractor and the image segmentation decoder based on the training and prediction samples. The training of the parametric prediction model is considered complete when the association loss function reaches a predetermined value.

[0020] It should be noted that for the face training images in the training samples, the face region in the image may or may not include occluded areas. By training the parameter prediction model on face training images with different occlusion conditions, the stability and reliability of the model prediction can be improved.

[0021] The embodiments of this application design an association loss function to make the predicted sample gradually approach the original training sample. When the association loss function is in a set state, it means that the similarity between the training sample and the predicted sample meets the model prediction standard and can be applied to 3D face reconstruction.

[0022] For example, such as Figure 3 As shown, by inputting a face training image I T Facial landmark information from face training images (Im) T and face occlusion segmentation region M T Based on the above model training process, the corresponding 3D face prediction image I is output. R Face prediction key point information Im R Face prediction occlusion segmentation region M R It should be noted that the 3D face prediction image I output by the parameter prediction model... RIt only reconstructs the unobstructed parts of the face; in scenarios with occlusion, it does not reconstruct the corresponding occluding objects. For example... Figure 3 As shown, the predicted 3D face image I R It does not reconstruct the eyes or sunglasses. This avoids interference from occluded areas of the face in the reconstructed 3D face, optimizing the subsequent post-processing of the 3D face reconstruction.

[0023] Furthermore, the associated loss functions of the parameter prediction model include a segmentation loss function, a segmentation scaling loss function, and a face reconstruction loss function; among which, the segmentation loss function is used to measure the difference between the face occlusion segmentation region and the corresponding face prediction occlusion segmentation region; the segmentation scaling loss function is used to scale and adjust the face prediction occlusion segmentation region; and the face reconstruction loss is used to measure the difference between the face training image and the corresponding 3D face prediction image.

[0024] This application combines the relevant loss function for face region segmentation with the relevant loss function for face reconstruction, and introduces a segmentation scaling loss function to establish the connection between 3D face reconstruction and face occlusion segmentation. This makes face reconstruction parameter prediction more stable in the presence of occlusion, while face occlusion segmentation can also be more accurate. This allows the image feature extractor and image segmentation decoder to complete model training in a mutually reinforcing manner.

[0025] Furthermore, the segmentation scaling loss function includes a segmentation region magnification function and a segmentation region shrinkage function; the segmentation region magnification function is used to magnify the segmentation region obscured by face prediction, and the segmentation region shrinkage function is used to shrink the segmentation region obscured by face prediction.

[0026] Specifically, the segmentation loss function is expressed as: (1) (2) The segmentation scaling loss function is expressed as: (3) (4) (5) (6) in, M represents the segmentation loss function. T M represents the segmented area where the face is occluded. R This represents the predicted occlusion segmentation region of the face, and the difference between the occluded segmentation region and the corresponding predicted occlusion segmentation region is represented by a segmentation loss function; I T Represents face training images, I R This represents a 3D face prediction image. The number of pixels in the segmented region of face occlusion. Let represent the number of pixels in the predicted occluded segmentation region, and x represent the pixel value. Formulas (3) and (4) represent the segmentation region magnification function, which utilizes the characteristic that whether an image is occluded or not does not affect its perceptual characteristics, and maximizes the ratio between the number of pixels in the predicted occluded segmentation region and the predicted occluded segmentation region, so that the predicted predicted occluded segmentation region tends to expand outward as much as possible; at the same time, formulas (5) and (6) represent the segmentation region shrinkage function, where formula (5) indicates that when comparing pixel differences, the face training image I can be allowed to shrink. T With 3D face prediction image I R There is a slight displacement error between them; Formula (6) represents the face training image I under the predicted occlusion segmentation region of the predicted face. T With 3D face prediction image I R The perceptual errors between the rendered images should be as similar as possible; Formulas (5) and (6) tend to make the face prediction occlusion segmentation region ignore the parts with large pixel-level and perceptual layer errors, so that the predicted face prediction occlusion segmentation region tends to be as small as possible.

[0027] Combining formulas (1) to (6) above, the cross-entropy loss can be used to refine the face prediction occlusion segmentation region while maintaining the basic outline of the occlusion segmentation region. Furthermore, by establishing a connection between the segmentation scaling loss function and the face reconstruction part, and considering the applied face prediction occlusion segmentation region, the perceptual layer and pixel-level errors of the reconstructed face can be addressed. This allows 3D face reconstruction and face occlusion segmentation tasks to be performed in parallel, while simultaneously promoting mutual improvement in their effects.

[0028] Furthermore, the face reconstruction loss function is expressed as: (7) (8) (9) Wherein, formula (7) represents the face training image I T With 3D face prediction image I R In the unobstructed areas, the pixel level should be similar; Formula (8) represents the facial key point information Im. T and face prediction key point information Im R It should fit as closely as possible; Formula (9) represents the face training image I. T With 3D face prediction image I R The same should apply to model-aware systems.

[0029] The parameter prediction model is trained based on the correlation loss function of the above image feature extractor and image segmentation decoder until the correlation loss function reaches the set state. If the above correlation loss function formulas (1)-(9) converge to the set value, it means that the parameter prediction model training is completed and the prediction result of the parameter prediction model reaches the expected standard.

[0030] This application's embodiments couple 3D face reconstruction and face segmentation tasks together, simultaneously outputting 3D face reconstruction parameters and face occlusion regions within a single model, thereby reducing model deployment overhead. Furthermore, by combining three types of loss functions—segmentation loss function, segmentation scaling loss function, and face reconstruction loss function—the model learns the intrinsic relationship between 3D face reconstruction and face occlusion segmentation tasks, thus significantly reducing model redundancy and enabling the use of a smaller model to complete both tasks. This reduces model computation and improves processing efficiency.

[0031] Furthermore, based on the constructed parameter prediction model, when performing 3D face reconstruction and face occlusion segmentation of the target face image, this embodiment of the application preprocesses the target face image and inputs the preprocessed target face image into the parameter prediction model to perform 3D face reconstruction and face occlusion segmentation.

[0032] Reference Figure 4 The preprocessing steps for the target face image include: S1101. Based on the face landmark detector and the template face landmark registration of the target face image, the stretching and translation parameters of the target face image are obtained. S1102 crops the target face image based on stretching and translation parameters to make the target face image conform to the standard face size.

[0033] The preprocessing model for the target face image primarily involves filtering and correcting the image data input to the parameter prediction model. By registering the target face image with a facial landmark detector and template facial landmarks, the stretching and translation parameters of the preprocessed target face image are obtained. The target face image can then be processed and cropped using these parameters to conform to standard face dimensions, facilitating its use in subsequent parameter prediction models. It's understandable that the size of the face region varies across different target face images. To ensure the prediction accuracy of the parameter prediction model, the model standardizes the target face image, requiring the face region to be adjusted to standard face dimensions.

[0034] Then, for the preprocessed target face image, parameter prediction is performed based on the pre-trained parameter prediction model described above.

[0035] S120. Output the target face reconstruction parameters and the target face occlusion area of ​​the target face image based on the parameter prediction model, and perform three-dimensional face reconstruction post-processing based on the target face reconstruction parameters and the target face occlusion area.

[0036] The parameter prediction model in this embodiment receives a preprocessed target face image as input, predicts it through the model, and outputs the corresponding target face reconstruction parameters and the target face occlusion area.

[0037] Specifically, in the parameter prediction model, the target face image is input into the image feature extractor to obtain the corresponding feature map, the feature map is integrated to obtain the target face reconstruction parameters, and the feature map is input into the image segmentation decoder. Based on the image segmentation decoder, image segmentation is performed to obtain the target face occlusion region.

[0038] The overall framework of the parameter prediction model is as follows: Figure 5 As shown, this embodiment utilizes an improved lightweight Mobilenet-v3 network as an image-level feature extractor to learn the complete 3D facial structure geometry from image pixels, making it more suitable for deployment on mobile devices. Simultaneously, this embodiment connects the lightweight image segmentation decoder LR-ASPP to the image-level feature extractor, enabling it to efficiently extract deep features and detailed information, thereby achieving efficient image segmentation. Figure 5 The detailed structure of the parameter prediction model of this application embodiment is shown. The parameter prediction model takes the preprocessed target face image as input, passes through a series of bneck blocks to obtain a series of yellow feature maps, and then integrates the extracted features through 1x1 convolution to output the parameter prediction vector, that is, the target face reconstruction parameters.

[0039] The core component of the image-level feature extractor is the bneck module, which mainly implements channel-separable convolution, SE channel attention mechanism, and residual connections. Channel-separable convolution allows the model to achieve better feature extraction results with fewer parameters, the SE channel attention mechanism is used to adjust the weights of each channel, and combined with residual connections, the model can better combine high- and low-level features, laying the foundation for the model to learn 3D face parameters.

[0040] It should be noted that, in this embodiment, an image-level feature extractor capable of capturing 3D facial features is connected to an image segmentation decoder LR-ASPP that performs face segmentation tasks, thereby simultaneously outputting target face reconstruction parameters and the target face occlusion segmentation region. The LR-ASPP image segmentation decoder takes 56x56 and 7x7 feature maps as input, and uses an SE channel attention mechanism for further feature recalibration on the high-level feature map (7x7). Then, the high- and low-resolution features are respectively processed using 1x1 convolutional classes and then mixed. Through multi-level mixed feature learning, accurate segmentation of mobile images is achieved. Finally, the target face occlusion segmentation region is obtained.

[0041] Optionally, the parameter dimension of the target face reconstruction parameters output by the image feature extractor and the model computing power configuration of the parameter prediction model corresponding to the number of channels of the image segmentation decoder are specified. In this embodiment, the dimension of the final output target face reconstruction parameters can be freely defined by the user, who can specify it freely during the training phase based on the required model size and effect. The final output dimension is equal to the sum of identity (face ID) + expression (face expression) + albedo (face texture) + illumination (27 dimensions) + pose (3 dimensions) + translation (3 dimensions).

[0042] Furthermore, the computational load of the entire model structure can be controlled by a `width` parameter, which controls the number of channels in the model. Actual measurements show that when `width=0.5`, the entire model can be compressed to 20 MFLOPS, and excellent results are achieved in face reconstruction and occluded region segmentation on the evaluation set. This allows for deployment on various low-end devices and meets the module requirements of different computing power configurations. By specifying the required dimensions of the `identity` (face ID), `expression` (face facial expression), and `albedo` (face texture) parameters based on the actual application scenario and deployment environment, and controlling the computational load of the final model through the `width` parameter, customized models tailored to specific needs can be generated, improving the flexibility of model design.

[0043] Reference Figure 6 The provided steps a1-a5 are based on the above parameter prediction model. The preprocessed target face image is input into the parameter prediction model, and the image feature extractor based on the parameter prediction model obtains image features. The image features are used to generate target face reconstruction parameters on the one hand, and input into the image segmentation decoder on the other hand to generate the target face occlusion region. Then, the target face reconstruction parameters and the target face occlusion region are output for three-dimensional face reconstruction post-processing to complete the parameter prediction.

[0044] Furthermore, in the post-processing of 3D face reconstruction, the embodiments of this application perform 3D face reconstruction based on the target face reconstruction parameters to generate a target 3D face model. The target 3D face model includes the target 3D face shape and the target 3D face texture. Based on the occlusion area of ​​the target face, the occluded area on the target 3D face model is rendered using the target face image, and the unoccluded area on the target 3D face model is rendered using the target material.

[0045] Based on the target face reconstruction parameters, the three-dimensional face shape and three-dimensional face texture are reconstructed by combining the pre-generated face model base, and the target three-dimensional face model is generated.

[0046] The formula for constructing a target 3D face model is:

[0047]

[0048] Where S represents the three-dimensional shape of a human face. 3D face texture, For average human face shape, For average facial texture, , and These are PCA bases for face ID, facial expressions, and facial textures. as well as These are the corresponding coefficient vectors used to generate the 3D face model, which are obtained from the target face reconstruction parameters predicted by the parameter prediction model.

[0049] Other parameters output by the parametric prediction model, such as pose and translation parameters, can be used to correct the pose of the reconstructed 3D face model; while lighting parameters can be used to perform spherical harmonic lighting processing on the reconstructed face texture, thereby making the results more vivid and detailed.

[0050] It should be noted that the post-processing of the 3D face reconstruction in this application embodiment is described using 3D makeup in a live streaming scenario as an example. During live 3D makeup streaming, a 3D face model reconstructed from the user's face image can be used to apply 3D makeup materials. Then, the final rendering texture can be calculated based on the target face occlusion area predicted by the parameter prediction model. For occluded areas, the original image (and the target face image) captured by the camera can be used for rendering. For unoccluded areas, 3D makeup materials can be rendered according to normal logic. This achieves the goal of "using the original image captured by the camera for rendering occluded areas and using the reconstructed 3D makeup for rendering unoccluded areas," allowing users to enjoy a beautiful makeup visual effect while avoiding the makeup effect appearing to float on the occlusion, thus optimizing the 3D makeup effect.

[0051] In practical applications, the 3D face reconstruction method of this application embodiment can also be used in any 3D face reconstruction application scenario that requires real-time processing of potentially occluded input, such as 3D beauty and makeup, 3D special effects, and medical plastic surgery modeling in scenarios like live streaming and conferences. The specific application environment steps of this application embodiment are fixed and limited, and will not be elaborated here.

[0052] As described above, the target face image is input into a pre-constructed parametric prediction model, which includes an image feature extractor and an image segmentation decoder. The model is trained based on multiple face training images, facial keypoint information from these images, and face occlusion segmentation regions until the association loss function of the image feature extractor and image segmentation decoder reaches a predetermined state. Then, the model outputs the target face reconstruction parameters and the target face occlusion region from the target face image, and performs 3D face reconstruction post-processing based on these parameters and the occlusion segmentation region. By employing this technique, and training a parametric prediction model containing an image feature extractor and image segmentation decoder until their association loss function reaches a predetermined state, the model can integrate 3D face reconstruction and face occlusion segmentation functions. This reduces the computational resource consumption of the model deployment, decreases model redundancy, compresses computational load, and improves the efficiency of 3D face reconstruction.

[0053] Furthermore, by customizing the parameter dimensions of the target face reconstruction parameters and the number of channels of the image segmentation decoder, the computational load of the parameter prediction model can be adaptively configured to adapt the parameter prediction model to deployment environments supported by different computing power.

[0054] Based on the above embodiments, Figure 7 A schematic diagram of a 3D face reconstruction system based on occlusion segmentation provided in this application. (Reference) Figure 7 The 3D face reconstruction system based on occlusion segmentation provided in this embodiment specifically includes an input module and an output module.

[0055] The input module 21 is configured to input the target face image into a pre-constructed parameter prediction model. The parameter prediction model includes an image feature extractor and an image segmentation decoder. The parameter prediction model is trained based on multiple face training images, the face key point information of the face training images, and the face occlusion segmentation region until the association loss function of the image feature extractor and the image segmentation decoder reaches a set state. The output module 22 is configured to output the target face reconstruction parameters and the target face occlusion region of the target face image based on the parameter prediction model, and perform three-dimensional face reconstruction post-processing based on the face reconstruction parameters and the face occlusion segmentation region.

[0056] Specifically, the training process for the parameter prediction model includes: Multiple face training images, facial landmark information of the face training images, and face occlusion segmentation regions were used as training samples. The training parameters prediction model is trained based on training samples. The corresponding face prediction key point information is output through the image feature extractor. The face prediction occlusion segmentation region is output through the image segmentation decoder. Based on the face prediction key point information, three-dimensional face reconstruction is performed to generate a three-dimensional face prediction image. Using 3D face prediction images, face prediction key point information, and face prediction occlusion segmentation regions as prediction samples, the association loss function of the image feature extractor and image segmentation decoder is calculated based on the training samples and prediction samples. When the association loss function reaches the set state, the training process of the parameter prediction model is completed.

[0057] The associated loss functions include the segmentation loss function, the segmentation scaling loss function, and the face reconstruction loss function. The segmentation loss function measures the difference between the segmented region of the occluded face and the corresponding predicted segmented region of the occluded face. The segmentation scaling loss function is used to scale and adjust the predicted segmented region of the occluded face. The face reconstruction loss function measures the difference between the face training image and the corresponding 3D face prediction image.

[0058] The segmentation scaling loss function includes a segmentation region magnification function and a segmentation region shrinkage function; the segmentation region magnification function is used to magnify the segmentation region obscured by face prediction, and the segmentation region shrinkage function is used to shrink the segmentation region obscured by face prediction.

[0059] Specifically, the input module 21 is configured to input the target face image into the image feature extractor to obtain the corresponding feature map, integrate the feature map to obtain the target face reconstruction parameters, and input the feature map into the image segmentation decoder to perform image segmentation based on the image segmentation decoder to obtain the target face occlusion area.

[0060] The parameter dimensions of the target face reconstruction parameters output by the image feature extractor, and the model computing power configuration of the parameter prediction model corresponding to the number of channels of the image segmentation decoder.

[0061] Specifically, before inputting the target face image into the pre-built parameter prediction model, the following steps are also included: Based on the face landmark detector and the template face landmark registration of the target face image, the stretching and translation parameters of the target face image are obtained; The target face image is cropped based on stretching and translation parameters to make it conform to the standard face size.

[0062] Specifically, the output module 22 is configured to perform 3D face reconstruction based on the target face reconstruction parameters to generate a target 3D face model, which includes the target 3D face shape and the target 3D face texture; the target face image is used to render the occluded area on the target 3D face model based on the target face occlusion area, and the target material is used to render the unoccluded area on the target 3D face model.

[0063] As described above, the target face image is input into a pre-constructed parametric prediction model, which includes an image feature extractor and an image segmentation decoder. The model is trained based on multiple face training images, facial keypoint information from these images, and face occlusion segmentation regions until the association loss function of the image feature extractor and image segmentation decoder reaches a predetermined state. Then, the model outputs the target face reconstruction parameters and the target face occlusion region from the target face image, and performs 3D face reconstruction post-processing based on these parameters and the occlusion segmentation region. By employing this technique, and training a parametric prediction model containing an image feature extractor and image segmentation decoder until their association loss function reaches a predetermined state, the model can integrate 3D face reconstruction and face occlusion segmentation functions. This reduces the computational resource consumption of the model deployment, decreases model redundancy, compresses computational load, and improves the efficiency of 3D face reconstruction.

[0064] Furthermore, by customizing the parameter dimensions of the target face reconstruction parameters and the number of channels of the image segmentation decoder, the computational load of the parameter prediction model can be adaptively configured to adapt the parameter prediction model to deployment environments supported by different computing power.

[0065] The 3D face reconstruction system based on occlusion segmentation provided in this application embodiment can be configured to execute the 3D face reconstruction method based on occlusion segmentation provided in the above embodiment, and has corresponding functions and beneficial effects.

[0066] Based on the above practical examples, this application also provides a three-dimensional face reconstruction device based on occlusion segmentation, referring to... Figure 8The occlusion-segmentation-based 3D face reconstruction device includes a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The memory 32, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the occlusion-segmentation-based 3D face reconstruction method described in any embodiment of this application (e.g., input modules and output modules in an occlusion-segmentation-based 3D face reconstruction system). The communication module 33 is configured to perform data transmission. The processor 31 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory, thereby realizing the aforementioned occlusion-segmentation-based 3D face reconstruction method. The input device 34 can be configured to receive input digital or character information and generate key signal inputs related to user settings and function control of the device. The output device 35 may include a display screen or other display device. The occlusion-segmentation-based 3D face reconstruction device provided above can be configured to execute the occlusion-segmentation-based 3D face reconstruction method provided in the above embodiments, possessing corresponding functions and beneficial effects.

[0067] Based on the above embodiments, this application also provides a computer-readable storage medium storing computer-executable instructions. These instructions, when executed by a computer processor, are configured to perform a 3D face reconstruction method based on occlusion segmentation. The storage medium can be any type of memory device or storage device. Of course, the computer-readable storage medium provided in this application is not limited to the 3D face reconstruction method based on occlusion segmentation described above; it can also perform related operations in the 3D face reconstruction method based on occlusion segmentation provided in any embodiment of this application.

[0068] Based on the above embodiments, this application also provides a computer program product. The technical solution of this application, in essence or in other words, the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions to cause a computer device, mobile terminal, or processor therein to execute all or part of the steps of the occlusion segmentation-based three-dimensional face reconstruction method described in the various embodiments of this application.

Claims

1. A three-dimensional face reconstruction method based on occlusion segmentation, characterized in that, include: The target face image is input into a pre-constructed parameter prediction model, which includes an image feature extractor and an image segmentation decoder. The parameter prediction model is trained based on multiple face training images, facial keypoint information from the training images, and face occlusion segmentation regions until the association loss function of the image feature extractor and the image segmentation decoder reaches a set state. The association loss function includes a segmentation loss function, a segmentation scaling loss function, and a face reconstruction loss function. The segmentation loss function measures the difference between the face occlusion segmentation region and the corresponding predicted face occlusion segmentation region. The segmentation scaling loss function adjusts the scale of the predicted face occlusion segmentation region. The face reconstruction loss function measures the difference between the face training image and the corresponding 3D face prediction image. The target face reconstruction parameters and target face occlusion region of the target face image are output based on the parameter prediction model, and three-dimensional face reconstruction post-processing is performed based on the target face reconstruction parameters and target face occlusion region; the parameter dimension of the target face reconstruction parameters output by the image feature extractor and the number of channels of the image segmentation decoder correspond to the model computing power configuration of the parameter prediction model. Wherein, the segmentation region scaling function in the segmentation scaling loss function is: ; Among them, I T M represents face training images. R This represents the predicted occlusion segmentation region for faces. This indicates the number of pixels in the segmented region where the face is occluded. This represents the number of pixels in the predicted occluded segmentation region of the face. The segmentation region shrinkage function in the segmentation scaling loss function is: ; Among them, I R This represents a 3D face prediction image.

2. The 3D face reconstruction method based on occlusion segmentation according to claim 1, characterized in that, The training process for the parameter prediction model includes: Multiple face training images, facial key point information of the face training images, and face occlusion segmentation regions are used as training samples. The parameter prediction model is trained based on the training samples, the corresponding face prediction key point information is output through the image feature extractor, the face prediction occlusion segmentation region is output through the image segmentation decoder, and three-dimensional face reconstruction is performed based on the face prediction key point information to generate a three-dimensional face prediction image. Using the 3D face prediction image, the face prediction key point information, and the face prediction occlusion segmentation region as prediction samples, the association loss function of the image feature extractor and the image segmentation decoder is calculated based on the training samples and the prediction samples. When the association loss function reaches a set state, the training process of the parameter prediction model is completed.

3. The 3D face reconstruction method based on occlusion segmentation according to claim 1, characterized in that, The segmentation scaling loss function includes a segmentation region magnification function and a segmentation region shrinkage function; the segmentation region magnification function is used to magnify the face prediction occlusion segmentation region, and the segmentation region shrinkage function is used to shrink the face prediction occlusion segmentation region.

4. The 3D face reconstruction method based on occlusion segmentation according to claim 1, characterized in that, The step of outputting the target face reconstruction parameters and the target face occlusion region of the target face image based on the parameter prediction model includes: The target face image is input into the image feature extractor to obtain the corresponding feature map. The feature map is integrated to obtain the target face reconstruction parameters. The feature map is then input into the image segmentation decoder. Image segmentation is performed based on the image segmentation decoder to obtain the target face occlusion region.

5. The 3D face reconstruction method based on occlusion segmentation according to claim 1, characterized in that, Before inputting the target face image into the pre-built parametric prediction model, the following steps are also included: Based on the face landmark detector and the template face landmark registration, the stretching and translation parameters of the target face image are obtained. The target face image is cropped based on the stretching and translation parameters to make the target face image conform to the standard face size.

6. The 3D face reconstruction method based on occlusion segmentation according to claim 1, characterized in that, The post-processing for 3D face reconstruction based on the target face reconstruction parameters and the target face occlusion area includes: Based on the target face reconstruction parameters, a three-dimensional face reconstruction is performed to generate a target three-dimensional face model, which includes the target three-dimensional face shape and the target three-dimensional face texture. Based on the target face occlusion area, the occlusion area on the target 3D face model is rendered using the target face image, and the unoccluded area on the target 3D face model is rendered using the target material.

7. A three-dimensional face reconstruction system based on occlusion segmentation, characterized in that, include: The input module is configured to input a target face image into a pre-constructed parameter prediction model. The parameter prediction model includes an image feature extractor and an image segmentation decoder. The model is trained based on multiple face training images, facial keypoint information from the training images, and face occlusion segmentation regions until the association loss function of the image feature extractor and the image segmentation decoder reaches a set state. The association loss function includes a segmentation loss function, a segmentation scaling loss function, and a face reconstruction loss function. The segmentation loss function measures the difference between the face occlusion segmentation region and the corresponding predicted face occlusion segmentation region. The segmentation scaling loss function adjusts the scale of the predicted face occlusion segmentation region. The face reconstruction loss function measures the difference between the face training image and the corresponding 3D face prediction image. The output module is configured to output the target face reconstruction parameters and the target face occlusion region of the target face image based on the parameter prediction model, and perform three-dimensional face reconstruction post-processing based on the face reconstruction parameters and the face occlusion segmentation region; the parameter dimension of the target face reconstruction parameters output by the image feature extractor and the number of channels of the image segmentation decoder correspond to the model computing power configuration of the parameter prediction model. Wherein, the segmentation region scaling function in the segmentation scaling loss function is: ; Among them, I T M represents face training images. R This represents the predicted occlusion segmentation region for faces. This indicates the number of pixels in the segmented region where the face is occluded. This represents the number of pixels in the predicted occluded segmentation region of the face. The segmentation region shrinkage function in the segmentation scaling loss function is: ; Among them, I R This represents a 3D face prediction image.

8. A three-dimensional face reconstruction device based on occlusion segmentation, characterized in that, include: Memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the occlusion segmentation-based three-dimensional face reconstruction method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a computer processor, are configured to perform the occlusion-based three-dimensional face reconstruction method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes instructions that, when executed on a computer or processor, cause the computer or processor to perform the occlusion-based three-dimensional face reconstruction method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Face texture feature extraction method and device, 3D face reconstruction method and device and storage medium

    CN113111861A

  • Three-dimensional face reconstruction model training method and system and readable storage medium

    CN115115784A