Head three-dimensional model generation method based on 2D facial expression analysis

By combining facial expression recognition and neutralization, occlusion removal, and multi-view image generation with differentiable rendering technology, the shape and texture of the 3D head model are optimized, solving the problem of generating high-precision 3D head models in existing technologies. This method is applicable to fields such as virtual reality, game development, and film and television production.

CN119516122BActive Publication Date: 2025-11-18TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411797801.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-11-18
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently generate high-precision 3D head models, especially when dealing with facial expressions and occlusions. This results in insufficient geometric structure and texture fit of the model, making it difficult to generate neutral facial expression images suitable for different angles.

Method used

By employing techniques such as facial expression recognition and neutralization, automatic occlusion identification and removal, multi-view image generation, and differentiable rendering, the shape parameters and textures of the 3D head model are optimized. Combined with deep learning and generative adversarial networks, neutral facial expression images from multiple angles are generated, and projection errors are iteratively optimized to generate a high-precision 3D model.

Benefits of technology

It significantly improves the generation accuracy and consistency of 3D head models, reduces geometric distortion caused by facial expressions and occlusions, and generates high-fidelity 3D head models suitable for fields such as virtual reality, game development, film and television production, and medical imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516122B_ABST
    Figure CN119516122B_ABST
Patent Text Reader

Abstract

A three-dimensional head model generation method based on 2D facial expression analysis includes four steps of expression analysis, occlusion processing, multi-angle image generation and 3D model optimization. First, the expression parameters are extracted by the expression recognition model, and the expression is neutralized to generate expressionless images. Then, the image segmentation and deep learning inpainting technology are used to automatically identify and remove the occlusion, and the facial texture continuity is maintained. Then, multi-angle views are generated from the unoccluded neutral expression images to ensure the consistency of facial geometry and texture. Finally, the differentiable rendering technology is used to optimize the shape and map of the 3D model, realizing high consistency with the 2D image, so as to obtain high-precision and high-fidelity 3D head model. The multi-angle optimization method combines differentiable rendering with deep learning to improve the accuracy and consistency of 3D model generation. The neutral expression generation reduces geometric distortion, and the automatic removal of occlusion effectively improves the accuracy of facial reconstruction. It has strong practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and image processing, and in particular to a method for generating a 3D head model based on 2D facial expression analysis. Background Technology

[0002] With the increasing demand for 3D facial reconstruction in the fields of virtual reality (VR), augmented reality (AR), and film and entertainment, efficient and high-precision 3D head model generation technology is becoming increasingly important given limited computing resources. However, existing technologies cannot efficiently generate 3D head models from only 2D facial images with expressions, especially when dealing with occlusions (such as hair, glasses, and beards). Current technologies cannot accurately remove occlusions while preserving the fidelity of facial features and generating neutral expression images suitable for different angles. This directly affects the quality of subsequent 3D model generation.

[0003] Currently, with the rapid development of computer image processing technology, the generation and application of 3D models has become an important part of many fields. Especially in areas such as virtual reality (VR), augmented reality (AR), film and television entertainment, game development, medical imaging, and remote communication, the demand for 3D character modeling is increasing. Faced with these application scenarios, how to efficiently and accurately generate 3D character head models from 2D images has become a key research focus in this field.

[0004] In existing 2D-to-3D face modeling research, the common approach is to use deep learning techniques for image-to-geometry conversion. These techniques construct neural networks to map input 2D images into depth information or 3D geometric structures. While some progress has been made in the coarse reconstruction of textureless 3D geometry, generating high-fidelity, realistic 3D face models, especially in complex situations such as varying facial expressions, occlusions (e.g., glasses, hair), and multi-angle matching, traditional 2D-to-3D conversion techniques still face numerous challenges. Specifically, existing technologies are insufficient in addressing the following issues:

[0005] 1. The issue of handling facial expression changes:

[0006] In most everyday portrait photographs, people often display some kind of expression (such as smiling, frowning, or surprise). However, in 3D portrait modeling and many practical applications, a 3D head model with a neutral expression is generated. Therefore, generating a neutral-expression 3D model from an expressive 2D image has become a technical challenge. Existing technologies typically rely on directly generating depth information, but they haven't effectively solved how to automatically neutralize complex and diverse facial expressions. Using expressive images for 3D modeling can lead to deviations in facial geometry, especially when the expression is intense; this error can even affect the overall geometric accuracy of the model.

[0007] 2. The impact of obstructions:

[0008] In real-world shooting scenarios, people often wear glasses, hats, or have hair, beards, or other obstructions. These obstructions interfere with the geometric information of different areas of the face, causing significant errors when directly extracting facial geometry from 2D images. For example, glasses may affect the geometric accuracy of the eyes, hair may obscure the head contour, and beards may cover the mouth and chin area. Current technologies for handling obstructions typically rely on manual removal or outright removal of the obstructed areas, resulting in a lack of detail in the generated 3D model. How to automatically identify and remove these obstructions, and still restore high-fidelity, seamless facial textures after removing the obstructed areas, has become a pressing problem to be solved.

[0009] 3. Limitations of multi-view imaging:

[0010] Generating 3D head models typically requires views from different angles to reconstruct the head's true geometry. While this problem can be solved with specialized equipment in traditional reconstruction techniques such as photogrammetry and structured light, most 2D images (like ordinary portrait photographs) only contain one view of the head. Existing 2D-to-3D conversion techniques have limited ability to recover geometric details from single-view images, resulting in models that often exhibit coarseness and distortion in 3D structure. Furthermore, input images from a single angle can contain ambiguities (especially under complex lighting conditions), affecting model accuracy. Therefore, generating a 3D head model with high-precision geometric details from a single-view image is a challenging task.

[0011] 4. The issue of combining model generation and texture generation:

[0012] In existing 3D modeling techniques based on 2D images, the generated geometry and texture maps are often separate. The model shape and texture map need to be optimized separately, which easily leads to insufficient geometry-texture fit in the final model from different viewpoints, resulting in poor visual effects. Achieving joint optimization of the geometry model and texture maps, as well as generating high-quality, high-fidelity 3D textures, remains a technical bottleneck in this field.

[0013] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0014] The main objective of this invention is to overcome the deficiencies in the aforementioned background technology and provide a method for generating a three-dimensional head model based on 2D facial expression analysis.

[0015] To achieve the above objectives, the present invention adopts the following technical solution:

[0016] A method for generating a 3D head model based on 2D facial expression analysis includes the following steps:

[0017] S1: Use an expression recognition model to analyze the expression of the input 2D face image and extract expression parameters; based on the extracted expression parameters, use an expression editing model to neutralize the expression in the input image and generate a neutral expression image;

[0018] S2: Combine image segmentation algorithms to automatically identify occluded areas in neutral expression images; use deep learning-based inpainting technology to remove occluded areas and automatically fill in the removed areas to maintain the consistency of facial texture;

[0019] S3: Based on unobstructed neutral expression images, a face image generation model is used to generate multiple neutral expression images from different angles to ensure the continuity and consistency of facial geometry and texture under different viewpoints.

[0020] S4: Combining multi-view neutral facial expression images, differentiable rendering technology is used to reverse optimize the shape parameters and textures of the parameterized 3D head model; by iteratively optimizing the projection error, the generated 3D head model is ensured to be consistent with the 2D images under different viewpoints, and high-precision geometry and high-fidelity textures are obtained.

[0021] Furthermore, step S1 specifically includes:

[0022] The expression parameter model is used to perform expression analysis on the input 2D face image, and the output expression vector describes the activation intensity of each action unit on the face.

[0023] Based on this expression vector, the original image with expression is edited using an expression editing model to generate a neutral expression image;

[0024] The generated neutral expression image is re-detected by the expression recognition model. If the difference between the weight of any action unit and the neutral expression vector exceeds a preset threshold, the image is re-input into the expression editing model for iterative optimization until the generated image is close to a neutral expression.

[0025] Furthermore, step S2 specifically includes:

[0026] An image segmentation algorithm is used to process the neutral expression image to generate a segmentation mask. This mask marks each pixel to identify the occluded area.

[0027] Based on the labeling results of the segmentation mask, the color values ​​of pixels identified as occluders in the image are set to zero, and corresponding occluder masks are created.

[0028] The occlusion mask and the image are input into a deep learning image generation model. The model's inpainting function is used to remove the occlusion and intelligently fill in the original occlusion position to maintain the continuity and consistency of facial texture.

[0029] Furthermore, step S3 specifically includes:

[0030] Using a human-image-generated-image model, neutral expression images from multiple angles are generated from unoccluded neutral expression images;

[0031] By applying perspective transformation and geometric deformation techniques, facial images from multiple angles are generated from a standard image taken from a single angle.

[0032] This ensures that images generated from different viewpoints retain neutral facial expression features and maintain the continuity and consistency of facial geometry and texture.

[0033] Furthermore, the human-generated image model is an image generation model based on Generative Adversarial Network (GAN) or Stable Diffusion.

[0034] Furthermore, the facial images from multiple angles include viewing angles from the left, right, top, bottom, and 15, 30, and 45 degrees.

[0035] Furthermore, step S4 specifically includes:

[0036] Using the multi-view, neutral, and unobstructed 2D neutral expression images generated in step S3, the 3D parametric model is iteratively optimized using differentiable rendering technology to generate a 3D head model and texture.

[0037] Randomly initialize the shape parameters and textures of the 3D model, extract vertex positions from the 3D head model, and adjust the vertex positions according to the shape parameters;

[0038] The 3D model is rendered to a viewpoint corresponding to the multi-angle 2D image, and iterative optimization is performed to minimize the difference between the rendered image and the actual 2D image.

[0039] Furthermore, the iterative optimization includes calculating a loss function, which is the sum of squares of the differences between the rendered image and the corresponding 2D image for all views, wherein the difference between the rendered result and the actual 2D image for each view is measured by Euclidean distance.

[0040] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for generating a 3D head model based on 2D facial expression analysis.

[0041] A computer program product includes a computer program that, when executed by a processor, implements the method for generating a 3D head model based on 2D facial expression analysis.

[0042] The present invention has the following beneficial effects:

[0043] This invention proposes a fully optimized 3D head model generation method, particularly suitable for 2D face images in complex scenes, including images with facial expressions and facial occlusions. This method utilizes automated image processing techniques, such as expression recognition and neutralization, multi-view generation, automatic occlusion identification and removal, and differentiable rendering optimization, to generate high-precision 3D head models from single frames or a small number of 2D images, significantly improving modeling efficiency and accuracy. This process effectively addresses facial geometric distortions caused by expression changes, automatically removes occlusions to reduce interference with modeling, and generates high-precision image templates from multiple angles, providing multiple safeguards for 3D reconstruction.

[0044] This invention first analyzes the input 2D facial image with facial expressions using an expression recognition model to extract expression parameters and obtain a neutral expression. Then, an expression editing model is used to automatically neutralize the expression, generating an expressionless image. Next, the system combines image segmentation algorithms to automatically identify and mark occluded areas, such as glasses, hair, and beards, and uses deep learning-based inpainting technology to intelligently remove these occluders while maintaining the consistency of facial texture. After processing the occluders, the system uses a face image generation model to generate multiple images from different angles from the unoccluded neutral expression image, ensuring the continuity and consistency of facial geometry and texture from different perspectives. Finally, the system employs differentiable rendering technology, iteratively optimizing the shape parameters and textures of the parameterized 3D head model through reverse optimization, continuously iterating to optimize projection errors to ensure that the generated 3D head model is consistent with the 2D images from multiple perspectives, achieving high-precision geometry and high-fidelity textures.

[0045] The significant innovation of this invention lies in providing a multi-angle optimization method that combines differentiable rendering and deep learning to improve the accuracy and consistency of 3D model generation. This method reduces geometric distortion caused by facial expression changes and occlusions, thus improving the accuracy of facial reconstruction. The fully automated system does not rely on specific hardware and is applicable to multiple fields such as virtual reality, game development, film and television production, and medical imaging, demonstrating strong practical application value. Furthermore, the efficiency and versatility of this invention make it suitable for various scenarios. Users only need to provide conventional 2D photos to generate reliable 3D models, applicable to fields such as virtual characters, game development, and medical imaging. It is a fully automated computational process that improves generation efficiency and accuracy, suitable for various application scenarios such as large-scale image processing and virtual human character generation.

[0046] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description

[0047] Figure 1 This is a flowchart of a method for generating a 3D head model based on 2D facial expression analysis according to an embodiment of the present invention. Detailed Implementation

[0048] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0049] Existing technologies for generating 3D head models from 2D portrait photos often encounter two major challenges: First, most everyday photos typically feature people with various facial expressions, while practical applications of 3D head model generation often require generating 3D models with neutral expressions. This poses a challenge to traditional 3D model generation methods, as these methods generally struggle to handle complex and diverse facial expressions and cannot directly generate accurate neutral-expression 3D models from 2D photos with facial expressions. Second, in many images, people wear glasses or have facial occlusions such as beards, which interfere with the accurate capture of facial geometric features, thus affecting the generation of the 3D model.

[0050] Therefore, the present invention aims to propose an innovative process to overcome these two major problems, providing an optimized 3D head model generation method for 2D facial images with expressions and occlusions. Through expression recognition and neutralization processing, automatic identification and removal of occlusions (such as glasses, hair, beards, etc.), and generation of neutral, unoccluded facial images from multiple angles, and further generation of a 3D head model, this method can solve the challenges of expression and occlusion in model generation while preserving facial details and geometric features.

[0051] See Figure 1 This invention provides a method for generating a 3D head model based on 2D facial expression analysis, comprising the following steps:

[0052] Step S1: Automatic Expression Recognition and Automatic Expression Neutrality: The expression recognition model is used to analyze the expression of the input 2D face image and extract expression parameters; based on the extracted expression parameters, the expression editing model is used to neutralize the expression in the input image and generate a neutral expression image.

[0053] In a preferred embodiment, step S1 specifically includes: performing expression analysis on the input 2D face image using an expression parameter model, and outputting an expression vector describing the activation intensity of each action unit on the face; based on the expression vector, editing the original expression image using an expression editing model to generate a neutral expression image; and re-detecting the generated neutral expression image using an expression recognition model. If the difference between the weight of any action unit and the neutral expression vector exceeds a preset threshold, the image is re-inputted into the expression editing model for iterative optimization until the generated image is close to a neutral expression.

[0054] Step S2: Automatic Occlusion Identification and Removal: Combine image segmentation algorithms to automatically identify occlusion regions in neutral expression images; use deep learning-based inpainting technology to remove occlusions and automatically fill in the removed areas to maintain the consistency of facial texture.

[0055] In a preferred embodiment, step S2 specifically includes: processing the neutral expression image using an image segmentation algorithm to generate a segmentation mask, which marks each pixel to identify occlusion areas; setting the color values ​​of pixels identified as occlusions in the image to zero according to the marking results of the segmentation mask, and creating corresponding occlusion masks; inputting the occlusion mask and the image into a deep learning image generation model (such as deep learning algorithms like convolutional neural networks (CNN), using the model's inpainting function to remove occlusions, and intelligently filling in the original occlusion positions to maintain the continuity and consistency of facial textures.

[0056] This step combines image segmentation and deep learning inpainting techniques to achieve accurate identification and effective removal of occlusions, while maintaining the natural transition and consistency of facial textures, thus improving the accuracy and realism of 3D model generation.

[0057] Step S3: Automatic generation of neutral expression images from multiple angles: Based on the unobstructed neutral expression image, a face image generation model is used to generate multiple neutral expression images from different angles to ensure the continuity and consistency of facial geometry and texture under different viewpoints.

[0058] In a preferred embodiment, step S3 specifically includes: using a human-image-generated image model to generate neutral expression images from multiple angles from an unobstructed neutral expression image, which may include, but is not limited to, viewing angles from left, right, top, bottom, and directions such as 15, 30, and 45 degrees; applying perspective transformation and geometric deformation techniques to generate facial images from multiple angles from a standard image at a single angle; thereby ensuring that the images generated at different viewing angles maintain neutral expression features and maintain the continuity and consistency of facial geometry and texture.

[0059] This step maintains consistency in lighting and facial features by expanding a single-angle image to multiple angle images, while ensuring the preservation of neutral facial expressions, providing rich input data for generating high-precision 3D head models.

[0060] In different embodiments, the human-generated image model can be an image generation model based on generative adversarial networks (GANs) or stable diffusion.

[0061] Step S4: Automatic generation and optimization of 3D head model: Combining multi-view neutral expression images, differentiable rendering technology is used to reverse optimize the shape parameters and textures of the parameterized 3D head model; by iteratively optimizing the projection error, it is ensured that the generated 3D head model is consistent with the 2D images under different viewpoints, and high-precision geometry and high-fidelity textures are obtained.

[0062] In a preferred embodiment, step S4 specifically includes: using the multi-view, neutral, and unobstructed 2D neutral expression images generated in step S3, employing differentiable rendering technology to iteratively optimize the 3D parametric model to generate a 3D head model and textures; randomly initializing the shape parameters and textures of the 3D model, extracting vertex positions from the 3D head model, and adjusting the vertex positions according to the shape parameters; rendering the 3D model to the viewpoints corresponding to the multi-angle 2D images, and performing iterative optimization to minimize the difference between the rendered image and the actual 2D image. The iterative optimization includes calculating a loss function, which is the sum of squares of the differences between the rendered image and the corresponding 2D image for all viewpoints, where the difference between the rendering result of each viewpoint and the actual 2D image is measured by Euclidean distance.

[0063] This step, through an iterative process of multi-view optimization and differentiable rendering, precisely matches the input multi-angle 2D images to achieve fine geometry and high-fidelity textures, thereby providing higher accuracy and consistency for the generation of 3D models.

[0064] This invention proposes a full-process optimization method for 2D face images with expressions and occlusions. From image expression analysis, occlusion recognition and removal, neutral expression generation, to multi-angle image generation and final 3D model generation, it provides a more efficient, accurate and compatible solution.

[0065] The following further describes specific embodiments of the present invention and examples of its algorithm implementation.

[0066] A method for generating 3D head models based on 2D face images and expression analysis is proposed. This method optimizes and generates high-precision 3D head models with textures through a series of processing steps, including expression analysis, image segmentation, occlusion handling, and multi-angle image generation. This method can automatically generate accurate 3D head models from 2D images with different expressions and facial occlusions, including texture optimization, effectively solving the problem in existing methods where images with expressions and occlusions cannot generate high-precision 3D models.

[0067] First, the system analyzes the input 2D portrait image with facial expressions using an expression recognition model and extracts expression parameters to obtain a neutral expression for the person. Then, an expression editing model is used to automatically neutralize the expression, generating an image with a neutral expression. Second, to reduce interference from occlusions in the 3D model generation, the system combines image segmentation algorithms to automatically identify and mark occlusion areas (such as glasses, hair, and beards), and uses deep learning-based inpainting technology for intelligent removal and filling, ensuring the consistency of facial texture after occlusion removal.

[0068] After generating unobstructed neutral expression images, the system generates multiple neutral images from different angles based on these images. It then uses a face image generation model (such as GAN or Stable Diffusion) to extend the multi-view generated images, ensuring consistency in facial geometry and texture when viewed from different perspectives. Finally, combining these multi-view neutral expression images, the system employs differentiable rendering technology to inversely optimize the shape parameters and textures of the parameterized 3D head model. By continuously iterating and optimizing the projection error, the system ensures that the generated 3D head model is consistent with the 2D images from different viewpoints, thereby obtaining high-precision geometry and high-fidelity textures.

[0069] This invention provides an innovative multi-angle optimization method that improves the accuracy and consistency of 3D model generation by combining differentiable rendering and deep learning. Neutral expression generation reduces geometric distortion, and automatic removal of occlusions effectively improves the accuracy of facial reconstruction. This fully automated system requires no specific hardware and is widely applicable to various fields such as virtual reality, game development, film and television production, and medical imaging, demonstrating strong practical application value.

[0070] A method for generating high-precision 3D head models from 2D face images, specifically including: Figure 1 The four steps shown:

[0071] Step 1: Automatic Facial Expression Recognition and Automatic Facial Expression Neutrality

[0072] 1.1 The system first employs an expression recognition model to analyze the expressions of the input 2D face image and output expression parameters. Specifically, for the input image I∈R... H×W×C Facial expression parameter model f AU Transform the input image into a set of k-dimensional expression vectors n∈R k It is used to accurately describe the activation intensity of each Action Unit (AU) on the human face.

[0073] 1.2 Using an expression editing model, the original image with facial expressions is edited based on the expression parameters to convert the original image into a neutral expression image. Specifically, given an input image I∈R with facial expressions... H×W×C And the expression vector n∈R obtained in step 1.1 above k The expression editing model is then used to approximate the neutral expression, generating a new image with the expression removed. This generated image is then re-detected by the expression recognition model. If the weight of any action unit in the output expression parameter vector differs from the neutral expression vector n0 by more than a preset threshold ∈, the generated image is re-input into the expression editing model for further optimization until the generated image approximates a neutral expression.

[0074] Initialize I0 = I, and iterate cyclically for t = 0, 1, 2, ...:

[0075] I t+1 =E(I t ,n)

[0076] n t+1 =f AU (I t+1 )

[0077] Termination conditions:

[0078]

[0079] In this invention, expression neutralization does not rely solely on the deletion of a single expression. Instead, it combines captured expression parameters with a neutral expression generation method that optimizes image quality and uses high-precision positioning to protect facial details and texture information.

[0080] Step 2: Automatic identification and removal of obstructions

[0081] 2.1 Using an image segmentation algorithm, the neutral expression image Ineutral∈R generated in step 1.2 above is processed. H ×W×C Generate a segmentation mask (M∈R) H×W The mask assigns a label m to each pixel. i,j ∈{0,1,2,…,N}, mark the occluded areas (such as hair, glasses, beard, etc.) pixel by pixel.

[0082] 2.2 Based on the pixel markings in 2.1 above, the color values ​​of pixels marked as occluders in the image are set to 0, and an additional image mask is generated for the occluder markings. The mask and the image are input into Stable Diffusion or other powerful deep learning-based image generation models, and their inpainting methods are used to remove the occluders and intelligently fill in the original occluded positions to ensure that the generated facial textures have consistency.

[0083] Unlike existing occlusion handling methods that rely on a single specific model, this invention proposes a dual approach based on image segmentation and intelligent generation. First, a segmentation model accurately identifies different types of occlusions. Second, an inpainting model intelligently fills in facial textures as needed, achieving a general workflow that can handle various complex situations without prior training.

[0084] Step 3: Automatic generation of neutral facial expression images from multiple angles

[0085] 3.1 In this method, based on the generated unoccluded neutral expression images, the system further uses a human-image-generated image model (such as an image generation model based on GAN or Stable Diffusion) to generate neutral expression images from different angles. First, through the preliminary expression recognition and occlusion processing, an unoccluded, expressionless standard image is obtained as the basic input. Next, the human-image-generated image model uses viewpoint transformation and geometric deformation techniques to generate facial images from multiple angles, such as images observed from the left, right, top, and bottom at 15, 30, and 45 degrees. The key to this step is to ensure that the generated images maintain neutral expression features while ensuring the continuity and consistency of facial geometry and texture from different viewpoints. These expressionless images from different angles will provide more data support for the subsequent high-precision generation of 3D head models.

[0086] In this invention, an image generation model is used to augment an image from a single angle into an image from multiple angles, while maintaining consistency in lighting, facial features, and neutral facial expressions, thus providing richer input for the subsequent generation of high-precision 3D head models.

[0087] Step 4: Automatic generation and optimization of 3D avatar models

[0088] 4.1 Multi-view, neutral, unobstructed 2D neutral facial expression images generated according to step 3.1 (i) ∈R H×W×C Where i = 1, 2, ..., N, this method uses differentiable rendering to iteratively optimize a 3D parametric model to generate the 3D head model and texture. Specifically, the shape parameter β and texture I are first randomly initialized. tex Extracting vertex positions from a 3D head model Then, this 3D model is rendered to the angle corresponding to each image in step 3.1 above, and then iteratively optimized to minimize the difference between the projected image and the multi-angle image generated in step 3.1 above. The optimization function is as follows:

[0089]

[0090] Where T (i) Let R(V,T) be the viewpoint transformation matrix. (i) ) is the rendering function.

[0091] In this invention, the aforementioned steps amplify a single image into a multi-angle neutral expression image, providing sufficient information for differentiable rendering and reverse optimization. This iterative process based on multi-view optimization and differentiable rendering helps to more accurately match the input image, achieving detailed geometry and high-fidelity textures, thus providing higher accuracy and consistency for 3D model generation.

[0092] In summary, this invention provides a method for generating a 3D head model based on 2D facial expression analysis. This method can generate a high-precision 3D head model from a single frame or a small number of 2D facial images with expressions through a fully automated process. By employing expression recognition and neutralization processing, this method effectively reduces facial geometric distortion caused by expression changes, improving the accuracy of shape and expression reconstruction in the 3D model. Automatic removal of occlusions, such as glasses, hair, and beards, eliminates interference with facial geometry, reducing reconstruction errors and enhancing the accuracy of facial region reconstruction. Furthermore, by utilizing multi-angle image generation technology combined with differentiable rendering technology, the shape and texture of the 3D model are iteratively optimized, significantly improving the geometric details and texture realism of the 3D model. This method not only improves modeling efficiency and accuracy but also has broad applicability, requiring no specific hardware and applicable to multiple fields such as virtual reality, game development, film and television production, and medical imaging, demonstrating strong practical application value.

[0093] Compared with traditional technologies, the main advantages of this invention are:

[0094] This invention effectively improves the accuracy and consistency of 3D head model generation. First, using neutral facial expressions to generate images significantly reduces geometric distortion caused by facial expressions, resulting in more accurate shape and expression reproduction in the reconstructed 3D model. By automatically removing occlusions (such as glasses, hair, and beards) from the images, the interference of these occlusions on facial geometry is eliminated, greatly reducing reconstruction errors caused by hair and glasses occlusion and improving the accuracy of facial region reproduction. Simultaneously, the system uses multi-angle generated neutral images, combined with differentiable rendering technology, to iteratively optimize the shape and texture of the 3D model. By projecting the model from different viewpoints onto a 2D image matrix, this method effectively reduces information ambiguity and blurring caused by single-viewpoint rendering, thereby generating higher-quality, more detailed 3D head models. These can be widely applied to 3D image restoration in virtual reality, film and television production, and the medical field.

[0095] This invention combines high efficiency and versatility, making it widely applicable to various scenarios. It requires no specific camera equipment; users only need to provide a standard 2D photograph to generate a reliable 3D model, suitable for fields such as virtual characters, game development, and medical imaging. This invention is a fully automated computational process that improves generation efficiency and accuracy, making it suitable for large-scale image processing, virtual human character generation, and many other applications.

[0096] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.

[0097] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.

[0098] This invention also provides a processor that executes a computer program, at least performing the methods described above.

[0099] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk drive or magnetic tape drive. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0100] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0101] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0102] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0103] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0105] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0106] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0107] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0108] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.

Claims

1. A method for generating a 3D head model based on 2D facial expression analysis, characterized in that, Includes the following steps: S1: Use an expression recognition model to analyze the expression of the input 2D face image and extract expression parameters; Based on the extracted facial expression parameters, an expression editing model is used to neutralize the facial expressions in the input image, generating a neutral facial expression image. Specifically, an expression recognition model is used to analyze the facial expressions of the input 2D face image, outputting an expression vector describing the activation intensity of each action unit on the face. Based on this expression vector, the expression editing model is used to edit the original image with facial expressions to generate a neutral facial expression image. S2: Combine image segmentation algorithms to automatically identify occluded areas in neutral expression images; use deep learning-based inpainting technology to remove occluded areas and automatically fill in the removed areas to maintain the consistency of facial texture; S3: Based on unobstructed neutral expression images, a face image generation model is used to generate multiple neutral expression images from different angles to ensure the continuity and consistency of facial geometry and texture under different viewpoints. S4: Combining multi-view neutral facial expression images, differentiable rendering technology is used to reverse optimize the shape parameters and textures of the parameterized 3D head model; by iteratively optimizing the projection error, the generated 3D head model is ensured to be consistent with the 2D images under different viewpoints, and high-precision geometry and high-fidelity textures are obtained.

2. The method for generating a 3D head model based on 2D facial expression analysis as described in claim 1, characterized in that, Step S1 also includes: The generated neutral expression image is re-detected by the expression recognition model. If the difference between the weight of any action unit and the neutral expression vector exceeds a preset threshold, the image is re-input into the expression editing model for iterative optimization until the generated image is close to a neutral expression.

3. The method for generating a 3D head model based on 2D facial expression analysis as described in claim 1 or 2, characterized in that, Step S2 specifically includes: An image segmentation algorithm is used to process the neutral expression image to generate a segmentation mask. This mask marks each pixel to identify the occluded area. Based on the labeling results of the segmentation mask, the color values ​​of pixels identified as occluders in the image are set to zero, and corresponding occluder masks are created. The occlusion mask and the image are input into a deep learning image generation model. The model's inpainting function is used to remove the occlusion and intelligently fill in the original occlusion position to maintain the continuity and consistency of facial texture.

4. The method for generating a 3D head model based on 2D facial expression analysis as described in any one of claims 1 to 2, characterized in that, Step S3 specifically includes: Using a human-image-generated-image model, neutral expression images from multiple angles are generated from unoccluded neutral expression images; By applying perspective transformation and geometric deformation techniques, facial images from multiple angles are generated from a standard image taken from a single angle. This ensures that images generated from different viewpoints retain neutral facial expression features and maintain the continuity and consistency of facial geometry and texture.

5. The method for generating a 3D head model based on 2D facial expression analysis as described in claim 4, characterized in that, The image generation model is based on Generative Adversarial Network (GAN) or Stable Diffusion.

6. The method for generating a 3D head model based on 2D facial expression analysis as described in claim 4, characterized in that, The facial images from multiple angles include perspectives from the left, right, top, bottom, and at 15, 30, and 45 degrees.

7. The method for generating a 3D head model based on 2D facial expression analysis as described in any one of claims 1 to 2, characterized in that, Step S4 specifically includes: Using the multi-view, neutral, and unobstructed 2D neutral expression images generated in step S3, the 3D parametric model is iteratively optimized using differentiable rendering technology to generate a 3D head model and texture. Randomly initialize the shape parameters and textures of the 3D model, extract vertex positions from the 3D head model, and adjust the vertex positions according to the shape parameters; The 3D model is rendered to a viewpoint corresponding to the multi-angle 2D image, and iterative optimization is performed to minimize the difference between the rendered image and the actual 2D image.

8. The method for generating a 3D head model based on 2D facial expression analysis as described in claim 7, characterized in that, The iterative optimization includes calculating a loss function, which is the sum of squares of the differences between the rendered image and the corresponding 2D image for all views, wherein the difference between the rendered result and the actual 2D image for each view is measured by Euclidean distance.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating a three-dimensional head model based on 2D facial expression analysis as described in any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating a three-dimensional head model based on 2D facial expression analysis as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for reconstructing a three-dimensional facial expression model based on a monocular video

    CN109584353A

  • 3D face model construction method

    US20100134487A1