A 3D digital human face detail enhancement method and device based on neural network

Through the face generation model based on a generative adversarial network, the problem of low efficiency in generating high-definition and highly realistic and touching face images is solved, and the reality and clarity of faces in virtual reality scenes are improved.

CN115797389BActive Publication Date: 2025-07-11ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211376034.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-07-11
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

The prior art is inefficient in generating high-definition, high-realistic face images, and the faces generated in virtual reality scenes lack realism.

Method used

A face generation model based on a generative adversarial network is adopted to improve the clarity and reality of the face image through face detection, cutout, alignment, scaling, generation, background fusion and other steps.

Benefits of technology

It quickly generates high-definition and high-reality face images, solving the problem of lack of realism in virtual reality scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797389B_ABST
    Figure CN115797389B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for enhancing three-dimensional digital human facial details based on a neural network. A face detector is used to obtain a face image in an original RGB image, and a face generation model based on a generative adversarial network is used to convert the extracted face image into a clearer and more realistic face image. Subsequently, the face in the original image is removed, the background of the newly generated face is removed, and the newly generated face is fused with the original image after the face is removed to obtain an RGB image with a high sense of reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and computer graphics, and in particular to a method and device for enhancing three-dimensional digital human facial details based on a neural network. Background Art

[0002] Face image generation is an important issue in the fields of computer graphics and computer vision. Since NVIDIA proposed PGGAN, the generation of high-definition faces (1024*1024 resolution) has become a reality, with huge potential application spaces in games, virtual reality, computer graphics, etc.

[0003] In games, the fun of creating a unique character with one's own hands is self-evident. But sometimes it is often "five minutes of playing games, two hours of face sculpting". At present, most existing methods for creating and customizing game characters require players to manually adjust the facial features of the character in order to recreate their own face or sculpt someone else's face. A player usually needs to patiently manually adjust hundreds of parameters (such as face shape, eyes) for several hours to create a character similar to a specified portrait. Using state-of-the-art generative adversarial networks, such as StyleGAN, can generate highly realistic editable attribute faces within seconds, and at the same time allow players to generate simulated faces from provided photos within seconds.

[0004] The same technology can also be applied to virtual reality and video games, such as quickly replacing a large number of generated virtual faces lacking realism, facial reconstruction based on real face photos, or quickly generating realistic faces using hand-drawn sketches.

[0005] In the field of computer graphics, the process of 3D modeling requires a large amount of resources. Using a face generation network can accelerate the acquisition of 2D images before modeling. Currently, some advanced networks, such as MeInGame, can reconstruct a 3D face model based on a 2D face and then transfer it to the corresponding template network, thereby accelerating the construction of the 3D model. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for enhancing three-dimensional digital human facial details based on a neural network in view of the deficiencies of the prior art.

[0007] The purpose of the present invention is achieved through the following technical solutions: A method for enhancing three-dimensional digital human facial details based on a neural network, comprising the following steps:

[0008] S1, obtaining an original RGB image;

[0009] S2. Use a face detector to detect the original RGB image. If a face image is detected in the original RGB image, extract N square-shaped face images I1…I from the original RGB image: i …I N , where I i represents the i-th face image, i = 1,…i…N, N≥1; obtain the set of boundary vertex coordinates P i of each face image I i , where P i represents the set of boundary vertex coordinates of the face image I i ; the P i = {p i,1 , p i,2 , p i,3 , p i,4}, where p i,1 represents the upper-left vertex coordinate of the face image I i , p i,2 represents the upper-right vertex coordinate of the face image I i , p i,3 represents the lower-right vertex coordinate of the face image I i , p i,4 represents the lower-left vertex coordinate of the face image I i ;

[0010] S3. Use a face matting algorithm to filter the background of any face image I i extracted in step S2, and obtain a face image mask I i,1 and a face image background I i,2 ;

[0011] S4. Use a face alignment algorithm to process the obtained face image mask I i,1 to obtain an aligned face image I i,3 ;

[0012] S5. Scale the face image I i,3 to a size of 1024*1024 to obtain a face image I i,4 ;

[0013] S6. Input the face image I i,4 into a face generation model based on a generative adversarial network, and obtain a feature vector W i corresponding to it in the latent code space through the encoder of the face generation model;

[0014] S7. Use the decoder of the face generation model to convert the obtained feature vector W i into a face image I i,5 ;

[0015] S8, scale the face image I i,5 to the same size as the face image I i,3 to obtain the face image I i,6 ;

[0016] S9, use the face matting algorithm to remove the background of the face image I i,6 to obtain the face image I i,7 ;

[0017] S10, use the edge fusion algorithm to fuse the face image I i,7 and the face image background map I i,2 to obtain the face image I i,8 ;

[0018] S11, repeat steps S3 to S10 to obtain each face image I i of the face image I i,8 ;

[0019] S12, according to the set P of boundary vertex coordinates of each face image I i , use the face image I i to cover the face image I i,8 in the original RGB image to obtain a new RGB image. i Further, the face detector is a face detector in the dlib library or a face detector in the opencv library.

[0020] Furthermore, the generative adversarial network includes a generative model and a discriminative model; the generative model includes an encoder and a decoder.

[0021] The present invention also provides a three-dimensional digital human face detail enhancement device based on a neural network, including one or more processors for implementing the above-mentioned three-dimensional digital human face detail enhancement method based on a neural network.

[0022] The present invention also provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it is used to implement the above-mentioned three-dimensional digital human face detail enhancement method based on a neural network.

[0023]

[0024] ​The beneficial effects of the present invention are as follows: For any picture that can detect a human face and has deficiencies in the human face, the realism and clarity of the human face can be improved; for areas where a human face cannot be detected but is required, such as when the human face in the original picture becomes blurred for some reason, a new human face can be generated using a neural network, or the area in the original picture can be replaced with a given clear human face, so as to add a new realistic human face to the original picture. For example, in a virtual reality scenario, the human faces generated by software often lack realism, and this technology can solve this problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of a method for enhancing the facial details of a 3D digital human based on a neural network;

[0026] Figure 2 It is a schematic diagram of an original RGB image;

[0027] Figure 3 It is a schematic diagram of a human face image I i ;

[0028] Figure 4 It is a schematic diagram of a human face image I i,6 ;

[0029] Figure 5 It is a schematic diagram of a human face image I i,7 ;

[0030] Figure 6 It is a schematic diagram of a new RGB image;

[0031] Figure 7 It is a schematic diagram of a device for enhancing the facial details of a 3D digital human based on a neural network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] To address the problem of the inability to quickly generate high-definition and highly realistic face images mentioned above, the present invention proposes a method for enhancing the facial realism and clarity of 3D digital humans in a virtual scene using a neural network. A face generation model is used, which is based on a generative adversarial network with an encoder. This network is a deep learning model that includes a generative model and a discriminative model, incorporating the idea of game theory. The generation quality of the model is improved through the continuous confrontation between the generative model and the discriminative model. The quality of the images finally generated by the model depends to a large extent on the data quality and training techniques. The generative model used in this method takes and outputs RGB images of 1024*1024, including an encoder and a decoder. The encoder can convert the input image into a 512-dimensional real number vector, and adjusting on this vector can affect the image attributes output by the decoder. The decoder can convert this real number vector into an RGB image of 1024*1024.

[0034] Embodiment 1

[0035] As Figure 1 shown, the present invention provides a method for enhancing the facial details of 3D digital humans based on a neural network, including the following steps:

[0036] S1. Obtain an original RGB image, as Figure 2 shown.

[0037] S2. Use a face detector to detect the original RGB image. If a face image is detected in the original RGB image, extract N square-shaped face images I1…I i …I N in the original RGB image, where I i represents the i-th face image, i = 1,…i…N, N≥1; obtain the set of boundary vertex coordinates P i of each face image I i through the face detector, where P i represents the set of boundary vertex coordinates of the face image I i ; the P i ={p i,1 ,p i,2 ,p i,3 ,p i,4}, where p i,1 represents the upper left vertex coordinate of the face image I i , p i,2 represents the upper right vertex coordinate of the face image I i , p i,3 represents the lower right vertex coordinate of the face image I i , p i,4 represents the lower left vertex coordinate of the face image I iThe lower left vertex coordinates; as Figure 3 shown, this figure is a square-shaped face image extracted from Figure 2 , which is face image I i .

[0038] The face detector is the face detector of the dlib library or the face detector of the opencv library.

[0039] S3. Use the face matting algorithm (matting network) to filter the background of any face image I extracted in step S2 i to obtain the face image mask I i,1 and the face image background map I i,2 .

[0040] The face image mask and the face image background map obtained by the face matting algorithm are both RGBA images containing an alpha channel.

[0041] The purpose of this step is to remove the background in the face image, making it cleaner when input into the face generation model, and improving the generation effect of the generation model.

[0042] S4. Use the face alignment algorithm (alignment) to process the obtained face image mask I i,1 to obtain the aligned face image I i,3 .

[0043] The significance of alignment is to make the face more centered in the image, adjust the tilt angle of the face, making the face present a frontal view of the screen and the angle tend to be horizontal.

[0044] S5. Scale the face image I i,3 to a size of 1024 * 1024 to obtain the face image I i,4 .

[0045] S6. Input the face image I i,4 into the face generation model based on the generative adversarial network, and obtain the corresponding feature vector W in the latent code space through the encoder of the face generation model i .

[0046] The generative adversarial network includes a generation model and a discriminant model; the generation model includes an encoder and a decoder.

[0047] The generative model is a new model fine-tuned based on the open-source model pixel2style2pixel. The new model uses an Asian face dataset, which is a set of Asian face pictures screened from the open-source face dataset FFHQ. The new model has a better decoding effect on Asian-style face images. At the same time, the generated face images also better conform to the facial features of the Asian race. The network structure of Pixel2style2pixel is divided into an encoder and a decoder. The input of the encoder is an RGB image of 1024*1024*3, and the output is a latent code of 18*512 dimensions, with a total of 18 layers. Its backbone is of the feature pyramid type and adopts the skip structure of ResNet, including three feature maps of different sizes, named small, medium, and largest respectively. Each feature map generates multiple vectors. The small one corresponds to the vectors of layers 0-2, the medium one corresponds to layers 3-6, and the largest one corresponds to layers 7-18. Each layer of vectors is transformed into a 1*512 vector through the map2style network (where map2style is a multi-layer fully convolutional network with a stride of 2 and a LeakyReLU activation function). After affine transformation and combination into an 18*512-dimensional latent code, it is input into the decoder. The decoder is the Synthesis Network of the open-source model StyleGAN2, which consists of 18 synthesis blocks and 1 ToRGB layer. Among them, a synthesis block consists of 2 style blocks and 1 ToRGB layer. A style block includes 1 fully connected layer, 1 convolutional layer, 1 activation function layer, and a learnable bias; the ToRGB layer consists of 1 fully connected layer, 1 convolutional layer, 1 activation function layer, and a learnable bias.

[0048] S7, use the decoder of the face generation model to convert the obtained feature vector W i into a face image I i,5 . The face image I i,5 is similar to the face image I i,4 but is clearer and more realistic than the face image I i,4 .

[0049] S8, scale the face image I i,5 to the same size as the face image I i,3 to obtain the face image I i,6 , as Figure 4 shown.

[0050] S9, use a face matting algorithm to process the face image I i,6Background removal to obtain the face image I i,7 The face image I i,7 is an RGBA image, as Figure 5 shown.

[0051] S10. Use the edge fusion algorithm to fuse the face image I i,7 and the face image background map I i,2 to obtain the face image I i,8 . The face image I i,8 contains both the background of the face image I i , and the face in the face image I i,8 is clearer than the face in the face image I i ;

[0052] S11. Repeat steps S3 - S10 to obtain the face image I i for each face image I i,8 .

[0053] S12. According to the set of boundary vertex coordinates P i of each face image I i , use the face image I i,8 to cover the face image I i in the original RGB image to obtain a new RGB image, as Figure 6 shown.

[0054] Embodiment 2

[0055] Refer to Figure 7 . An apparatus for enhancing the facial details of a 3D digital human based on a neural network provided by an embodiment of the present invention includes one or more processors for implementing the method for enhancing the facial details of a 3D digital human based on a neural network in the above embodiment.

[0056] Embodiments of the apparatus for enhancing the facial details of a 3D digital human based on a neural network of the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The apparatus embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non - volatile memory into the memory for running. From the hardware level, as Figure 7 shown, it is a hardware structure diagram of any device with data processing capabilities where the apparatus for enhancing the facial details of a 3D digital human based on a neural network of the present invention is located. Except for Figure 7In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0057] For the specific implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0058] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0059] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for enhancing the facial details of a 3D digital human based on a neural network in the above embodiment.

[0060] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by the device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for enhancing three-dimensional digital human facial details based on a neural network, characterized in that, including the following steps: S1, obtaining an original RGB image; S2, use the face detector to detect the original RGB image. If a face image is detected in the original RGB image, extract N square face images in the original RGB image: I1…I i …I N , where I i represents the i-th face image, i=1,…i…N, N≥1; each face image I is obtained by the face detector i The boundary vertex coordinate set P i , where P i Represents a face image I i The boundary vertex coordinate set of P i ={p i,1 ,p i,2 ,p i,3 ,p i,4 }, where p i,1 Represents a face image I i The coordinates of the upper left vertex, p i,2 Represents a face image I i The coordinates of the upper right vertex, p i,3 Represents a face image I i The coordinates of the lower right vertex, p i,4 Represents a face image I i The coordinates of the lower left vertex; S3. Use a face matting algorithm to process any face image I extracted in step S2 i to filter the background and obtain a face image mask I i,1 and a face image background map I i,2 ; S4, use a face alignment algorithm to process the obtained face image mask I i,1 to obtain the aligned face image I i,3 ; S5, scale the face image I i,3 to a size of 1024*1024 to obtain the face image I i,4 ; S6. Input the face image I i,4 into the face generation model based on the generative adversarial network, and obtain the corresponding feature vector W in the latent code space through the encoder of the face generation model i ; S7, use the decoder of the face generation model to convert the obtained feature vector W i into a face image I i,5 ; S8, scale the face image I i,5 to the same size as the face image I i,3 to obtain the face image I i,6 ; S9. Use the face matting algorithm to remove the background of the face image I i,6 to obtain the face image I i,7 ; S10, use the edge fusion algorithm to fuse the face image I i,7 and the background image I of the face image i,2 to obtain the face image I i,8 ; S11. Repeat steps S3 - S10 to obtain each face image I i of the face image I i,8 ; S12. According to the set of boundary vertex coordinates P i of each face image I i , use the face image I i,8 to cover the face image I in the original RGB image i , and obtain a new RGB image.

2. The three-dimensional digital human face detail enhancement method based on a neural network according to claim 1, wherein The face detector is a face detector of the dlib library or a face detector of the opencv library.

3. A method for enhancing the facial details of a 3D digital human based on a neural network according to claim 1, characterized in that, The generative adversarial network includes a generative model and a discriminative model; the generative model includes an encoder and a decoder.

4. A three-dimensional digital human face detail enhancement device based on a neural network, characterized in that, including one or more processors for implementing a three-dimensional digital human face detail enhancement method according to any one of claims 1-3.

5. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it is used to implement a three-dimensional digital human face detail enhancement method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Image generation method and device

    CN111275784A

  • Image enhancement method and device of face image, equipment and storage medium

    CN114862716A