Three-dimensional virtual image expression animation generation method based on realistic rendering technology

By using voice-driven deformation of the 3D virtual character mesh model and skin rendering model, combined with specular reflection, subsurface scattering and transmission components, highly realistic 3D virtual character facial animations are generated. This solves the problem of insufficient realism in skin rendering in existing technologies and achieves high-fidelity and beautiful facial animation generation.

CN115937387BActive Publication Date: 2026-02-03TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211703449.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-02-03
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing 3D virtual avatar facial animation generation technology neglects the simulation of the physical realism of human skin during the rendering process, resulting in an insufficient user experience.

Method used

A highly realistic 3D virtual avatar facial animation sequence with matching skin effect is generated by using a single face image and voice signal. The voice signal drives the deformation of the head mesh model, and the skin rendering model is combined with specular reflection, subsurface scattering and transmission components for rendering to generate highly realistic 3D virtual avatar facial animation.

Benefits of technology

It can generate high-fidelity, believable 3D virtual character facial animations without the need for professional artists. The animations are rich in detail, realistic and beautiful, and enhance the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937387B_ABST
    Figure CN115937387B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional virtual image expression animation generation method based on a real feeling rendering technology, and comprises the following steps: inputting a speech segment and a three-dimensional virtual image head grid model, driving the head grid model to deform by using a speech signal, and generating a three-dimensional virtual image grid model sequence corresponding to the speech signal and containing expressions; inputting a single face image, and generating a diffuse reflection map conforming to a topological structure of the three-dimensional virtual image head grid model; respectively acquiring a specular reflection component, a subsurface scattering component and a transmission component, and constructing a skin rendering model based on the specular reflection component, the subsurface scattering component and the transmission component; rendering the three-dimensional virtual image grid model sequence based on the skin rendering model, and outputting a three-dimensional virtual image expression animation sequence with a high real feeling skin effect. The application meets various needs in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of graphics and image processing technology, and in particular to a method for generating facial expression animations of three-dimensional virtual characters based on realistic rendering technology. Background Technology

[0002] Virtual avatars have been widely used in education, healthcare, military defense, and public services. Among them, three-dimensional virtual avatars, which have a higher degree of realism, have received common attention from the academic and industrial communities at home and abroad as an important support for emerging industries such as virtual reality and digital twins.

[0003] However, existing technical solutions mainly focus on the geometric structure and texture reconstruction of 3D virtual character models, ignoring the simulation of the physical realism of human skin during the rendering process, which limits the user experience. Summary of the Invention

[0004] This invention provides a method for generating 3D virtual character facial expression animations based on realistic rendering technology. Addressing the shortcomings of existing virtual character facial expression animation generation technologies, this invention generates a 3D virtual character facial expression animation sequence that matches the voice signal and includes highly realistic skin effects, using a single face image and a voice signal. This meets various needs in practical applications, as detailed below:

[0005] A method for generating facial expression animations of 3D virtual characters based on realistic rendering technology, the method comprising:

[0006] Input a speech segment and a 3D virtual avatar head mesh model, use the speech signal to drive the head mesh model to deform, and generate a sequence of 3D virtual avatar mesh models that correspond to the speech signal and include facial expressions;

[0007] Input a single face image and generate a diffuse texture that conforms to the topology of the 3D virtual avatar's head mesh model;

[0008] The specular reflection component, subsurface scattering component, and transmission component are obtained respectively, and a skin rendering model is constructed based on the specular reflection component, subsurface scattering component, and transmission component;

[0009] The three-dimensional virtual character mesh model sequence is rendered based on the skin rendering model to output a three-dimensional virtual character expression animation sequence with highly realistic skin effects.

[0010] The three-dimensional virtual image head mesh model includes: vertex coordinates, normal coordinates, texture coordinates, and triangle vertex indices.

[0011] Furthermore, the diffuse texture conforms to the topology of the 3D virtual avatar head mesh model and includes diffuse color information of the facial skin obtained from a single face image.

[0012] Wherein, obtaining the specular reflection component is:

[0013] The normal distribution function, Fresnel reflection function, and geometric masking function are calculated based on the half-angle vector, plane normal, incident ray direction, line of sight, and roughness. The specular reflection component is then calculated using the light source color.

[0014] Furthermore, the acquisition of the subsurface scattering component is as follows:

[0015] The diffusion profile is used to describe how light scatters and propagates in the skin; based on the Gaussian sum-fit algorithm, the diffusion profile is fitted with the sum of multiple Gaussian functions to calculate the subsurface scattered light distribution in the skin.

[0016] The subsurface scattering component is obtained by calculating the diffuse reflection color based on the subsurface scattering light distribution in the skin.

[0017] Wherein, the acquisition of the transmission component is:

[0018] The transmission component is calculated based on the incident direction of light, the normal direction, the diffuse map, and the light distribution in the skin described by the diffusion profile.

[0019] The beneficial effects of the technical solution provided by this invention are:

[0020] 1. This invention can obtain the mesh model of each frame of the virtual character's facial expression animation without the participation of professional artists and modelers, and retains the accurate facial animation that matches the voice segments;

[0021] 2. The present invention obtains a three-dimensional virtual image expression animation sequence with highly realistic skin effect. Compared with the use of simple texture mapping, the present invention considers more complex physical effects, presents richer details, and renders more realistic and beautiful skin effects.

[0022] In summary, this method can generate 3D virtual character facial expression animations and realistic skin rendering without the involvement of professional artists. It can generate high-fidelity 3D virtual character facial expression animations with believable skin effects using only voice clips, 3D virtual character head mesh models, and single face images. Attached Figure Description

[0023] Figure 1 A flowchart of a method for generating facial expression animations of 3D virtual characters based on realistic rendering technology;

[0024] Figure 2 This is a diagram illustrating the specular reflection effect of the skin in a more specific embodiment;

[0025] Figure 3 This is a diagram illustrating the subsurface scattering effect of the skin in a more specific embodiment;

[0026] Figure (a) shows the effect of subsurface scattering without using it, and Figure (b) shows the effect of subsurface scattering with it.

[0027] Figure 4 This is a diagram illustrating the skin translucency effect in a more specific embodiment.

[0028] Figure (a) shows the effect without using the translucency effect, and Figure (b) shows the effect with the translucency effect, illustrating the special orange-red appearance that appears on thin skin such as the nose and ears after the translucency is turned on.

[0029] Table 1 is a schematic diagram of a set of Gaussian function parameters that can be used to fit a diffusion profile in a more specific embodiment. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.

[0031] Example 1

[0032] A method for generating facial expression animations of 3D virtual characters based on realistic rendering technology, see [link to relevant documentation]. Figure 1 The method includes the following steps:

[0033] 101: Input a speech segment and a 3D virtual avatar head mesh model, use the speech signal to drive the head mesh model to deform, and generate a sequence of 3D virtual avatar mesh models that correspond to the speech signal and include facial expressions;

[0034] The 3D virtual avatar head mesh model file contains vertex coordinates, normal coordinates, texture coordinates, and triangle vertex indices.

[0035] 102: Input a single face image and generate a diffuse texture that conforms to the topology of the 3D virtual avatar's head mesh model;

[0036] The diffuse texture conforms to the topology of the 3D virtual character's head mesh model and contains diffuse color information of the facial skin obtained from a single face image.

[0037] 103: Construct a skin rendering model that includes specular reflection, subsurface scattering and transmission components. Based on this skin rendering model, render a sequence of three-dimensional virtual character mesh models to obtain a sequence of three-dimensional virtual character facial expression animations with highly realistic skin effects.

[0038] The skin rendering model includes a specular reflection component, which is calculated as follows:

[0039] The normal distribution function, Fresnel reflection function, and geometric masking function are calculated based on the half-angle vector, plane normal, incident ray direction, line of sight, and roughness. The specular reflection component is then calculated using the light source color.

[0040] The skin rendering model includes subsurface scattering components, which are calculated as follows:

[0041] Diffusion profiles are used to describe how light scatters and propagates in the skin;

[0042] Based on the Gaussian sum-fitting algorithm, the diffusion profile is fitted using the sum of multiple Gaussian functions to calculate the subsurface scattered light distribution and obtain the subsurface scattered light distribution in the skin.

[0043] The subsurface scattering component is obtained by calculating the diffuse reflection color based on the subsurface scattering light distribution in the skin.

[0044] The skin rendering model includes a transmission component, and its calculation method is as follows:

[0045] The transmission component is calculated based on the incident direction of light, the normal direction, the diffuse map, and the light distribution in the skin described by the diffusion profile.

[0046] Example 2

[0047] The following is combined with Figures 1-4 Table 1 and the calculation formulas further illustrate the scheme in Example 1, as detailed below:

[0048] 201: Input a speech segment and a 3D virtual avatar head mesh model, use the speech to drive the deformation of the head mesh model, and generate a sequence of 3D virtual avatar mesh models that correspond to the speech signal and include facial expressions;

[0049] It should be noted that, in the specific implementation of generating a sequence of three-dimensional virtual character mesh models through voice-driven generation, algorithms such as phoneme-visme mapping and deep neural networks can be used. This embodiment of the invention does not limit this, as long as it achieves the generation of a sequence of three-dimensional virtual character mesh models through voice-driven generation.

[0050] The type of the 3D virtual avatar head mesh model can be selected from blendshape models, parametric 3D head models, etc., and the embodiments of the present invention do not limit this.

[0051] As a preferred approach, the FLAME model (Faces Learned with an Articulated Model and Expressions) is used as the head mesh model for the 3D virtual avatar.

[0052] As a preferred approach, Voice Operated Character Animation (VOCA) is used to generate a sequence of three-dimensional virtual character mesh models that correspond to voice signals and include facial expressions.

[0053] 202: Input a single face image to generate a diffuse texture that conforms to the topology of the 3D virtual avatar's head mesh model;

[0054] In specific implementation, the single face image in the embodiments of the present invention can be an image that can be captured by face image acquisition devices such as cameras, smart terminals, and computers; the captured face image can also be an image obtained by electronic devices such as computers, such as a face image downloaded using computer devices, and the embodiments of the present invention do not limit this.

[0055] It should be noted that the embodiments of the present invention do not limit the method of generating diffuse texture, as long as the generated diffuse texture conforms to the topology of the three-dimensional head mesh model used.

[0056] As a preferred approach, the diffuse map generation method provided by the FLAME model is selected to obtain a diffuse map that conforms to the topological structure of the FLAME model from a single input face image.

[0057] 203: Construct a skin rendering model that includes specular reflection, subsurface scattering and transmission components. Based on this skin rendering model, render a sequence of three-dimensional virtual character mesh models to obtain a sequence of three-dimensional virtual character facial expression animations with highly realistic skin effects.

[0058] Specifically, step 203 includes:

[0059] To improve the realism of the final rendering result, physically based rendering methods can be used to calculate the specular reflection component of the skin, such as the Cook / Torrance BRDF model, the Ward BRDF model, the Kelemen / Szirmay-Kalos BRDF model, etc. This embodiment of the invention does not limit this, as long as the specular reflection component calculation adopts a physically based rendering method.

[0060] As a preferred method: the specular reflection component of the skin is calculated using the improved Kelemen / Szirmay-Kalos BRDF; in this step, the Kelemen / Szirmay-Kalos BRDF is defined by the following formula:

[0061]

[0062] Among them, P h (h) is the probability density of the direction of the half-angle vector being h; F(l,h) is the Fresnel reflection function; l is the direction vector from the shading point to the light source; 1 / (h·h) is the simplified geometric masking function in the Kelemen / Szirmay-Kalos mirror BRDF, which describes the self-occlusion phenomenon of each microplane in the microplane model.

[0063] As a preferred approach, the Beckmann distribution function is used for the normal distribution function of the microfacet.

[0064]

[0065] Where a is the angle between the macroscopic plane normal vector n and the half-angle vector h, and m represents the roughness of the plane.

[0066] As a preferred approach: to accelerate the calculation of the final skin texture, the pre-calculated result of the Beckmann function is exponentially scaled to within [0,1], and the processed result is saved as an 8-bit texture. During specular reflection calculation, samples are taken from the pre-calculated Beckmann texture, and the sampled values ​​are then used in the calculation after undergoing the appropriate inverse transformation.

[0067] As a preferred approach, the Fresnel reflection function uses the Schlick approximation equation:

[0068]

[0069] Where F0 represents the reflectivity when the light is incident normally, v represents the line-of-sight vector, l represents the light direction vector, and h represents the half-angle vector.

[0070] The effect of the skin specular reflection component obtained by the above method is as follows: Figure 2 As shown.

[0071] It should be noted that the subsurface scattering component can be calculated using texture space blurring method, path tracing subsurface scattering method, pre-integrated skin shading method, separable subsurface scattering method, screen space blurring method, etc. The embodiments of the present invention do not limit this, as long as the subsurface scattering component of the skin is calculated.

[0072] As a preferred approach, the subsurface scattering effect is calculated using a screen-space blurring method. In this step, a diffusion profile is fitted based on a Gaussian and fitting algorithm, and the diffusion profile is used to describe the subsurface scattering light distribution within the skin.

[0073] It should be noted that, due to the varying diffusion capabilities of different colors, when using Gaussian fitting algorithms to fit the diffusion profile, the relevant parameters of the fitting algorithm for each color channel need to be set. This allows for the use of Gaussian fitting algorithms corresponding to multiple color channels to fit the diffusion profile and obtain the light distribution within the skin. Due to the physical properties of skin, a single Gaussian function cannot accurately fit the complex light distribution of human skin. However, practice shows that using multiple Gaussian function distributions can more accurately approximate the diffusion profile and simulate the precise subsurface scattering effect of human facial skin.

[0074] In this embodiment of the invention, there is no limit to the number of Gaussian functions used, as long as the subsurface scattering components of the skin are simulated.

[0075] As a preferred approach, six Gaussian functions can be used to fit each diffusion profile:

[0076]

[0077] Where R(r) is the diffusion profile, w i The weights and coefficients v corresponding to each Gaussian function i Let be the variance, and r be the distance between the incident and exit points of the light ray.

[0078] As a preferred approach, the variance of the Gaussian function is defined as follows:

[0079]

[0080] It should be noted that the parameters and weights of the Gaussian function used are not limited, as long as the subsurface scattering and reflection components of the skin are simulated.

[0081] As a preferred approach, a set of Gaussian function parameters and weights for each Gaussian function used in the skin model are shown in Table 1.

[0082] Table 1

[0083]

[0084] The illumination distribution of a color channel can be obtained by summing multiple Gaussian functions. Similarly, the illumination distributions of multiple color channels can be obtained. Based on the Gaussian sum parameters defined in Table 1, the diffusion profile fitting result for the red channel can be defined (the green and blue channels require color-corresponding parameters): R represents the diffusion profile of the red channel fitted using the sum of six Gaussian functions, r represents the distance between the incident and exit points of the light, and G represents the Gaussian function.

[0085] R(r)=0.233*G(0.0064,r)+0.1*G(0.0484,r)+0.118*G(0.187,r)+0.113*G(0.567,r)+0.358*G(1.99,r)+0.078*G(7.41,r) (6)

[0086] The subsurface scattering component can be obtained by calculating the light distribution of the skin's subsurface scattering based on the above steps and the sampled colors in the diffuse reflection map.

[0087] The effect of obtaining the skin subsurface scattering component using the above method is as follows: Figure 3 As shown.

[0088] It should be noted that the transmission component calculation can use improved semi-transparent shadow mapping methods, screen space methods, etc., and the embodiments of the present invention do not limit this.

[0089] As a preferred approach, the transmission component is calculated using a screen-space method, defined by the following formula:

[0090]

[0091] Where E is the irradiance, R is the diffusion profile (calculated using the same method as the diffusion profile method for calculating the subsurface scattering component), r is the distance between the incident point and other points on the incident surface, and d is the medium thickness.

[0092] It should be noted that for thin-layered objects with a large area, the reverse normal of the current shading point can be used instead of the normal of the light incident point on the back of the object. Secondly, in thinner areas such as ears and the wings of the nose, the two sides maintain similar skin color, and the albedo on their surface will not change much. Therefore, the albedo of the front side can be used to estimate the irradiance on the back of the object.

[0093] As a preferred method, irradiance is defined by the following formula:

[0094] E(x,y)=E=a c max(-N c ·L,0.0) (9)

[0095] Where E is the irradiance, a c For albedo, N c Let L be the reverse normal of the shading point, and L be the direction of the ray. The diffusion profile is calculated using the same method as that used for calculating the subsurface scattering components:

[0096]

[0097] Where r is the distance between the incident point and other points on the incident surface, d is the thickness of the medium, k is the number of Gaussian functions used to fit the diffusion profile, and G is the Gaussian function.

[0098] The effect of the skin transmission component obtained by the above method is as follows: Figure 4 As shown, the feasibility of this method is verified, and it meets various needs in practical applications.

[0099] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0100] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating facial expression animations of 3D virtual characters based on realistic rendering technology, characterized in that, The method includes: Input a speech segment and a 3D virtual avatar head mesh model, use the speech signal to drive the head mesh model to deform, and generate a sequence of 3D virtual avatar mesh models that correspond to the speech signal and include facial expressions; Input a single face image and generate a diffuse texture that conforms to the topology of the 3D virtual avatar's head mesh model; The specular reflection component, subsurface scattering component, and transmission component are obtained respectively, and a skin rendering model is constructed based on the specular reflection component, subsurface scattering component, and transmission component; The three-dimensional virtual character mesh model sequence is rendered based on the skin rendering model to output a three-dimensional virtual character expression animation sequence with highly realistic skin effect; The diffuse texture conforms to the topological structure of the three-dimensional virtual image head mesh model and includes diffuse color information of the facial skin obtained from a single face image. The method for obtaining the specular reflection component is as follows: The normal distribution function, Fresnel reflection function, and geometric masking function are calculated based on the half-angle vector, plane normal, incident ray direction, line of sight, and roughness. The specular reflection component is then calculated using the light source color. The method for obtaining the subsurface scattering component is as follows: The diffusion profile is used to describe how light scatters and propagates in the skin; based on the Gaussian sum-fit algorithm, the diffusion profile is fitted with the sum of multiple Gaussian functions to calculate the subsurface scattered light distribution in the skin. The diffuse reflection color is calculated based on the distribution of subsurface scattered light in the skin to obtain the subsurface scattered component; The acquisition of the transmission component is as follows: The transmission component is calculated based on the incident direction of light, the normal direction, the diffuse map, and the light distribution in the skin described by the diffusion profile. The specular reflection component of the skin is calculated using the improved Kelemen / Szirmay-Kalos BRDF; the Kelemen / Szirmay-Kalos BRDF is defined by the following formula: Among them, P h (h) is the probability density of the direction of the half-angle vector being h; F(l,h) is the Fresnel reflection function; l is the direction vector from the shading point to the light source; 1 / (h·h) is the simplified geometric masking function in the Kelemen / Szirmay-Kalos mirror BRDF, which describes the self-occlusion phenomenon of each microplane in the microplane model. The normal distribution function of the microplane uses the Beckmann distribution function: Where α is the angle between the macroscopic plane normal vector n and the half-angle vector h, and m represents the roughness of the plane; The pre-calculated result of the Beckmann function is exponentially scaled to be limited to [0,1]. The processed calculation result is saved as an 8-bit texture. When performing specular reflection calculation, samples are taken from the pre-calculated Beckmann texture, and the sampled values ​​are used in the calculation after undergoing the corresponding inverse transformation. The Fresnel reflection function uses the Schlick approximation equation: Where F0 represents the reflectivity when the light is incident normally, v represents the line-of-sight vector, l represents the light direction vector, and h represents the half-angle vector; For each diffusion profile, six Gaussian functions were used for fitting: Where R(r) is the diffusion profile, w i The weights and coefficients v corresponding to each Gaussian function i Let r be the variance, and r be the distance between the incident and exit points of the ray. The variance of the Gaussian function is defined as follows: The parameters and weights of the Gaussian function used are not limited, as long as the subsurface scattering and reflection components of the skin are simulated. The transmission component is calculated using a screen-space method, defined by the following formula: Where E is the irradiance, R is the diffusion profile, calculated in the same way as the diffusion profile method for calculating the subsurface scattering component, r is the distance between the incident point and the other points on the incident surface, and d is the thickness of the medium. Irradiance is defined by the following formula: E(x,y)=E=αcmax(-Nc·L,0.0) Where, α c Let N be the albedo. c The reverse normal to the shading point is L, where L is the direction of the ray. The diffusion profile is calculated using the same method as that used for calculating the subsurface scattering components: Where k is the number of Gaussian functions used to fit the diffusion profile, and G is the Gaussian function.

2. The method for generating facial expression animations of a 3D virtual character based on realistic rendering technology according to claim 1, characterized in that, The 3D virtual avatar head mesh model includes: vertex coordinates, normal coordinates, texture coordinates, and triangle vertex indices.

Citation Information

Patent Citations

  • Method for realizing real-time face interaction animation based on monocular camera

    CN110599573A

  • Three-dimensional model rendering method and device, electronic equipment and storage medium

    CN114155335A