Virtual image expression generation method and device, computer device, and storage medium

CN117475042BActive Publication Date: 2026-09-15GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210833670.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2026-09-15
Estimated Expiration
2042-07-14

AI Technical Summary

Technical Problem

[0004]本发明实施例提供一种虚拟形象的表情生成方法、装置、计算机设备和存储介质,以改善现有虚拟形象的表情不够生动的问题

Benefits of technology

[0015] The virtual avatar expression generation method, apparatus, computer device, and storage medium provided in this invention pre-adjust the initial expression base based on the semantic error between the reference expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thereby obtaining an expression base. This makes the expression base more closely match the sample face image and improves the realism of the expression base. Thus, after obtaining the target face image and the target expression coefficient corresponding to the target face image, the target expression base corresponding to the target expression coefficient in the face expression base data can be directly determined based on the target expression coefficient. Then, based on the target expression coefficient and the target expression base, the expression of the virtual avatar corresponding to the target face image is generated, making the expression of the virtual avatar more vivid and realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117475042B_ABST
    Figure CN117475042B_ABST
Patent Text Reader

Abstract

The application discloses a virtual image expression generation method and device, computer equipment and a storage medium, relates to the technical field of computer vision, and adjusts an initial expression base according to the semantic error between a reference expression encoding vector of a sample face image and a conversion expression encoding vector corresponding to the initial expression base of the sample face image, so that the expression base is more suitable for the sample face image, the expression base is improved in fidelity, and thus, after a target face image and a target expression coefficient corresponding to the target face image are acquired, the target expression base corresponding to the target expression coefficient in the face expression base data can be directly determined according to the target expression coefficient, and then, according to the target expression coefficient and the target expression base, the expression of the virtual image corresponding to the target face image is generated, so that the expression of the virtual image is more vivid and lifelike.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a method, apparatus, computer device, and storage medium for generating facial expressions for virtual avatars. Background Technology

[0002] Facial morphing is a commonly used method for compositing facial expressions. This method linearly fuses a predefined set of facial bases that cover most facial details to create a new expression. Facial morphing has been widely applied in fields such as 3D animation and film production.

[0003] In facial deformation, an important task is to bind facial expressions to expression bases. Existing automatic binding of expression bases is based on deformation transfer. However, simply using deformation transfer can lead to a low degree of fit between expression bases and facial images, resulting in less vivid expressions of virtual characters driven by expression bases. Summary of the Invention

[0004] This invention provides a method, apparatus, computer device, and storage medium for generating facial expressions for virtual avatars, in order to improve the problem that the facial expressions of existing virtual avatars are not vivid enough.

[0005] In a first aspect, embodiments of the present invention provide a method for generating facial expressions for a virtual avatar, the method comprising:

[0006] Obtain the target face image to be converted, and the target expression coefficient corresponding to the target face image;

[0007] A target expression base corresponding to the target expression coefficient is determined in the facial expression base data; the facial expression base data includes multiple expression bases and expression coefficients corresponding to each expression base, and the expression base is obtained by adjusting the initial expression base based on the semantic error between the baseline expression encoding vector of the sample facial image and the transformed expression encoding vector of the initial expression base corresponding to the sample facial image;

[0008] Based on the target expression base, generate the expression of a virtual avatar corresponding to the target face image.

[0009] Secondly, according to an embodiment of the present invention, a virtual avatar expression generation device is provided, the device comprising:

[0010] The acquisition module is used to acquire the target face image to be converted, and the target expression coefficients corresponding to the target face image;

[0011] The expression base determination module is used to determine the target expression base in the facial expression base data that corresponds to the target expression coefficient; the facial expression base data includes multiple expression bases and expression coefficients corresponding to each expression base, and the expression base is obtained by adjusting the initial expression base based on the semantic error between the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image;

[0012] The expression generation module is used to generate expressions of a virtual image corresponding to the target face image based on the target expression base.

[0013] Thirdly, embodiments of the present invention provide a computer device, the computer device including a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to perform the operations in the virtual avatar expression generation method.

[0014] Fourthly, embodiments of the present invention provide a storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in the virtual avatar expression generation method.

[0015] The virtual avatar expression generation method, apparatus, computer device, and storage medium provided in this invention pre-adjust the initial expression base based on the semantic error between the reference expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thereby obtaining an expression base. This makes the expression base more closely match the sample face image and improves the realism of the expression base. Thus, after obtaining the target face image and the target expression coefficient corresponding to the target face image, the target expression base corresponding to the target expression coefficient in the face expression base data can be directly determined based on the target expression coefficient. Then, based on the target expression coefficient and the target expression base, the expression of the virtual avatar corresponding to the target face image is generated, making the expression of the virtual avatar more vivid and realistic. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a structural block diagram of the computer device provided in the embodiments of the present invention;

[0018] Figure 2 This is a flowchart illustrating the method for generating facial expressions for virtual avatars provided in an embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of a target facial expression base provided in an embodiment of the present invention;

[0020] Figure 4 This is a schematic diagram of the structure of the virtual image expression generation device provided in an embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0024] As mentioned in the background, current virtual avatar generation methods rely on deformation transfer to bind faces to facial expression bases. However, simply using deformation transfer often results in subtle discrepancies between the expression base and the facial image, leading to a larger error between the generated expression base and the input facial image, making it less realistic and affecting the expression-driven effect of facial animation. Therefore, to improve the matching degree between expression bases and facial images, existing technologies mainly rely on extensive manual fine-tuning. However, this heavily depends on the modeler's experience, requiring significant time and increasing costs.

[0025] Based on this, embodiments of the present invention provide a method for generating expressions for virtual avatars. This method encodes facial images and corresponding initial expression bases, and obtains a target expression base based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector. The expression encoding vector is used to measure the similarity between the initial expression base and the facial image, thereby enabling better adjustment of the target expression base, making the target expression base more consistent with the facial image, and thus making the expressions of the virtual avatar more vivid, avoiding the need for designers to make extensive fine-tuning adjustments in the later stages.

[0026] The virtual avatar expression generation method provided in this invention can be applied to various scenarios of virtual avatar generation and 3D animation production, such as news broadcasting, weather forecasting, game commentary, and game scenes that allow the creation of game characters with faces identical to the user's own. It can also be used in scenarios where virtual avatars provide personalized services, such as one-on-one services like psychologists and virtual assistants. Furthermore, it can be applied to animation production scenarios, such as 3D animation production. In these scenarios, the method provided in this invention can determine the target expression base of a target facial image, construct a facial model using the target expression base, and then drive the virtual avatar based on this facial model and expression base.

[0027] To facilitate understanding of the technical solution of the present invention, the method for generating facial expressions of virtual characters provided in the embodiments of the present invention will be introduced below in conjunction with actual application scenarios.

[0028] Please see Figure 1 , Figure 1 This is a structural block diagram of a computer device provided in an embodiment of the present invention. The computer device 100 may include a virtual avatar expression generation device 10, a memory 20, a processor 30, and a communication unit 40. The memory 20 stores machine-readable instructions that can be executed by the processor 30. When the computer device 100 is running, the processor 30 and the memory 20 communicate via a bus. The processor 30 executes the machine-readable instructions and performs the virtual avatar expression generation method.

[0029] The memory 20, processor 30, and communication unit 40 are electrically connected directly or indirectly to each other to achieve signal transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The virtual avatar expression generation device 10 includes at least one software function module that can be stored in the memory 20 in the form of software or firmware. The processor 30 is used to execute the executable module (e.g., the software function module or computer program included in the virtual avatar expression generation device 10) stored in the memory 20.

[0030] The memory 20 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0031] In some embodiments, processor 30 is used to perform one or more functions described in this embodiment. In some embodiments, processor 30 may include one or more processing cores (e.g., a single-core processor (S) or a multi-core processor (S)). By way of example only, processor 30 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction-set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computing (RISC) computer, or a microprocessor, or any combination thereof.

[0032] For ease of explanation, only one processor is described in computer device 100. However, it should be noted that computer device 100 in this embodiment may also include multiple processors, and therefore the steps performed by one processor as described in this embodiment may also be performed jointly or individually by multiple processors. For example, if the server's processor performs steps A and B, it should be understood that steps A and B may also be performed jointly by two different processors or individually by one processor. For example, one processor performs step A, and a second processor performs step B, or the first and second processors jointly perform steps A and B.

[0033] In this embodiment, the memory 20 is used to store the program, and the processor 30 is used to execute the program after receiving the execution instruction. The process definition method disclosed in any implementation of this embodiment can be applied to the processor 30, or implemented by the processor 30.

[0034] The communication unit 40 is used to establish a communication connection between the computer device 100 and other devices via a network, and to send and receive data via the network.

[0035] In some implementations, the network can be any type of wired or wireless network, or a combination thereof. By way of example only, the network may include wired networks, wireless networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, or near field communication (NFC) networks, or any combination thereof.

[0036] In this embodiment, the computer device 100 may be, but is not limited to, a laptop computer, a mobile terminal, a personal computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and other computer devices. This embodiment does not impose any restrictions on the specific type of computer device.

[0037] Understandably, Figure 1The structure shown is for illustrative purposes only. The computer device 100 may also have... Figure 1 Showing more or fewer components, or having with Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0038] based on Figure 1 The implementation architecture of this invention provides a method for generating expressions for virtual avatars, which is based on the following: Figure 1 The computer device shown performs the following based on Figure 1 The structural diagram of the computer device shown illustrates in detail the steps of the virtual avatar expression generation method provided in this embodiment. Please refer to... Figure 2 The virtual avatar expression generation method provided in this embodiment of the invention includes steps S201 to S203:

[0039] Step S201: Obtain the target face image to be converted, and the target expression coefficient corresponding to the target face image.

[0040] The facial expression coefficient represents the state of a person's face, including but not limited to opening the mouth, closing the eyes, puffing out the cheeks, and squinting.

[0041] In some embodiments of the present invention, there are multiple ways to obtain the target expression coefficients corresponding to the target face image, including, for example:

[0042] (1) It can acquire facial key points in the target face image, identify the expression corresponding to the target face image based on the facial key points, and obtain the target expression coefficient corresponding to the target face image by querying the preset expression data based on the expression corresponding to the target face image. Among them, the preset expression data stores multiple face expressions and the expression coefficient corresponding to each face expression.

[0043] (2) The expression recognition model can be used to identify the target face image to be transformed, obtain the expression corresponding to the target face image, and query the preset expression data based on the expression corresponding to the target face image to obtain the target expression coefficient corresponding to the target face image. Among them, the expression recognition model can be a neural network-based expression recognition model, such as one based on Convolutional Neural Networks (CNN), De-Convolutional Networks (DN), Deep Neural Networks (DNN), Deep Convolutional Inverse Graphics Networks (DCIGN), Region-based Convolutional Networks (RCNN), Faster Region-based Convolutional Networks (Faster RCNN), and Bidirectional Encoder Representations from Transformers (BERT) model;

[0044] (3) The facial expression coefficient recognition model can be used to identify the facial expression coefficients of the target face image and obtain the target facial expression coefficients corresponding to the target face image. Among them, the facial expression coefficient recognition model can use a lightweight network to ensure the efficiency of face recognition. For example, considering the efficiency of the model, the facial expression coefficient recognition model can adopt a network structure such as MobileNet v2.

[0045] It should be noted that the above-described method for obtaining the target facial expression coefficient corresponding to the target face image is merely an illustrative example, and the embodiments of the present invention do not limit the method for obtaining the target facial expression coefficient corresponding to the target face image.

[0046] In some embodiments of the present invention, the target face image to be converted can be a face image uploaded by the user, such as a face image captured by a camera in a live broadcast scene, or a face image captured during image shooting; the target face image to be converted can also be a created face image, such as a pre-created face image in 3D animation production.

[0047] In some embodiments of the present invention, in order to make the expressions of the generated virtual image more vivid, in the acquisition of target expression coefficients, data normalization, facial key point detection and image cropping can be performed on the target face image to be converted, the background in the target face image is removed, a portrait image is obtained from the target face image, and expression coefficient recognition is performed on the portrait image to obtain the target expression coefficients corresponding to the target face image.

[0048] Step S202: Determine the target expression base in the facial expression base data that corresponds to the target expression coefficient.

[0049] The facial expression base data includes multiple expression bases and corresponding expression coefficients for each base. The expression bases are obtained by adjusting the initial expression base based on the semantic error between the baseline expression encoding vector of the sample facial image and the transformed expression encoding vector of the initial expression base corresponding to the sample facial image. The expression encoding vector is used to characterize the expression semantics of the facial image, and the distance error is used to quantify the similarity between the facial image and the initial expression base corresponding to the facial image.

[0050] In some embodiments of the present invention, the target expression base can be a single expression base or multiple expression bases. Each target expression base is used to identify an expression, such as whether the mouth is open, how wide the mouth is open, whether the eyes are squinting, and the angle of squinting. It is understood that in embodiments of the present invention, there can be multiple target expression bases, and multiple sets of target expression coefficients are applied, with each set of target expression coefficients corresponding to one target expression base.

[0051] In some embodiments of the present invention, the target expression base can be a general human face expression base or an expression base corresponding to a virtual avatar. The following explanation uses a general human face expression base as an example. Figure 3 As shown, Figure 3 This is a schematic diagram of a target expression base provided in an embodiment of the present invention. Each target expression base has the same three-dimensional mesh structure, wherein the three-dimensional mesh structure is a 3D model storage format containing fixed points and edges. Each point and edge corresponds to a semantic information, such as the xxth key point representing the left corner of the eye.

[0052] In some embodiments of the present invention, step S202 includes: querying pre-stored facial expression base data; comparing the target expression coefficient with the expression coefficients in the pre-stored facial expression base data; setting the expression coefficients in the pre-stored facial expression base data that are consistent with or similar to the target expression coefficient as the expression coefficients in the pre-stored facial expression base data that match the target expression coefficient; and setting the expression base corresponding to the expression coefficients in the pre-stored facial expression base data that match the target expression coefficient as the target expression base corresponding to the target expression coefficient. The pre-stored facial expression base data includes multiple expression bases and expression coefficients corresponding to each expression base. Expression coefficients similar to the target expression coefficient can be expression coefficients where the difference between the target expression coefficient and the expression coefficient is less than or equal to a preset difference threshold.

[0053] In some embodiments of the present invention, step S202 includes: identifying the target expression coefficient to obtain the target expression base corresponding to the target expression coefficient. The target expression coefficient can be identified using a preset recognition model to obtain the target expression base corresponding to the target expression coefficient, wherein the model can be a machine learning model, a probabilistic model, a classification model, etc.

[0054] Step S203: Generate the expression of the virtual image corresponding to the target face image based on the target expression base.

[0055] In some embodiments of the present invention, when multiple target expression bases exist, step S203 includes: combining the target expression bases according to the target expression coefficients to generate the expression of a virtual image corresponding to the target face image. Specifically, combining the target expression bases corresponding to the target expression coefficients according to the target expression coefficients to generate the expression of a virtual image corresponding to the target face image. Combining the target expression bases corresponding to the target expression coefficients according to the target expression coefficients includes: adjusting the target expression bases corresponding to the target expression coefficients to obtain adjusted target expression bases, and combining each adjusted target expression base to generate the expression of a virtual image corresponding to the target face image. For example, adjusting the target expression bases corresponding to the target expression coefficients according to the strabismus angle of the eyes, the upward angle of the mouth, etc., in the target expression coefficients.

[0056] In some embodiments of the present invention, when there are multiple target expression bases, step S203 includes: setting the target expression coefficient as the weight of the corresponding target expression base, superimposing according to the weight of each target expression base, driving the target face image onto the corresponding virtual image, and generating the expression of the virtual image corresponding to the target face image.

[0057] In some embodiments of this application, when the target expression base is single, step S203 includes: driving the target expression base according to the target expression coefficient to generate the expression of a virtual image corresponding to the target face image.

[0058] In some embodiments of this application, when the target expression base is single, step S203 includes: generating a face model corresponding to the target face image based on the target expression base, determining a virtual image based on the target expression base and the face image, and generating an expression of the virtual image corresponding to the target face image.

[0059] In this embodiment of the invention, the initial expression base is pre-adjusted by the semantic error between the base expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thus obtaining the expression base. This makes the expression base more consistent with the sample face image and improves the realism of the expression base. In this way, when the target face image to be converted and the target expression coefficient corresponding to the target face image are obtained, the target expression base corresponding to the target expression coefficient can be directly determined. Based on the target expression coefficient and the target expression base, the expression of the virtual image corresponding to the target face image is generated, which makes the face image and the expression base more consistent, and thus makes the expression of the virtual image generated based on the target expression base more vivid and realistic.

[0060] In some embodiments of the present invention, in order to improve the vividness of the virtual avatar's expressions, it is necessary to improve the matching degree between the sample face images and the expression base. In the binding of sample face images and expression bases, the present invention adjusts the expression base according to the similarity between each sample face image in the sample face image set and the expression base, thereby ensuring the matching degree between the sample face images and the expression base, and thus improving the vividness of the virtual avatar's expressions. Specifically, the method for binding sample face images and expression bases includes steps a1 to a4:

[0061] Step a1: Obtain the sample face image set, as well as the initial expression basis and expression coefficients corresponding to each sample face image in the sample face image set.

[0062] Step a2: For each sample face image, perform expression encoding on the sample face image and the initial expression base corresponding to the sample face image to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image.

[0063] Step a3: Adjust the initial expression base according to the semantic error between the baseline expression encoding vector and the transformed expression encoding vector to obtain the expression base corresponding to the sample face image.

[0064] Step a4: Associate the expression base and expression coefficients corresponding to each sample face image to obtain the face expression base data.

[0065] The sample face image set can be a pre-collected training image set, which includes multiple sample face images, each corresponding to a different expression. In some embodiments of the present invention, the method of acquiring the sample face image set is not limited. For example, the sample face image set can be acquired by an image acquisition device, such as sample face images with different expressions extracted from a captured video.

[0066] The initial expression base is obtained by transferring expressions from a template expression base. The template expression base can be a pre-stored general expression base for a virtual avatar. In some embodiments of this invention, the template expression base can be created by an artist or obtained from an open-source website. This invention does not limit the method of obtaining the template expression base.

[0067] In some embodiments of the present invention, the expression coefficients corresponding to each sample face image can be obtained according to the target expression coefficient acquisition method in step S201.

[0068] In some embodiments of the present invention, the expressions in a sample face image may be mapped to a template expression base to obtain an initial expression base corresponding to the sample face image. Specifically, the method for obtaining the initial expression base corresponding to the sample face image through expression transfer includes:

[0069] (1) Perform face detection on each sample face image in the sample face image set to obtain the location of facial feature points in each sample face image.

[0070] (2) For each sample face image, project the position of the facial feature points in the sample face image to obtain the projection result of the sample face image.

[0071] (3) Fit the projection result of the sample face image to the position of the facial feature points in the pre-stored template expression base to obtain the expression data corresponding to the sample face image.

[0072] (4) Adjust the template expression base according to the expression data corresponding to the sample face image to obtain the initial expression base corresponding to the sample face image.

[0073] The projection result can be the projection of the positions of facial feature points in the sample face image onto a pre-stored template face model, or it can be the projection of the positions of facial feature points in the sample face image onto a two-dimensional space. The template face model includes neutral expressions and different template expression bases. This embodiment of the invention does not limit the template face model; for example, the template face model can be a Basel face model.

[0074] The number of facial feature points in each sample face image can be 68. The positions of the 68 facial feature points in each sample face image can be calculated using a trained 69 feature point alignment model, or the position of the 68 facial feature points in each sample face image can be output by a trained feature point detection model.

[0075] In some embodiments of the present invention, when the positions of facial feature points in the sample face image are projected onto a two-dimensional space, a projection matrix is ​​calculated based on the positions of facial feature points in the sample face image and the positions of facial feature points in a pre-stored template face model. The template expression base in the pre-stored template face model is then mapped onto the two-dimensional space based on the projection matrix to obtain the projection result corresponding to the sample face image.

[0076] In some embodiments of the present invention, when the positions of facial feature points in the sample face image are projected onto a pre-stored template face model, a projection matrix is ​​calculated based on the positions of facial feature points in the sample face image and the positions of facial feature points in the pre-stored template face model, and the projection matrix is ​​set as the projection result corresponding to the sample face image.

[0077] The expression data corresponding to the sample face image can be the weight coefficients of the template expression base, where the weight coefficients are used to quantify the weight of each template expression base in the template face model when representing the expression corresponding to the sample face image.

[0078] In some embodiments of the present invention, a system of linear equations can be constructed based on the projection result corresponding to the sample face image and the position of facial feature points in the pre-stored template expression base. The expression data corresponding to the sample face image can be obtained by fitting the system of linear equations. Specifically, a system of linear equations is constructed based on the projection result and the position of facial feature points in the pre-stored template expression base. The weight coefficients of the template expression base are obtained by fitting the system of linear equations, and the weight coefficients of the template expression base are set as the expression data corresponding to the sample face image.

[0079] In some embodiments of the present invention, the weight of each template expression base is changed according to the expression coefficient, and the neutral expression and the template expression base are weighted and summed to obtain the initial expression base corresponding to the sample face image.

[0080] In some embodiments of the present invention, a trained expression coding model can be used to encode the sample face image and the corresponding initial expression base to obtain the baseline expression coding vector of the sample face image and the transformed expression coding vector of the corresponding initial expression base. The expression coding model can be a neural network-based model or a machine learning-based model. For example, a 3D face mesh can be generated based on the initial expression base corresponding to the sample face image, the 3D face mesh can be rendered to obtain a restored face image, and the restored face image and the sample face image can be input into the trained expression coding model for expression coding to obtain the baseline expression coding vector of the sample face image and the transformed expression coding vector of the corresponding initial expression base.

[0081] In some embodiments of the present invention, an expression recognition model can be used to perform expression recognition on the sample face image and the initial expression base corresponding to the sample face image, respectively, to obtain the expression information corresponding to the sample face image and the expression information corresponding to the initial expression base corresponding to the sample face image. The expression information corresponding to the sample face image is encoded according to a preset encoding rule to obtain the base expression encoding vector of the sample face image. The expression information corresponding to the initial expression base corresponding to the sample face image is encoded according to a preset encoding rule to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image. For example, a 3D face mesh can be generated based on the initial expression basis corresponding to the sample face image. The 3D face mesh can be rendered to obtain a restored face image. The restored face image and the sample face image can be input into a trained expression recognition model for expression recognition to obtain the expression information corresponding to the restored face image and the expression information corresponding to the sample face image. The expression information corresponding to the restored face image and the sample face image can be encoded according to a preset encoding rule to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector corresponding to the restored face image. The transformed expression encoding vector corresponding to the restored face image can be set as the transformed expression encoding vector of the initial expression basis corresponding to the sample face image.

[0082] In some embodiments of the present invention, the L2 norm between the baseline expression encoding vector and the transformed expression encoding vector can be calculated, and the value of the L2 norm can be set as the semantic error between the baseline expression encoding vector and the transformed expression encoding vector. By minimizing the semantic error, the initial expression base corresponding to the sample face image can be adjusted to obtain the expression base corresponding to the sample face image.

[0083] In some embodiments of the present invention, in order to improve the matching degree between the sample face image and the corresponding expression base, and reduce the error of the expression base corresponding to the sample face image, a deformation error between the initial expression base and the pre-stored original expression base can be added on the basis of semantic error. By minimizing the semantic error and deformation error, the initial expression base is adjusted to obtain the expression base corresponding to the sample face image. Specifically, the adjustment method of the initial expression base includes steps b1 to b4:

[0084] Step b1: Determine the deformation error between the initial expression base corresponding to the sample face image and the pre-stored original expression base.

[0085] Step b2: Based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector, as well as the deformation error, determine whether the initial expression base corresponding to the sample face image matches the sample face image.

[0086] Step b3: If there is no match, adjust the initial expression base corresponding to the sample face image according to the semantic error and deformation error to obtain a new initial expression base corresponding to the sample face image.

[0087] Step b4: Based on the new deformation error between the new initial expression base and the original expression base, and the new semantic error between the transformed expression encoding vector of the new initial expression base and the reference expression encoding vector, determine whether the new initial expression base matches the sample face image. If they do not match, adjust the new initial expression base according to the new semantic error and the new deformation error until the adjusted new initial expression base matches the sample face image, and obtain the expression base corresponding to the sample face image.

[0088] Deformation error is used to quantify the deformation similarity between the initial expression base corresponding to the sample face image and the pre-stored original expression base. In some embodiments of the present invention, the L2 norm between the initial expression base and the pre-stored original expression base can be calculated, and the value of the L2 norm between the initial expression base corresponding to the sample face image and the pre-stored original expression base can be set as the deformation error between the initial expression base corresponding to the sample face image and the pre-stored original expression base. In some embodiments of the present invention, in order to make the obtained expression base cover the facial expression details in each sample face image in the training images, that is, to make a small deformation on the obtained expression base to achieve the effect of covering most expression details, the deformation error between the initial expression base corresponding to the sample face image and the pre-stored original expression base can be increased. By minimizing the deformation error, the deformation between the initial expression base corresponding to the sample face image and the pre-stored original expression base is minimized, thereby obtaining an expression base that can cover most expression details.

[0089] In some embodiments of the present invention, the pre-stored original expression base may be a pre-selected neutral expression base, or it may be the initial expression base obtained from the previous round of adjustment of the sample face image, or it may be the expression base corresponding to the previous sample face image.

[0090] In some embodiments of the present invention, the sum of semantic error and deformation error can be compared with a preset error threshold. If the sum of semantic error and deformation error is less than or equal to the preset error threshold, it is determined that the initial expression base corresponding to the sample face image matches the sample face image. If the sum of semantic error and deformation error is greater than the preset error threshold, it is determined that the initial expression base corresponding to the sample face image does not match the sample face image.

[0091] In some embodiments of the present invention, if the initial expression base corresponding to the sample face image matches the sample face image, then the initial expression base corresponding to the sample face image is set as the expression base corresponding to the sample face image.

[0092] In some embodiments of the present invention, if the initial expression base corresponding to the sample face image does not match the sample face image, the initial expression base corresponding to the sample face image is iteratively adjusted by minimizing the sum of semantic error and deformation error. After each round of adjustment, based on the obtained new initial expression base, the new deformation error between the new initial expression base corresponding to the sample face image and the original expression base, and the new semantic error between the transformed expression encoding vector of the new initial expression base corresponding to the sample face image and the reference expression encoding vector are obtained. The sum of the new deformation error and the new semantic error is compared with a preset error threshold. When the sum of the new deformation error and the new semantic error corresponding to the sample face image is less than or equal to the preset error threshold, that is, when the new initial expression base corresponding to the sample face image matches the sample face image, the adjustment is stopped, and the current new initial expression base corresponding to the sample face image is set as the expression base corresponding to the sample face image.

[0093] In some embodiments of the present invention, steps a1 to a3 are performed on each sample face image in the sample face image set to obtain the expression base corresponding to each sample face image, obtain the expression coefficient of each sample face image, and associate and store the expression coefficient of each sample face image with the corresponding expression base to obtain face expression base data.

[0094] In some embodiments of the present invention, in step a2, a restored face image can be obtained by rendering based on the initial expression base corresponding to the sample face image. Expression encoding is then performed on both the sample face image and the restored face image to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image. Specifically, the expression encoding steps include steps c1 to c2:

[0095] Step c1: Obtain the restored face image based on the expression coefficients corresponding to the sample face image and the initial expression base.

[0096] Step c2: Perform expression encoding on the restored face image to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image, and perform expression encoding on the sample face image to obtain the base expression encoding vector of the sample face image.

[0097] In some embodiments of the present invention, the method for obtaining a restored face image includes:

[0098] Based on the expression coefficients corresponding to the sample face image, a face restoration operation is performed on the initial expression base corresponding to the sample face image to obtain restored face mesh data.

[0099] The reconstructed face mesh data is rendered to obtain the reconstructed face image.

[0100] The reconstructed face mesh data can be a three-dimensional face mesh, which contains fixed points and edges. Each point and edge corresponds to semantic information, such as the xxth key point representing the left corner of the eye. In some embodiments of the present invention, the reconstructed face mesh data can be rendered using a differentiable renderer to obtain a reconstructed face image.

[0101] In some embodiments of the present invention, the facial expression encoding method includes:

[0102] Extract the facial expression information corresponding to the restored face image and the facial expression information corresponding to the sample face image.

[0103] The facial expression information corresponding to the restored face image is encoded to obtain the transformed facial expression encoding vector of the initial facial expression base corresponding to the sample face image.

[0104] The facial expression information corresponding to the sample face image is encoded to obtain the baseline facial expression encoding vector of the sample face image.

[0105] In some embodiments of the present invention, the expression information corresponding to the restored face image and the expression information corresponding to the sample face image can be extracted according to the expression information extraction method in step a2. The expression information corresponding to the restored face image is then encoded using a preset encoding rule to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image. Finally, the baseline expression encoding vector of the sample face image is obtained by applying the preset encoding rule to the expression information corresponding to the sample face image. It should be noted that the preset encoding rule is not limited in the embodiments of the present invention, and the encoding rule can be set according to the actual application scenario.

[0106] In some embodiments of the present invention, similar to step a2, the sample face image can be encoded with expressions using a trained expression encoding model to obtain the baseline expression encoding vector of the sample face image, and the restored face image can be encoded with expressions using a trained expression encoding model to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image.

[0107] In some embodiments of the present invention, semantic errors exist between the baseline expression encoding vector and the transformed expression encoding vector. These errors may stem from errors in the extraction methods of the baseline and transformed expression encoding vectors, as well as errors introduced during the rendering process. Therefore, in step a3, when adjusting the initial expression base corresponding to the sample face image based on the semantic error between the baseline and transformed expression encoding vectors, the differentiable renderer and expression encoding model can be adjusted according to the semantic error between the baseline and transformed expression encoding vectors. The initial expression base corresponding to the sample face image is adjusted based on the semantic error between the baseline and transformed expression encoding vectors and the deformation error between the initial expression base and the pre-stored original expression base, thus obtaining the expression base corresponding to the sample face image. In some embodiments of the present invention, the differentiable renderer, expression encoding model, and initial expression base can be adjusted using an automatic differentiation method with reverse gradient propagation.

[0108] In this embodiment of the invention, after obtaining the initial expression base corresponding to the sample face image through expression transfer, the initial expression base can be better adjusted by using the expression encoding vector for similarity measurement. This makes the expression base corresponding to the sample face image more closely matched with the sample face image. Furthermore, deformation error is added during the adjustment of the initial expression base to obtain an expression base that can cover more expression details, thereby making the expressions of the virtual image generated based on the expression base more vivid and realistic.

[0109] In some embodiments of the present invention, after obtaining the facial expression base data, when the target facial image to be converted is received, the target expression coefficient corresponding to the target facial image is obtained according to step S201, the facial expression base data is queried to determine the target expression base corresponding to the target expression coefficient, and according to the target expression coefficient and the target expression base, the expression of the virtual image corresponding to the target facial image is generated according to the target expression base according to step S203.

[0110] In some embodiments of this application, when the number of target expression bases is greater than 1, there are multiple sets of corresponding target expression coefficients, and each set of target expression coefficients corresponds to one target expression base. The number of target expression coefficients is equal to the number of target expression bases. The step of generating the expression of a virtual image corresponding to the target face image based on the target expression bases includes: determining the target expression coefficients corresponding to each target expression base, and combining each target expression base according to the target expression coefficients corresponding to each target expression base to obtain the expression of the virtual image corresponding to the target face image.

[0111] In this embodiment of the invention, the initial expression base is pre-adjusted by the semantic error between the base expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thus obtaining the expression base. This makes the expression base more consistent with the sample face image and improves the realism of the expression base. In this way, when the target face image to be converted and the target expression coefficient corresponding to the target face image are obtained, the target expression base corresponding to the target expression coefficient can be directly determined. Based on the target expression coefficient and the target expression base, the expression of the virtual image corresponding to the target face image is generated, which makes the sample face image more consistent with the expression base, and thus makes the expression of the virtual image generated based on the target expression base more vivid and realistic.

[0112] Based on the same inventive concept, please refer to the following: Figure 4 This invention also provides a virtual avatar expression generation device 10, which is applied to... Figure 1 The computer device shown, in this embodiment of the invention, allows the functional modules of the virtual avatar expression generation device 10 to be stored in the computer device's memory in the form of software or firmware, such as... Figure 4 As shown, the virtual avatar expression generation device 10 includes an acquisition module 11, an expression base determination module 12, an expression generation module 13, and an expression base binding module 14.

[0113] The acquisition module 11 is used to acquire the target face image to be converted, and the target expression coefficients corresponding to the target face image;

[0114] The expression base determination module 12 is used to determine the target expression base corresponding to the target expression coefficient in the facial expression base data. The facial expression base data includes multiple expression bases and the expression coefficients corresponding to each expression base. The expression base is obtained by adjusting the initial expression base based on the semantic error between the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image.

[0115] The expression generation module 13 is used to generate expressions of a virtual image corresponding to the target face image based on the target expression coefficient and the target expression base.

[0116] In some embodiments of the present invention, the expression base binding module 14 includes:

[0117] The initial expression basis generation unit is used to acquire the sample face image set, as well as the initial expression basis and expression coefficients corresponding to each sample face image in the sample face image set;

[0118] The encoding unit performs expression encoding on each sample face image and the initial expression base corresponding to the sample face image to obtain the base expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image.

[0119] The expression base binding unit is used to adjust the initial expression base according to the semantic error between the baseline expression encoding vector and the transformed expression encoding vector to obtain the expression base corresponding to the sample face image;

[0120] The facial expression data generation unit is used to associate the facial expression base and facial expression coefficients corresponding to each sample face image to obtain facial expression base data.

[0121] In some embodiments of the present invention, the encoding unit is used for:

[0122] Based on the expression coefficients and initial expression base corresponding to the sample face image, the restored face image is obtained;

[0123] The restored face image is encoded with facial expressions to obtain the transformed facial expression encoding vector of the initial facial expression base corresponding to the sample face image. The sample face image is then encoded with facial expressions to obtain the base facial expression encoding vector of the sample face image.

[0124] In some embodiments of the present invention, the encoding unit is used for:

[0125] Extract the facial expression information corresponding to the restored face image and the facial expression information corresponding to the sample face image;

[0126] Encode the facial expression information corresponding to the restored face image to obtain the transformed facial expression encoding vector of the initial facial expression base corresponding to the sample face image;

[0127] The facial expression information corresponding to the sample face image is encoded to obtain the baseline facial expression encoding vector of the sample face image.

[0128] In some embodiments of the present invention, the encoding unit is used for:

[0129] Based on the expression coefficients corresponding to the sample face image, a face restoration operation is performed on the initial expression base corresponding to the sample face image to obtain restored face mesh data;

[0130] The reconstructed face mesh data is rendered to obtain the reconstructed face image.

[0131] In some embodiments of the present invention, the expression base binding unit is used for:

[0132] Determine the deformation error between the initial expression base corresponding to the sample face image and the pre-stored original expression base;

[0133] Based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector, as well as the deformation error, determine whether the initial expression base corresponding to the sample face image matches the sample face image.

[0134] If there is no match, the initial expression base corresponding to the sample face image is adjusted according to the semantic error and deformation error to obtain a new initial expression base corresponding to the sample face image.

[0135] Based on the new deformation error between the new initial expression base and the original expression base, and the new semantic error between the transformed expression encoding vector of the new initial expression base and the baseline expression encoding vector, it is determined whether the new initial expression base matches the sample face image. If they do not match, the new initial expression base is adjusted according to the deformation error and the new semantic error until the adjusted new initial expression base matches the sample face image, thus obtaining the expression base corresponding to the sample face image.

[0136] In some embodiments of the present invention, the initial expression base generation unit is used for:

[0137] Face detection is performed on each sample face image in the sample face image set to obtain the location of facial feature points in each sample face image;

[0138] For each sample face image, the positions of facial feature points in the sample face image are projected to obtain the projection result of the sample face image;

[0139] The projection result of the sample face image is fitted with the position of the facial feature points in the pre-stored template expression base to obtain the expression data corresponding to the sample face image;

[0140] The template expression base is adjusted based on the expression data corresponding to the sample face image to obtain the initial expression base corresponding to the sample face image.

[0141] In some embodiments of the present invention, the expression generation module 13 is used for:

[0142] Determine the target expression coefficient corresponding to each target expression base, and combine each target expression base according to the target expression coefficient corresponding to each target expression base to obtain the expression of the virtual image corresponding to the target face image.

[0143] The virtual avatar expression generation device provided in this embodiment of the invention pre-adjusts the initial expression base based on the semantic error between the reference expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thereby obtaining the expression base. This makes the expression base more consistent with the sample face image and improves the realism of the expression base. Thus, after obtaining the target face image and the target expression coefficient corresponding to the target face image, the target expression base corresponding to the target expression coefficient in the face expression base data can be directly determined based on the target expression coefficient. Then, based on the target expression coefficient and the target expression base, the expression of the virtual avatar corresponding to the target face image is generated, making the expression of the virtual avatar more vivid and realistic.

[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the image processing device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0145] Based on the above, embodiments of the present invention provide a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the virtual image expression generation method described in any of the foregoing embodiments.

[0146] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the computer equipment described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0147] Based on the above, embodiments of the present invention provide a storage medium storing a computer program that, when executed by a processor, implements a method for generating facial expressions of a virtual avatar according to any of the aforementioned embodiments.

[0148] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the storage medium described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0149] In summary, the virtual avatar expression generation method, apparatus, computer device, and storage medium provided in this embodiment of the invention pre-adjust the initial expression base based on the semantic error between the reference expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image, thereby obtaining the expression base. This makes the expression base more closely match the sample face image and improves the realism of the expression base. Thus, after obtaining the target face image and the target expression coefficient corresponding to the target face image, the target expression base corresponding to the target expression coefficient in the face expression base data can be directly determined based on the target expression coefficient. Then, based on the target expression coefficient and the target expression base, the expression of the virtual avatar corresponding to the target face image is generated, making the expression of the virtual avatar more vivid and realistic.

[0150] The above descriptions are merely various embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating facial expressions for a virtual avatar, characterized in that, The method includes: Obtain the target face image and the target expression coefficient corresponding to the target face image; Obtain a sample face image set, as well as the initial expression base and expression coefficients corresponding to each sample face image in the sample face image set; For each sample face image, expression encoding is performed on the sample face image and the initial expression base corresponding to the sample face image to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image. Based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector, the initial expression base is adjusted to obtain the expression base corresponding to the sample face image; The facial expression base and expression coefficients corresponding to each sample face image are correlated to obtain facial expression base data; Determine the target expression base in the facial expression base data that corresponds to the target expression coefficient; the facial expression base data includes multiple expression bases and the expression coefficients corresponding to each expression base; Based on the target expression base, generate the expression of a virtual avatar corresponding to the target face image.

2. The method for generating facial expressions for virtual avatars as described in claim 1, characterized in that, The step of encoding expressions for the sample face image and the corresponding initial expression base to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the corresponding initial expression base includes: Based on the expression coefficients and initial expression base corresponding to the sample face image, the restored face image is obtained; The restored face image is subjected to expression encoding to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image. The sample face image is then subjected to expression encoding to obtain the base expression encoding vector of the sample face image.

3. The method for generating facial expressions for virtual avatars as described in claim 2, characterized in that, The process of encoding expressions on the restored face image to obtain the transformed expression encoding vector of the initial expression base corresponding to the sample face image, and encoding expressions on the sample face image to obtain the base expression encoding vector of the sample face image, includes: Extract the facial expression information corresponding to the restored face image and the facial expression information corresponding to the sample face image; The facial expression information corresponding to the restored face image is encoded to obtain the transformed facial expression encoding vector of the initial facial expression base corresponding to the sample face image; The facial expression information corresponding to the sample face image is encoded to obtain the baseline facial expression encoding vector of the sample face image.

4. The method for generating facial expressions for virtual avatars as described in claim 2, characterized in that, The process of obtaining the restored face image based on the expression coefficients and initial expression base corresponding to the sample face image includes: Based on the expression coefficients corresponding to the sample face image, a face restoration operation is performed on the initial expression base corresponding to the sample face image to obtain restored face mesh data; The restored face mesh data is rendered to obtain a restored face image.

5. The method for generating facial expressions for virtual avatars as described in claim 1, characterized in that, The step of adjusting the initial expression base based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector to obtain the expression base corresponding to the sample face image includes: Determine the deformation error between the initial expression base corresponding to the sample face image and the pre-stored original expression base; Based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector, as well as the deformation error, it is determined whether the initial expression base corresponding to the sample face image matches the sample face image. If there is no match, the initial expression base corresponding to the sample face image is adjusted according to the semantic error and the deformation error to obtain a new initial expression base corresponding to the sample face image. Based on the new deformation error between the new initial expression base and the original expression base, and the new semantic error between the transformed expression encoding vector of the new initial expression base and the reference expression encoding vector, it is determined whether the new initial expression base matches the sample face image. If they do not match, the new initial expression base is adjusted according to the deformation error and the new semantic error until the adjusted new initial expression base matches the sample face image, thus obtaining the expression base corresponding to the sample face image.

6. The method for generating facial expressions for virtual avatars as described in claim 1, characterized in that, The step of obtaining the initial expression base corresponding to each sample face image in the sample face image set includes: Face detection is performed on each sample face image in the sample face image set to obtain the position of facial feature points in each sample face image; For each sample face image, the positions of facial feature points in the sample face image are projected to obtain the projection result of the sample face image; The projection result of the sample face image is fitted with the position of the facial feature points in the pre-stored template expression base to obtain the expression data corresponding to the sample face image; The template expression base is adjusted based on the expression data corresponding to the sample face image to obtain the initial expression base corresponding to the sample face image.

7. The method for generating facial expressions for virtual avatars as described in any one of claims 1 to 6, characterized in that, The step of generating the expression of a virtual avatar corresponding to the target face image based on the target expression base includes: Determine the target expression coefficient corresponding to each target expression base, and combine each target expression base according to the target expression coefficient corresponding to each target expression base to obtain the expression of the virtual image corresponding to the target face image.

8. A device for generating facial expressions for virtual avatars, characterized in that, The device includes: The acquisition module is used to acquire the target face image to be converted, and the target expression coefficients corresponding to the target face image; An expression base binding module is used to acquire a set of sample face images, as well as the initial expression base and expression coefficients corresponding to each sample face image in the set; for each sample face image, expression encoding is performed on the sample face image and the initial expression base corresponding to the sample face image to obtain the baseline expression encoding vector of the sample face image and the transformed expression encoding vector of the initial expression base corresponding to the sample face image; based on the semantic error between the baseline expression encoding vector and the transformed expression encoding vector, the initial expression base is adjusted to obtain the expression base corresponding to the sample face image; the expression base and expression coefficients corresponding to each sample face image are associated to obtain face expression base data; The expression base determination module is used to determine the target expression base in the facial expression base data that corresponds to the target expression coefficient; the facial expression base data includes multiple expression bases and the expression coefficient corresponding to each expression base; The expression generation module is used to generate expressions of a virtual image corresponding to the target face image based on the target expression base.

9. A computer device, characterized in that, The computer device includes a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to perform the operations in the virtual avatar expression generation method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a plurality of instructions adapted for loading by a processor to execute the steps in the virtual avatar expression generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Expression redirection method and device, equipment and medium

    CN113066156A