An expression generation method and apparatus, an electronic device, and a storage medium
By constructing an expression generation model and a refinement model, and using the expression generation model to obtain transfer data and refine the expressions, the problem of poor 3D expression generation in existing technologies is solved, and more accurate and vivid expression generation is achieved.
Patent Information
- Application Number
- CN202510166820.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In existing technologies, voice-driven methods for generating 3D facial expressions cannot effectively address the poor quality of voice-driven 3D facial expression generation, resulting in insufficient facial expression information, simplistic and unnatural expressions, and limited versatility.
By extracting patents, collecting the same object, collecting several training facial expression samples, constructing an facial expression generation model, obtaining transfer data, updating the model, generating more accurate facial expression images, and further refining them through the facial expression refinement model to generate more vivid facial expression details.
It enables more efficient generation of 3D expressions, improves the naturalness and versatility of expression generation, and makes the generated expressions more accurate and vivid.
Smart Images

Figure CN120014094B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image data processing, and more particularly, to an expression generation method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the continuous development of technology, computer vision and artificial intelligence technology have made remarkable achievements in various fields. Among them, three-dimensional expression driving technology has wide application prospects in virtual reality, animation production, game development and other fields. As an important technology in three-dimensional expression driving, expression generation is increasingly attracting the attention of researchers. In the prior art, three-dimensional expressions are usually driven by voice, but the expression information contained in the voice is less, so the effect of driving expressions by voice is poor. SUMMARY
[0003] The present application aims to overcome at least one of the above-mentioned defects of the prior art, and provides an expression generation method, device, electronic equipment and storage medium, which can more effectively generate three-dimensional expressions.
[0004] According to an aspect of the present application, an expression generation method is provided, the generation method comprising:
[0005] Collecting a plurality of training expression samples, each of which contains a first expression image sample and a second expression image sample; the objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, and the expressions of the objects are different;
[0006] Constructing an expression generation model;
[0007] Obtaining the transfer data between the corresponding first expression image sample and second expression image sample in each training expression sample through the expression generation model;
[0008] Obtaining the first driving expression image of the first expression image sample and the second driving expression image of the second expression image sample according to the corresponding transfer data through the expression generation model;
[0009] Updating the expression generation model according to the first driving expression image and the second driving expression image to obtain a trained expression generation model;
[0010] Collecting a target basic image and a target expression driving image, inputting the target basic image and the target expression driving image into the trained expression generation model, and obtaining a target expression image.
[0011] Optionally, the obtaining, by the expression generation model, migration data between the corresponding first expression image sample and second expression image sample in each training expression sample comprises:
[0012] processing, by the expression generation model, the corresponding first expression image sample in each training expression sample to obtain first expressionless frontal face key points and first expression face key points;
[0013] processing, by the expression generation model, the corresponding second expression image sample in each training expression sample to obtain second expressionless frontal face key points and second expression face key points;
[0014] obtaining first expression migration data according to the second expression face key points and the first expressionless frontal face key points, and obtaining second expression migration data according to the first expression face key points and the second expressionless frontal face key points.
[0015] Optionally, the obtaining, by the expression generation model, migration data between the corresponding first expression image sample and second expression image sample in each training expression sample comprises:
[0016] inputting the corresponding first expression image sample in each training expression sample into the expression generation model to obtain corresponding first object face features;
[0017] inputting the corresponding second expression image sample in each training expression sample into the expression generation model to obtain corresponding second object face features; the expression generation model obtains the first driven expression image according to the first expression migration data and the first object face features;
[0018] the expression generation model obtains the second driven expression image according to the second expression migration data and the second object face features.
[0019] Optionally, after the collecting of the plurality of training expression samples, the method further comprises adding key point labels to the first expression image sample and the second expression image sample in each training expression sample.
[0020] The adding of the key point labels to the first expression image sample and the second expression image sample in each training expression sample comprises:
[0021] adding eye key point labels, eyebrow key point labels and mouth key point labels to the first expression image sample, and adding eye key point labels, eyebrow key point labels and mouth key point labels to the second expression image sample;
[0022] The migration data between the corresponding first expression image sample and second expression image sample in each training expression sample is obtained, and specifically is:
[0023] The migration data between the corresponding first expression image sample and second expression image sample in each training expression sample after adding the key point label is obtained.
[0024] Optionally, the expression generation model is updated according to the first driving expression image and the second driving expression image, and specifically includes:
[0025] An expressionless loss is calculated according to the first expressionless frontal key point and the second expressionless frontal key point.
[0026] A first eye part loss, a first eyebrow part loss and a first mouth part loss are respectively calculated according to the first expression facial key point and the corresponding eye part key point label, eyebrow part key point label and mouth part key point label.
[0027] A second eye part loss, a second eyebrow part loss and a second mouth part loss are respectively calculated according to the second expression facial key point and the corresponding eye part key point label, eyebrow part key point label and mouth part key point label.
[0028] A first reconstruction loss is calculated according to the first driving expression image and the second expression image sample.
[0029] A second reconstruction loss is calculated according to the second driving expression image and the first expression image sample.
[0030] An image reconstruction loss is obtained according to the first reconstruction loss and the second reconstruction loss.
[0031] The expression generation model is updated according to the expressionless loss, and / or the first eye part loss, and / or the first eyebrow part loss, and / or the first mouth part loss, and / or the second eye part loss, and / or the second eyebrow part loss, and / or the second mouth part loss, and / or the image reconstruction loss.
[0032] Optionally, before the target basic image and the target expression driving image are collected, the method further includes:
[0033] An expression refining model is constructed.
[0034] A first eye part coefficient, a first eyebrow part coefficient and a first mouth part coefficient are obtained according to the first expression facial key point, and a second eye part coefficient, a second eyebrow part coefficient and a second mouth part coefficient are obtained according to the second expression facial key point.
[0035] The first driving expression image, the first eye part coefficient, the first eyebrow part coefficient and the first mouth part coefficient are processed by using the expression refining model to obtain a first expression refining image.
[0036] processing the second driving expression image, the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient by using the expression refining model to obtain a second expression refined image;
[0037] updating the expression refining model according to the first expression refined image and the second expression refined image to obtain a trained expression refining model;
[0038] After the target expression image is obtained, the method further includes:
[0039] processing the target expression image by using the trained expression refining model to obtain a target refined expression image.
[0040] Optionally, the updating the expression refining model according to the first expression refined image and the second expression refined image specifically includes:
[0041] calculating a first refined loss according to the first expression refined image and the second expression image sample;
[0042] calculating a second refined loss according to the second expression refined image and the first expression image sample;
[0043] obtaining an image refined reconstruction loss according to the first refined loss and the second refined loss;
[0044] updating the expression refining model according to the image refined reconstruction loss.
[0045] According to a second aspect of the present application, an expression generation device is provided, and the device includes:
[0046] a sample collection module configured to collect a plurality of training expression samples, each of the training expression samples including a first expression image sample and a second expression image sample; the objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, and the expressions of the objects are different;
[0047] a model construction module configured to construct an expression generation model;
[0048] a data processing module configured to obtain transfer data between the corresponding first expression image sample and the second expression image sample in each of the training expression samples by using the expression generation model;
[0049] an expression driving module configured to obtain a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding transfer data by using the expression generation model;
[0050] The model updating module is configured to update the expression generation model according to the first driving expression image and the second driving expression image, and obtain a trained expression generation model.
[0051] The target expression generation module is configured to collect a target base image and a target expression driving image, input the target base image and the target expression driving image into the trained expression generation model, and obtain a target expression image.
[0052] According to a third aspect of the present application, an electronic device is provided, which includes a memory and a processor, the memory stores computer readable instructions, and the processor executes the computer readable instructions to implement the expression generation method of the first aspect.
[0053] According to a fourth aspect of the present application, a computer storage medium is provided, which stores a computer readable program, and the computer readable program is executed to implement the expression generation method of the first aspect.
[0054] According to any one of the above aspects, the expression generation method, device, electronic device and storage medium provided by the present application collect first expression image samples and second expression image samples containing the same object and different expressions of the object as training expression samples, and obtain transfer data between the first expression image samples and the second expression image samples through an expression generation model. On the one hand, the first expression image samples and the second expression image samples can be used as training samples at the same time, increasing the number of training samples of the expression generation model. On the other hand, the first expression image samples and the second expression image samples can be used as references to each other, enabling the expression generation model to learn sufficiently, so that the finally trained expression generation model can effectively and sufficiently utilize the expression information in the image containing the target expression to generate a target expression image corresponding to the expression on the base image.
[0055] Further, the expression generation method, device, electronic device and storage medium provided by the present application construct an expression refining model, and further refine the expression of the face through corresponding eye coefficients, eyebrow coefficients and mouth coefficients through the constructed expression refining model. The image refined through the expression refining model has more accurate and lively expression details. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Figure 1The application scenario diagram of the generation method provided in the embodiment.
[0057] Figure 2 The step flow chart of the generation method provided in the embodiment Figure 1 .
[0058] Figure 3 The step flow chart of the migration data acquisition provided in the embodiment.
[0059] Figure 4 The step flow chart of the first driving expression image and the second driving expression image acquisition provided in the embodiment.
[0060] Figure 5 The step flow chart of the expression generation model update provided in the embodiment.
[0061] Figure 6 The step flow chart of the generation method provided in the embodiment Figure 2 .
[0062] Figure 7 The step flow chart of the expression refining model update provided in the embodiment.
[0063] Figure 8 The device structure diagram of the generation device provided in the embodiment.
[0064] Figure 9 The device structure diagram of the electronic device provided in the embodiment.
[0065] The figure caption: server 100, terminal 200, sample collection module 11, generation model construction module 12, data processing module 13, expression driving module 14, generation model update module 15, target expression generation module 16, refining model construction module 17, sample expression refining module 18, refining model update module 19, target expression refining module 20, memory 31, processor 32, bus 33, communication interface 34. DETAILED DESCRIPTION
[0066] The drawings of the present application are only used for illustrative description, and cannot be understood as the limitation of the present application. In order to better illustrate the following embodiments, some components of the drawings will be omitted, enlarged or reduced, and the size of the actual product is not represented; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings can be omitted.
[0067] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should be within the scope of protection of the present application.
[0068] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0069] Embodiment 1
[0070] With the continuous development of technology, computer vision and artificial intelligence technology have achieved remarkable results in various fields. Among them, three-dimensional expression driving technology has wide application prospects in virtual reality, animation production, game development and other fields. As an important technology in three-dimensional expression driving, expression generation is increasingly attracting the attention of researchers.
[0071] Currently, expression generation technology is mainly based on speech generation, which analyzes the features of images and speech in videos, obtains corresponding expression information features according to the facial expressions of objects in corresponding images and speech, and reconstructs three-dimensional facial expressions according to the extracted expression information features. However, the method based on speech has the following limitations:
[0072] Limited signal: the expression information contained in the speech signal is limited, resulting in a single expression that is difficult to meet the needs of complex scenarios.
[0073] Poor effect: due to the lack of expression information in the speech signal, the generated expression has low naturalness, which is easy to cause discomfort to the user.
[0074] Lack of universality: different languages and different speakers have great differences in speech signals, making it difficult for the speech-based method to achieve universality.
[0075] The embodiment provides a technical solution capable of solving the above problems. The specific implementation of the application is described in detail below with reference to the drawings.
[0076] Exemplarily, an application scenario schematic diagram of an expression generation method provided by the embodiment of the application is shown in FIG. 1. Figure 1 As shown in the figure, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a data transmission function of video stream and audio stream. The terminal 200 has a streaming media playing function and can also have an image processing function.
[0077] It can be understood that the server 100 can be an independent electronic device or a cluster composed of multiple electronic devices. The terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.
[0078] In an implementable manner, the server 100 and the terminal 200 can respectively execute an expression generation method provided by the embodiment of the application, or alternatively, the expression generation method provided by the embodiment of the application is partially executed in the server 100 and partially executed in the terminal 200.
[0079] As shown in the figure, the embodiment provides an expression generation method, which can specifically include: Figure 2
[0080] S1: Collecting a plurality of training expression samples, each of which contains a first expression image sample and a second expression image sample; wherein the objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, and the expressions of the objects are different;
[0081] In the embodiment, the collection of the training expression sample can be performed by collecting a video containing facial expressions of the same object, taking two video images of the same object but different expressions as the first expression image sample and the second expression image sample, respectively. The training expression sample can also be directly collected by an image collection device such as a camera, and two facial images of the same object but different expressions are directly collected as the first expression image sample and the second expression image sample, respectively.
[0082] It can be understood that for the first expression image sample and the second expression image sample in the same training expression sample, the objects need to be the same, and for different training expression samples, the corresponding objects can be different. In order to improve the generality of the model expression generation, different objects of the training expression sample need to be collected. The more the number of collected objects, the higher the generality of the model obtained by subsequent training.
[0083] It can be understood that facial expressions are mostly transmitted through the three parts of eyes, eyebrows and mouth. In order to better learn the features of the eyes, eyebrows and mouth of the face, in the embodiment, after a plurality of training expression samples are collected, key point labels are added to the first expression image sample and the second expression image sample in each training expression sample.
[0084] Specifically, the key point labels are added to the first expression image sample and the second expression image sample in each training expression sample, specifically:
[0085] The eye key point label, the eyebrow key point label and the mouth key point label are added to the first expression image sample, and the eye key point label, the eyebrow key point label and the mouth key point label are added to the second expression image sample.
[0086] Specifically, for all the training expression samples, the eye key point label is added to the eye key point in each first expression image, the eyebrow key point label is added to the eyebrow key point, and the mouth key point label is added to the mouth key point. Correspondingly, the eye key point label is added to the eye key point in each second expression image sample, the eyebrow key point label is added to the eyebrow key point, and the mouth key point label is added to the mouth key point. The other key points of the first expression image and the second expression image are not labeled. By adding labels to the eye, eyebrow and mouth key points and not adding labels to other key points, the model can be semi-supervised based on the training expression samples. On the one hand, the model can effectively learn based on the eye, eyebrow and mouth key points, improve the accuracy of the eye, eyebrow and mouth parts, and further make the generated expression more accurate and lively. On the other hand, the model can improve the generalization ability for different object faces and effectively generate expressions of different objects.
[0087] S2: constructing an expression generation model;
[0088] S3: obtaining transfer data between the corresponding first expression image sample and second expression image sample in each training expression sample by the expression generation model;
[0089] As described above, in order to improve the accuracy of the eye, eyebrow and mouth parts and enable the expression generation model to perform semi-supervised learning training, preferably, the transfer data between the corresponding first expression image sample and second expression image sample in each training expression sample is obtained, specifically:
[0090] The transfer data between the corresponding first expression image sample and second expression image sample in each training expression sample after the key point labels are added is obtained.
[0091] In the embodiment, based on each training expression sample after adding the key point label, as shown in Figure 3 The step S3 can specifically include the following steps:
[0092] S31: processing the corresponding first expression image sample in each training expression sample by using the expression generation model to obtain first expressionless frontal face key points and first expression face key points;
[0093] It can be understood that the first expressionless frontal face key points represent the frontal face key point condition of the corresponding object without expression, and the first expression face key points represent the face key point condition of the corresponding object with expression.
[0094] S32: processing the corresponding second expression image sample in each training expression sample by using the expression generation model to obtain second expressionless frontal face key points and second expression face key points;
[0095] Corresponding to the obtained result based on the first expression image sample, the second expressionless frontal face key points represent the frontal face key point condition of the corresponding object without expression, and the second expression face key points represent the face key point condition of the corresponding object with expression.
[0096] And the objects of the first expression image sample and the second expression image sample are the same, so for the same training expression sample, the first expressionless frontal face key points and the second expressionless frontal face key points obtained should be the same.
[0097] S33: obtaining first expression migration data according to the second expression face key points and the first expressionless frontal face key points, and obtaining second expression migration data according to the first expression face key points and the second expressionless frontal face key points.
[0098] In the embodiment, the first expression migration data and the second expression migration data represent the change condition of expression. Specifically, the first expression migration data is the migrated key point, which can represent the second expression face key points, and the change condition compared with the first expressionless frontal face key points. Correspondingly, the second expression migration data is the migrated key point, which can represent the first expression face key points, and the change condition compared with the second expressionless key points.
[0099] Specifically, the calculation of the first expression migration data can be expressed as:
[0100]
[0101] The calculation of the second expression migration data can be expressed as:
[0102]
[0103] wherein, denotes the first expression transfer data, denotes the second expression transfer data; s is a scalar, denoting the size of the face, denotes the face size corresponding to the first expression image sample, denotes the face size corresponding to the second expression image sample; N is a vector, denoting the neutral face key points, denotes the first neutral face key points, denotes the second neutral face key points; denotes the first expression transfer data, corresponding to the expression of the second expression image sample, denotes the second expression transfer data, corresponding to the expression of the first expression image sample; denotes the displacement of the face relative to the neutral face in the first expression image sample, denotes the displacement of the face relative to the neutral face in the second expression image sample.
[0104] S4: obtaining, by the expression generation model, a first driven expression image of the first expression image sample according to the corresponding expression transfer data, and obtaining a second driven expression image of the second expression image sample;
[0105] Specifically, in the embodiment, as shown in Figure 4 , the step S4 can specifically include:
[0106] S41: inputting the corresponding first expression image sample in each training expression sample into the expression generation model to obtain a corresponding first object facial feature;
[0107] It can be understood that the first object facial feature represents the specific facial features of the corresponding object, such as skin color, hair color, etc.
[0108] S42: inputting the corresponding second expression image sample in each training expression sample into the expression generation model to obtain a corresponding second object facial feature;
[0109] Corresponding to the first object facial feature, the second object facial feature represents the specific facial features of the corresponding object.
[0110] S43: the expression generation model obtains the first driven expression image according to the first expression transfer data and the first object facial feature;
[0111] It can be understood that the first driven expression image represents a prediction of an expression contained in the second expression image sample based on the first expressionless frontal key point and using the first object facial feature.
[0112] S44: The expression generation model obtains the second driven expression image according to the second expression migration data and the second object facial feature.
[0113] It can be understood that the second driven expression image represents a prediction of an expression contained in the first expression image sample based on the second expressionless frontal key point and using the second object facial feature.
[0114] The obtaining of the first driven expression image can specifically include:
[0115] The first expression migration data is used to perform expression action migration calculation through the first expressionless frontal key point, and the first migration expression obtained after the calculation is input into the expression generation model together with the first object facial feature to generate the first driven expression image.
[0116] Correspondingly, the obtaining of the second driven expression image can specifically include:
[0117] The second expression migration data is used to perform expression action migration calculation through the second expressionless frontal key point, and the second migration expression obtained after the calculation is input into the expression generation model together with the second object facial feature to generate the second driven expression image.
[0118] Based on the above steps S3 and S4, it can be understood that in the embodiment, the expression generation model at least includes an expression prediction network and an expression generation network, and further, the expression prediction network can include an expressionless key point prediction unit, an expression key point prediction unit, and a facial feature encoding module.
[0119] Specifically, the expressionless key point prediction unit is configured to obtain the first expressionless frontal key point according to the first expression image sample and obtain the second expressionless frontal key point according to the second expression image sample.
[0120] The expression key point prediction unit is configured to obtain the first expression facial key point according to the first expression image sample and obtain the second expression facial key point according to the second expression image sample.
[0121] The facial feature encoding unit is configured to obtain the first object facial feature according to the first expression image sample and obtain the second object facial feature according to the second expression image sample.
[0122] The expression generation network is configured to generate the first driven expression image according to the first expression transfer data based on the first expressionless frontal face key points and the first object facial features, and generate the second driven expression image according to the second expression transfer data based on the second expressionless frontal face key points and the second object facial features.
[0123] In a specific implementation of the embodiment, the expressionless key point prediction unit, the expression key point prediction unit and the facial feature encoding module can all use an existing EffcientNet as a network backbone.
[0124] The expression generation network can be constructed using an existing neural network encoder Unet.
[0125] S5: updating the expression generation model according to the first driven expression image and the second driven expression image to obtain a trained expression generation model;
[0126] In this step, the expression generation model is updated according to the first driven expression image and the second driven expression image, as shown in the expression generation model, which can specifically include: Figure 5
[0127] S51: calculating an expressionless loss according to the first expressionless frontal face key points and the second expressionless frontal face key points;
[0128] As described above, since the first expression image sample and the second expression image sample are of the same object, theoretically, the first expressionless frontal face key points and the second expressionless frontal face key points should be the same, so by calculating the expressionless loss, the expression generation model can effectively obtain the expressionless frontal face key points of the input expression image.
[0129] S52: calculating a first eye loss, a first eyebrow loss and a first mouth loss according to the first expression facial key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively;
[0130] S53: calculating a second eye loss, a second eyebrow loss and a second mouth loss according to the second expression facial key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively;
[0131] By calculating the losses of the eye, eyebrow and mouth respectively, the expression generation model can focus on the learning of the eye, eyebrow and mouth features, and improve the accuracy of the eye, eyebrow and mouth of the generated expression image.
[0132] S54: calculating a first reconstruction loss according to the first driven expression image and the second expression image sample;
[0133] S55: calculating a second reconstruction loss according to the second driven expression image and the first expression image sample;
[0134] S56: obtaining an image reconstruction loss according to the first reconstruction loss and the second reconstruction loss;
[0135] It can be understood that the first reconstruction loss and the second reconstruction loss represent the overall loss of the expression generation model based on the input expression image to generate the driven expression image.
[0136] In this embodiment, the image reconstruction loss can be represented as:
[0137] +
[0138]
[0139] In the formula, represents the first driven expression image, represents the second driven expression image; represents a pre-trained image feature extraction model, represents extracting image features of the corresponding image; represents the image reconstruction loss; represents the first expression image sample, represents the second expression image sample.
[0140] S57: updating the expression generation model according to the expressionless loss, and / or the first eye part loss, and / or the first eyebrow part loss, and / or the first mouth part loss, and / or the second eye part loss, and / or the second eyebrow part loss, and / or the second mouth part loss, and / or the image reconstruction loss.
[0141] Preferably, the expression generation model is updated according to the expressionless loss, the first eye part loss, the first eyebrow part loss, the first mouth part loss, the second eye part loss, the second eyebrow part loss, the second mouth part loss, and the image reconstruction loss, so that the obtained expression generation model is more stable.
[0142] In this embodiment, after the expression generation model is updated, as shown in Figure 6 it further includes:
[0143] A1: constructing an expression refining model;
[0144] In this embodiment, the expression refining model can be constructed by using an existing neural network encoder Unet;
[0145] A2: obtaining a first eye coefficient, a first eyebrow coefficient and a first mouth coefficient according to the first facial expression key point, and obtaining a second eye coefficient, a second eyebrow coefficient and a second mouth coefficient according to the second facial expression key point;
[0146] Specifically, the first eye coefficient and the second eye coefficient are eye closing proportion coefficients, the first eyebrow coefficient and the second eyebrow coefficient are eyebrow movement coefficients, and the first mouth coefficient and the second mouth coefficient are mouth opening and closing coefficients, and the value range of the eye closing proportion coefficient, the eyebrow movement coefficient and the mouth opening and closing coefficient is 0-1; preferably, the eye closing proportion coefficient, the eyebrow movement coefficient and the eyebrow movement coefficient can be obtained by extracting the expression coefficients of the face in the corresponding first expression image sample and second expression image sample.
[0147] A3: processing the first driving expression image, the first eye coefficient, the first eyebrow coefficient and the first mouth coefficient by using the expression refining model to obtain a first expression refined image;
[0148] A4: processing the second driving expression image, the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient by using the expression refining model to obtain a second expression refined image;
[0149] The expression refining model refines the first driving expression image according to the first eye coefficient, the first eyebrow coefficient and the first mouth coefficient to obtain the refined first driving expression image as the first expression refined image, and refines the second driving expression image according to the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient to obtain the refined second driving expression image as the second expression refined image. By using the expression refining model, the specific expression coefficients of the regions such as eyes, eyebrows and mouth which can best reflect the facial expression are obtained to refine the facial expression, so that the final obtained facial expression has more accurate and vivid expression.
[0150] A5: updating the expression refining model according to the first expression refined image and the second expression refined image to obtain a trained expression refining model;
[0151] In this step, the updating of the expression refining model according to the first expression refined image and the second expression refined image can specifically include: Figure 7
[0152] A51: calculating a first refining loss according to the first expression refined image and the second expression image sample;
[0153] A52: calculating a second refining loss according to the second expression refined image and the first expression image sample;
[0154] A53: obtaining an image refinement reconstruction loss according to the first refinement loss and the second refinement loss;
[0155] In the embodiment, the image refinement reconstruction loss can be represented as:
[0156] +
[0157]
[0158] In the formula, denotes the first expression refinement image, denotes the second expression refinement image; denotes a pre-trained image feature extraction model, denotes extracting image features of the corresponding image; denotes the image refinement reconstruction loss; denotes the first expression image sample, denotes the second expression image sample.
[0159] A54: updating the expression refinement model according to the image refinement reconstruction loss.
[0160] In one specific implementation of the embodiment, the expression generation model and the expression refinement model can be integrated into an expression generation refinement large model. It can be understood that the training and updating of the expression generation refinement large model includes two stages, the first stage is an expression generation training stage, and the second stage is an expression refinement training stage. In the training process of the first stage, the expression refinement model part in the expression generation refinement large model is fixed, and the expression generation model part in the expression generation refinement large model is trained and updated through the collected training expression samples. The specific training steps can refer to the specific description of steps S1-S5 above. In the training process of the second stage, the expression generation model part is fixed, and the expression refinement model part is trained and updated through the first driving expression image and the second driving expression generated in the first stage. The specific training steps can refer to the above steps A1-A5.
[0161] S6: collecting a target basic image and a target expression driving image, inputting the target basic image and the target expression driving image into the trained expression generation model, and obtaining a target expression image.
[0162] It can be understood that in the embodiment, after the target expression image is obtained through the expression generation model, it further includes:
[0163] A6: processing the target expression image through the trained expression refinement model to obtain a target refined expression image.
[0164] In one specific implementation of the embodiment, the expression generation model obtains corresponding target expressionless frontal face key points and target object facial features according to the target base image, and obtains corresponding target expression face key points according to the target expression driving image; then, target migration data is obtained according to the target expressionless frontal face key points and the target expression face key points; the target migration data obtained is input into the expression generation model together with the target object facial features, so as to obtain a target driving expression image as the target expression image. Then, target eye coefficients, target eyebrow coefficients and target mouth coefficients of the target base image are obtained, and the target eye coefficients, target eyebrow coefficients and target mouth coefficients are input into the expression refining model together with the target expression image, so as to obtain a refined target expression refining image.
[0165] In the embodiment, by collecting first expression image samples and second expression image samples containing the same object and different expressions of the object as training expression samples, migration data between the first expression image samples and the second expression image samples is obtained through the expression generation model. On the one hand, the first expression image samples and the second expression image samples can be simultaneously used as training samples, thereby increasing the number of training samples of the expression generation model. On the other hand, the first expression image samples and the second expression image samples can be used as references for each other, so that the expression generation model can be fully learned. The finally trained expression generation model can effectively and fully utilize expression information in an image containing a target expression to generate a target expression image corresponding to the expression on a base image.
[0166] Meanwhile, the embodiment constructs an expression refining model to further refine the expression of the eye, eyebrow and mouth parts according to the expression coefficients, and the image refined through the expression refining model has more accurate and lively expression details.
[0167] Based on the same inventive concept, the embodiment further provides an expression generation device, as shown in Figure 8 The expression generation device can specifically include:
[0168] The sample collection module 11 is configured to collect a plurality of training expression samples, each of which contains a first expression image sample and a second expression image sample. The objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, and the expressions of the objects are different.
[0169] In the embodiment, the sample collection module 11 can be configured to perform step S1 as shown in Figure 2 The specific description of the sample collection module 11 can refer to the description of step S1.
[0170] The generation model construction module 12 is configured to construct an expression generation model.
[0171] In this embodiment, the generation model construction module 12 can be configured to perform Figure 2 The specific description of the generation model construction module 12 can refer to the description of step S2.
[0172] The data processing module 13 is configured to obtain, by the expression generation model, transfer data between the first expression image sample and the second expression image sample corresponding to each training expression sample.
[0173] In this embodiment, the data processing module 13 can be configured to perform Figure 2 Step S3, and Figure 3 Steps S31-S33, and the specific description of the data processing module 13 can refer to the description of step S3 and steps S31-S33.
[0174] The expression driving module 14 is configured to obtain, by the expression generation model, a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding transfer data.
[0175] In this embodiment, the expression driving module 14 can be configured to perform Figure 2 Step S4, and Figure 4 Steps S41-S42, and the specific description of the expression driving module 14 can refer to the description of step S3 and steps S41-S42.
[0176] The generation model updating module 15 is configured to update the expression generation model according to the first driving expression image and the second driving expression image, to obtain a trained expression generation model.
[0177] In this embodiment, the generation model updating module 15 can be configured to perform Figure 2 Step S5, and Figure 5 Steps S51-S57, and the specific description of the generation model updating module 15 can refer to the description of step S5 and steps S51-S57.
[0178] The target expression generation module 16 is configured to collect a target basic image and a target expression driving image, input the target basic image and the target expression driving image into the trained expression generation model, and obtain a target expression image.
[0179] In this embodiment, the target expression generation module 16 can be configured to performFigure 2 As shown in step S6, the specific description of the target expression generation module 16 can refer to the description of step S6.
[0180] The refining model construction module 17 is configured to construct an expression refining model.
[0181] In this embodiment, the refining model construction module 17 can be configured to perform Figure 6 As shown in step A1, the specific description of the refining model construction module 17 can refer to the description of step A1.
[0182] The sample expression refining module 18 is configured to refine the first driving expression image and the second driving expression image obtained according to the training expression sample.
[0183] In this embodiment, the sample expression refining module 18 can be configured to perform Figure 6 As shown in steps A2-A4, the specific description of the sample expression refining module 18 can refer to the description of steps A2-A4.
[0184] The refining model updating module 19 is configured to update the expression refining model.
[0185] In this embodiment, the refining model updating module 19 can be configured to perform Figure 6 As shown in step A5, and Figure 7 As shown in steps A51-A54, the specific description of the refining model updating module 19 can refer to the description of step A5 and steps A51-A54.
[0186] The target expression refining module 20 is configured to process the target expression image through the trained expression refining model to obtain a target refined expression image.
[0187] In this embodiment, the target expression refining module 20 can be configured to perform Figure 6 As shown in step A6, the specific description of the target expression refining module 20 can refer to the description of step A6.
[0188] The embodiment also provides an electronic device, Figure 9 A structural diagram of the electronic device of the embodiment is shown, which includes a memory 31 and a processor 32, the memory 31 stores computer readable instructions, and the processor 32 executes the computer readable instructions to realize the expression generation method of the embodiment.
[0189] Preferably, the electronic device further includes a bus 33 and a communication interface 34, and the processor 32, the communication interface 34 and the memory 31 are connected through the bus 33.
[0190] The memory 31 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 34 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 33 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus 33 can be divided into an address bus, a data bus, a control bus, etc. (not fully drawn in the figure).
[0191] The processor 32 can be an integrated circuit chip with signal processing capability. In the specific implementation process, the steps in the embodiments of the above method can be completed by the integrated logic circuit of hardware or the instructions in the form of software in the processor 32. The above processor 32 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, which can realize or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor 32 can also be any conventional processor 32, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory 31, and the processor 32 reads the information in the memory 31 and combines the hardware to complete the steps of the method in the above embodiments.
[0192] The embodiments of the present application also provide a computer readable storage medium, which stores computer executable instructions, and when the computer executable instructions are called and executed by the processor 32, the computer executable instructions cause the processor 32 to implement the above-mentioned expression generation method. For specific implementation, reference can be made to the embodiments described above, and will not be repeated here.
[0193] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0194] Obviously, the above embodiments of the present application are only examples for clearly illustrating the technical solutions of the present application, and are not intended to limit the specific embodiments of the present application. Any modification, equivalent replacement, and improvement within the spirit and principles of the claims of the present application shall be included in the protection scope of the claims of the present application.
Claims
1. A method for generating facial expressions, characterized in that, The generation method includes: Several training expression samples are collected, and each training expression sample contains a first expression image sample and a second expression image sample; the objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, but the expressions of the objects are different. Build an emoji generation model; The expression generation model is used to process the first expression image sample corresponding to each training expression sample to obtain the first expressionless frontal key points and the first expression facial key points. The expression generation model is used to process the second expression image sample corresponding to each training expression sample to obtain the second expressionless frontal key points and the second expression facial key points. First expression transfer data is obtained based on the second facial key points of the expression and the first expressionless frontal key points; second expression transfer data is obtained based on the first facial key points of the expression and the second expressionless frontal key points. The expression generation model obtains a first driving expression image of the first expression image sample based on the first expression transfer data, and obtains a second driving expression image of the second expression image sample based on the second expression transfer data. The expression generation model is updated based on the first driving expression image and the second driving expression image to obtain the trained expression generation model. Acquire a target base image and a target expression-driven image, and input the target base image and the target expression-driven image into the trained expression generation model to obtain the target expression image.
2. The facial expression generation method according to claim 1, characterized in that, The step of obtaining a first driving expression image of the first expression image sample based on the first expression transfer data and obtaining a second driving expression image of the second expression image sample based on the second expression transfer data through the expression generation model specifically includes: Input the first expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding first object facial features; The second expression image sample corresponding to each training expression sample is input into the expression generation model to obtain the corresponding second object facial features; the expression generation model obtains the first driving expression image based on the first expression transfer data and the first object facial features. The expression generation model obtains the second driving expression image based on the second expression transfer data and the facial features of the second object.
3. The facial expression generation method according to any one of claims 1-2, characterized in that, After collecting several training facial expression samples, the process also includes: Add eye key point labels, eyebrow key point labels, and mouth key point labels to the first expression image sample, and add eye key point labels, eyebrow key point labels, and mouth key point labels to the second expression image sample.
4. The facial expression generation method according to claim 3, characterized in that, The step of updating the expression generation model based on the first driving expression image and the second driving expression image specifically includes: Calculate the expressionless loss based on the first and second expressionless frontal face key points; Based on the first facial key points of the expression and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, calculate the first eye loss, the first eyebrow loss and the first mouth loss respectively; Based on the facial key points of the second expression and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, calculate the second eye loss, the second eyebrow loss and the second mouth loss respectively; Calculate the first reconstruction loss based on the first driving expression image and the second expression image samples; The second reconstruction loss is calculated based on the second driving expression image and the first expression image sample; Based on the first reconstruction loss and the second reconstruction loss, the image reconstruction loss is obtained; The expression generation model is updated based on the expressionless loss, and / or the first eye loss, and / or the first eyebrow loss, and / or the first mouth loss, and / or the second eye loss, and / or the second eyebrow loss, and / or the second mouth loss, and / or the image reconstruction loss.
5. The facial expression generation method according to claim 4, characterized in that, Before acquiring the target base image and the target expression-driven image, the process also includes: Build a facial expression refinement model; The first eye coefficient, the first eyebrow coefficient, and the first mouth coefficient are obtained based on the first facial key points of the first expression; the second eye coefficient, the second eyebrow coefficient, and the second mouth coefficient are obtained based on the second facial key points of the second expression. The first driven expression image, the first eye coefficient, the first eyebrow coefficient, and the first mouth coefficient are processed using the expression refinement model to obtain the first expression refinement image; The second driving expression image, the second eye coefficient, the second eyebrow coefficient, and the second mouth coefficient are processed using the expression refinement model to obtain the second expression refinement image; The expression refinement model is updated based on the first and second expression refinement images to obtain the trained expression refinement model. After acquiring the target facial expression image, the process further includes: The target facial expression image is processed by the trained facial expression retouching model to obtain the target retouched facial expression image.
6. The facial expression generation method according to claim 5, characterized in that, The step of updating the expression refinement model based on the first and second expression refinement images specifically includes: Calculate the first retouching loss based on the first retouched image and the second retouched image samples; The second retouching loss is calculated based on the second retouched image and the first retouching image sample; Based on the first refinement loss and the second refinement loss, obtain the image refinement and reconstruction loss; The expression refinement model is updated based on the image refinement and reconstruction loss.
7. An expression generation device, characterized in that, The generating apparatus includes: The sample acquisition module is used to acquire several training expression samples, each training expression sample containing a first expression image sample and a second expression image sample; the objects in the first expression image sample and the second expression image sample of the same training expression sample are the same, but the expressions of the objects are different. The generative model building module is used to build facial expression generation models; The data processing module is used to process the first expression image sample corresponding to each training expression sample using the expression generation model to obtain a first expressionless frontal key point and a first expression facial key point; to process the second expression image sample corresponding to each training expression sample using the expression generation model to obtain a second expressionless frontal key point and a second expression facial key point; to obtain first expression transfer data based on the second expression facial key point and the first expressionless frontal key point; and to obtain second expression transfer data based on the first expression facial key point and the second expressionless frontal key point. An expression-driven module is used to obtain a first driving expression image of the first expression image sample based on the first expression transfer data, and to obtain a second driving expression image of the second expression image sample based on the second expression transfer data, using the expression generation model. A model update module is used to update the expression generation model based on the first driving expression image and the second driving expression image to obtain the trained expression generation model. The target expression generation module is used to acquire a target base image and a target expression driving image, and input the target base image and the target expression driving image into the trained expression generation model to obtain the target expression image.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the expression generation method according to any one of claims 1-6.
9. A computer storage medium, characterized in that, It stores a computer-readable program thereon, which, when executed, implements an expression generation method according to any one of claims 1-6.
Citation Information
Patent Citations
Expression driving method and device, electronic equipment and storage medium
CN113870399A
Method for training expression driving generation model, expression driving method and device
CN115512014A