Expression generation method and device, electronic equipment and storage medium

By collecting image samples with different expressions of the same object and using migration data of the expression generation model, the problem of poor voice-driven three-dimensional expressions in the prior art is solved, and a more natural and general expression generation effect is achieved.

CN120014094AActive Publication Date: 2025-05-16GUANGZHOU HUYA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510166820.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In the prior art, the effect of driving three-dimensional expressions through voice is poor, mainly because the voice contains less expression information, resulting in less naturalness and insufficient universality.

Method used

The first and second expression images are collected that contain the same object but have different expressions, and the migration data between these samples is obtained through the expression generation model, and the model is updated to generate more accurate expression images.

Benefits of technology

By increasing the number of training samples and using expression migration data, the expression generation model can more effectively utilize expression information to generate target expression images of corresponding expressions on the basic images, improving the naturalness and universality of expression generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014094A_ABST
    Figure CN120014094A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image data processing, in particular to an expression generation method and device, electronic equipment and a storage medium, and the generation method comprises the steps: collecting a first expression image sample and a second expression image sample which comprise the same object and have different object expressions; obtaining migration data between the first expression image sample and the second expression image sample through the constructed expression generation model, obtaining a first driving expression image and a second driving expression image of the first expression image sample and the second expression image sample according to the migration data, and completing the training of the expression generation model. And obtaining a target expression image according to the target basic image and the target expression driving image by using the trained expression production model. Compared with the prior art, the method has the advantages that the first expression image sample and the second expression image sample are mutually referenced, the expression generation model is fully learned, and the expression information in the image can be effectively utilized to generate the target expression image of the corresponding expression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing, and more specifically, to an expression generation method, device, electronic equipment and storage medium. Background Art

[0002] With the continuous development of science and technology, computer vision and artificial intelligence technologies have achieved remarkable results in various fields. Among them, 3D expression driving technology has broad application prospects in virtual reality, animation production, game development and other fields. As an important technology in 3D expression driving, expression generation has attracted more and more attention from researchers. In the existing technology, 3D expressions are usually driven by voice, but the expression information contained in voice is relatively small, so the effect of driving expressions by voice is poor. Summary of the invention

[0003] The present invention aims to overcome at least one defect of the above-mentioned prior art and provide an expression generation method, device, electronic device and storage medium, which can more effectively generate three-dimensional expressions.

[0004] According to one aspect of the present application, a method for generating an expression is provided, the method comprising:

[0005] Collecting a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions of the object;

[0006] Build an expression generation model;

[0007] Obtaining migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model;

[0008] Obtaining, by means of the expression generation model, a first driven expression image of the first expression image sample according to the corresponding migration data, and obtaining a second driven expression image of the second expression image sample;

[0009] updating the expression generation model according to the first driving expression image and the second driving expression image to obtain the trained expression generation model;

[0010] A target basic image and a target expression driving image are collected, and the target basic image and the target expression driving image are input into the trained expression generation model to obtain a target expression image.

[0011] Optionally, obtaining the migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model specifically includes:

[0012] Using the expression generation model, the first expression image sample corresponding to each training expression sample is processed to obtain the first expressionless front face key points and the first expression facial key points;

[0013] Processing the second expression image sample corresponding to each training expression sample using the expression generation model to obtain second expressionless front face key points and second expression facial key points;

[0014] The first expression migration data is obtained according to the second expression facial key points and the first expressionless frontal face key points, and the second expression migration data is obtained according to the first expression facial key points and the second expressionless frontal face key points.

[0015] Optionally, the acquiring, through the expression generation model, a first driven expression image of the first expression image sample according to the corresponding migration data, and acquiring a second driven expression image of the second expression image sample specifically includes:

[0016] Inputting the first expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding first object facial features;

[0017] Inputting the second expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding second object facial features; the expression generation model obtains the first driving expression image according to the first expression migration data and the first object facial features;

[0018] The expression generation model obtains the second driving expression image according to the second expression migration data and the second object facial feature data.

[0019] Optionally, after collecting a number of training expression samples, the step further includes adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples;

[0020] The step of adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples is specifically as follows:

[0021] Add eye key point labels, eyebrow key point labels and mouth key point labels to the first expression image sample, and add eye key point labels, eyebrow key point labels and mouth key point labels to the second expression image sample;

[0022] The step of obtaining the migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples is specifically as follows:

[0023] Acquire migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples after adding the key point labels.

[0024] Optionally, updating the expression generation model according to the first drive expression image and the second drive expression image specifically includes:

[0025] Calculating the expressionless loss according to the first expressionless frontal face key point and the second expressionless frontal face key point;

[0026] According to the first facial expression key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively calculating a first eye loss, a first eyebrow loss and a first mouth loss;

[0027] According to the second facial expression key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively calculating the second eye loss, the second eyebrow loss and the second mouth loss;

[0028] Calculate a first reconstruction loss according to the first driving expression image and the second expression image sample;

[0029] Calculating a second reconstruction loss according to the second driving expression image and the first expression image sample;

[0030] Obtaining an image reconstruction loss according to the first reconstruction loss and the second reconstruction loss;

[0031] The expression generation model is updated according to the expressionless loss, and / or the first eye loss, and / or the first eyebrow loss, and / or the first mouth loss, and / or the second eye loss, and / or the second eyebrow loss, and / or the second mouth loss, and / or the image reconstruction loss.

[0032] Optionally, before acquiring the target basic image and the target expression driving image, the process further includes:

[0033] Construct expression refinement model;

[0034] Acquire a first eye coefficient, a first eyebrow coefficient, and a first mouth coefficient according to the first facial expression key points, and acquire a second eye coefficient, a second eyebrow coefficient, and a second mouth coefficient according to the second facial expression key points;

[0035] The first driving expression image, the first eye coefficient, the first eyebrow coefficient and the first mouth coefficient are processed by using the expression refinement model to obtain a first expression refinement image;

[0036] The second driving expression image, the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient are processed by using the expression refinement model to obtain a second expression refinement image;

[0037] Update the expression refinement model according to the first expression refinement image and the second expression refinement image to obtain the trained expression refinement model;

[0038] After obtaining the target expression image, the method further includes:

[0039] The target expression image is processed by the trained expression refinement model to obtain a target refined expression image.

[0040] Optionally, updating the expression refinement model according to the first expression refinement image and the second expression refinement image specifically includes:

[0041] Calculate a first refinement loss according to the first refined expression image and the second expression image sample;

[0042] Calculating a second refinement loss according to the second refined expression image and the first expression image sample;

[0043] Obtaining an image refinement and reconstruction loss according to the first refinement loss and the second refinement loss;

[0044] The expression refinement model is updated according to the image refinement reconstruction loss.

[0045] According to a second aspect of the present application, there is provided an expression generating device, the generating device comprising:

[0046] A sample collection module is used to collect a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions;

[0047] Generative model building module, used to build expression generation model;

[0048] A data processing module, used for obtaining migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model;

[0049] An expression driving module, used for obtaining a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding migration data through the expression generation model;

[0050] A generation model updating module, used for updating the expression generation model according to the first driving expression image and the second driving expression image to obtain the trained expression generation model;

[0051] The target expression generation module is used to collect a target basic image and a target expression driving image, input the target basic image and the target expression driving image into the trained expression generation model, and obtain a target expression image.

[0052] According to a third aspect of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement an expression generation method described in the first aspect above.

[0053] According to a fourth aspect of the present application, a computer storage medium is provided, on which a computer-readable program is stored. When the computer-readable program is executed, the expression generation method described in the first aspect is implemented.

[0054] According to any one of the above aspects, the present application provides an expression generation method, device, electronic device and storage medium, which collect a first expression image sample and a second expression image sample containing the same object and different expressions as training expression samples, and obtain the migration data between the first expression image sample and the second expression image sample through an expression generation model. On the one hand, the first expression image sample and the second expression image sample can be used as training samples at the same time, thereby increasing the number of samples trained by the expression generation model. On the other hand, the first expression image sample and the second expression image sample can be used as references to each other, thereby enabling the expression generation model to fully learn, so that the finally trained expression generation model can effectively and fully utilize the expression information in the image containing the target expression to generate a target expression image of the corresponding expression on the base image.

[0055] Furthermore, the expression generation method, device, electronic device and storage medium provided by the present application construct an expression refinement model, and further refine the facial expression through the corresponding eye coefficients, eyebrow coefficients and mouth coefficients through the constructed expression refinement model. The image refined by the expression refinement model has more accurate and vivid expression details. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 A schematic diagram of an application scenario of the generation method provided in this embodiment.

[0058] Figure 2 The steps of the generation method provided in this embodiment are as follows Figure 1 .

[0059] Figure 3 This is a flow chart of the steps for obtaining migration data provided in this embodiment.

[0060] Figure 4 This is a flow chart of the steps of obtaining the first driving expression image and the second driving expression image provided in this embodiment.

[0061] Figure 5 A flowchart of the steps for updating the expression generation model provided in this embodiment.

[0062] Figure 6 The steps of the generation method provided in this embodiment are as follows Figure 2 .

[0063] Figure 7 A flowchart of the steps for updating the expression refinement model provided in this embodiment.

[0064] Figure 8 This is a device structure diagram of the generating device provided in this embodiment.

[0065] Fig. 9 This is a device structure diagram of the electronic device provided in this embodiment.

[0066] Figures are labeled: server 100, terminal 200, sample collection module 11, generation model construction module 12, data processing module 13, expression driving module 14, generation model update module 15, target expression generation module 16, refinement model construction module 17, sample expression refinement module 18, refinement module update module 19, target expression refinement module 20, memory 31, processor 32, bus 33, communication interface 34. DETAILED DESCRIPTION

[0067] The drawings of this application are only used for illustrative purposes and should not be construed as limiting the present application. In order to better illustrate the following embodiments, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; it is understandable to those skilled in the art that some well-known structures and their descriptions in the drawings may be omitted.

[0068] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0069] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0070] Example 1

[0071] With the continuous development of science and technology, computer vision and artificial intelligence technologies have achieved remarkable results in various fields. Among them, 3D expression driving technology has broad application prospects in virtual reality, animation production, game development and other fields. As an important technology in 3D expression driving, expression generation has attracted more and more attention from researchers.

[0072] At present, expression generation technology is mainly based on speech generation. By analyzing the features of images and speech in the video, the corresponding expression information features are obtained according to the facial expressions and speech of the objects in the corresponding images, and the three-dimensional facial expressions are reconstructed based on the extracted expression information features. However, the speech-based method has the following limitations:

[0073] Very few signals: The voice signal contains limited expression information, resulting in a relatively simple expression that is difficult to meet the needs of complex scenarios.

[0074] Poor effect: Due to the lack of expression information in the voice signal, the generated expression is less natural and can easily make users feel uncomfortable.

[0075] Lack of universality: Speech signals of different languages ​​and different speakers vary greatly, making it difficult for speech-based methods to achieve universality.

[0076] This embodiment provides a technical solution that can solve the above-mentioned problem. The specific implementation methods of this application are described in detail below with reference to the accompanying drawings.

[0077] For example, the following is a schematic diagram of an application scenario of an expression generation method provided in an embodiment of the present application. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a data transmission function for video streams and audio streams; the terminal 200 has a streaming media playback function and can also have an image processing function.

[0078] It is understandable that the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a car terminal, etc., but is not limited thereto.

[0079] In an operative manner, the server 100 and the terminal 200 may respectively execute an expression generation method provided in an embodiment of the present application, or, optionally, the expression generation method provided in an embodiment of the present application is partially executed in the server 100 and partially executed in the terminal 200.

[0080] like Figure 2 As shown, this embodiment provides an expression generation method, and the expression generation method may specifically include:

[0081] S1: Collect a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; wherein the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions;

[0082] In this embodiment, the collection of the training expression samples can be carried out by collecting videos containing facial expressions of the same object, and taking two frames of video images of the same object but with different expressions as the first expression image sample and the second expression image sample, respectively; it is also possible to directly collect two facial images of the same object but with different expressions through an image acquisition device such as a camera, and use them as the first expression image sample and the second expression image sample, respectively.

[0083] It is understandable that for the first expression image sample and the second expression image sample in the same training expression sample, it is necessary to ensure that their objects are the same, and for different training expression samples, their corresponding objects may be different. In order to improve the versatility of model expression generation, it is necessary to collect the training expression samples of different objects. The more objects collected, the higher the versatility of the model obtained by subsequent training.

[0084] It is understandable that most facial expressions are transmitted through the eyes, eyebrows and mouth. In order to better learn the features of the eyes, eyebrows and mouth of the face, in this embodiment, after collecting a plurality of the training expression samples, it also includes adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples;

[0085] Specifically, adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples is specifically:

[0086] Eye key point labels, eyebrow key point labels and mouth key point labels are added to the first expression image sample, and eye key point labels, eyebrow key point labels and mouth key point labels are added to the second expression image sample.

[0087] Specifically, for all the training expression samples, eye key point labels are added to the eye key points in each of the first expression images, eyebrow key point labels are added to the eyebrow key points, and mouth key point labels are added to the mouth key points. Correspondingly, eye key point labels are added to the eye key points in each of the second expression image samples, eyebrow key point labels are added to the eyebrow key points, and mouth key point labels are added to the mouth key points; no labels are added to other key points of the first expression image and the second expression image. By adding labels to the key points of the eyebrow and mouth parts and not adding labels to the key points of other parts, the model can be semi-supervised based on the training expression samples. On the one hand, the model can effectively learn based on key points such as the eyebrow and mouth, improve the accuracy of the eyebrow and mouth part generation, and thus make the generated expressions more accurate and vivid. On the other hand, the model's generalization ability for faces of different objects can be improved, and expressions of different objects can be effectively generated.

[0088] S2: Construct expression generation model;

[0089] S3: obtaining migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model;

[0090] As described above, in order to improve the accuracy of eyebrow and mouth part generation and enable the expression generation model to perform semi-supervised learning training, preferably, the migration data between the first expression image sample and the second expression image sample corresponding to each training expression sample is obtained, specifically:

[0091] Acquire migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples after adding the key point labels.

[0092] In this embodiment, based on each of the training expression samples after adding the key point labels, such as Figure 3 As shown, the step S3 may specifically include:

[0093] S31: using the expression generation model to process the first expression image sample corresponding to each training expression sample to obtain first expressionless frontal face key points and first expression facial key points;

[0094] It can be understood that the first expressionless frontal facial key points represent the facial key points of the corresponding object when it has no expression, and the first expression facial key points represent the facial key points of the corresponding object when it has expression.

[0095] S32: using the expression generation model to process the second expression image sample corresponding to each training expression sample to obtain second expressionless front face key points and second expression facial key points;

[0096] Corresponding to the acquisition result based on the first expression image sample, the second expressionless frontal face key points represent the facial key points of the corresponding object when it has no expression, and the second expression facial key points represent the facial key points of the corresponding object when it has expression.

[0097] The first expression image sample and the second expression image sample are of the same object, so for the same training expression sample, the first expressionless frontal face key points and the second expressionless frontal face key points obtained therefrom should be the same.

[0098] S33: Acquire first expression transition data according to the second expression facial key points and the first expressionless frontal face key points, and acquire second expression transition data according to the first expression facial key points and the second expressionless frontal face key points.

[0099] In this embodiment, the first expression migration data and the second expression migration data represent changes in expression. Specifically, the first expression migration data is a key point after migration, which can represent changes in the second expression facial key point compared to the first expressionless frontal key point, and correspondingly, the second expression migration data is a key point after migration, which can represent changes in the first expression facial key point compared to the second expressionless key point.

[0100] Specifically, the calculation of the first expression migration data can be expressed as follows:

[0101]

[0102] The calculation of the second expression transition data can be expressed as:

[0103]

[0104] In the formula, X1 represents the first expression migration data, X2 represents the second expression migration data; s is a scalar, indicating the size of the face, s1 represents the face size corresponding to the first expression image sample, and s2 represents the face size corresponding to the second expression image sample; N is a vector, indicating the key points of the expressionless front face, N1 represents the first expressionless front face key points, and N2 represents the second expressionless front face key points; B1 represents the first expression migration data, corresponding to the expression of the second expression image sample, and B2 represents the second expression migration data, corresponding to the expression of the first expression image sample; T1 represents the displacement of the face in the first expression image sample relative to the front face, and T2 represents the displacement of the face in the second expression image sample relative to the front face.

[0105] S4: obtaining, by means of the expression generation model, a first driving expression image of the first expression image sample according to the corresponding migration data, and obtaining a second driving expression image of the second expression image sample;

[0106] Specifically, in this embodiment, Figure 4 As shown, the step S4 may specifically include:

[0107] S41: inputting the first expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding first object facial features;

[0108] It can be understood that the facial features of the first object represent specific facial features of the corresponding object, such as skin color, hair color, etc.

[0109] S42: inputting the second expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding second object facial features;

[0110] Corresponding to the facial features of the first object, the facial features of the second object represent specific facial features of the corresponding object.

[0111] S43: the expression generation model acquires the first driving expression image according to the first expression migration data and the first object facial feature data;

[0112] It can be understood that the first driven expression image represents a prediction of the expression included in the second expression image sample based on the first expressionless frontal face key points and using the first object facial features.

[0113] S44: The expression generation model obtains the second driving expression image according to the second expression migration data and the second object facial feature data.

[0114] It can be understood that the second driven expression image represents a prediction of the expression included in the first expression image sample based on the second expressionless front face key points and using the second object facial features.

[0115] The acquiring of the first driving expression image may specifically include:

[0116] The first expressionless frontal face key points are used to perform expression movement migration calculation according to the first expression migration data, and the first migration expression obtained after the calculation and the first object facial features are input into the expression generation model to generate the first driving expression image.

[0117] Correspondingly, the acquisition of the second driving expression image may specifically include:

[0118] The expression movement migration calculation is performed according to the second expression migration data through the second expression expression key points, and the second migration expression obtained after the calculation and the second object facial features are input into the expression generation model to generate the second driving expression image.

[0119] Based on the above steps S3 and S4, it can be understood that in this embodiment, the expression generation model at least includes an expression prediction network and an expression generation network. Further, the expression prediction network may include an expressionless key point prediction unit, an expression key point prediction unit and a facial feature encoding module.

[0120] Specifically, the expressionless key point prediction unit is used to obtain the first expressionless frontal face key points according to the first expression image samples, and to obtain the second expressionless frontal face key points according to the second expression image samples.

[0121] The expression key point prediction unit is used to obtain the first expression facial key points according to the first expression image sample, and to obtain the second expression facial key points according to the second expression image sample;

[0122] The facial feature encoding unit is used to obtain the facial features of the first object according to the first expression image sample, and to obtain the facial features of the second object according to the second expression image sample.

[0123] The expression generation network is used to generate the first driven expression image according to the first expression migration data based on the first expressionless frontal face key points and the first object facial features; and to generate the second driven expression image according to the second expression migration data based on the second expressionless frontal face key points and the second object facial features.

[0124] In a specific implementation of this embodiment, the expressionless key point prediction unit, the expression key point prediction unit and the facial feature encoding module can all be constructed using the existing EffcientNet as the network backbone;

[0125] The expression generation network can be constructed using the existing neural network encoder Unet.

[0126] S5: updating the expression generation model according to the first driving expression image and the second driving expression image to obtain the trained expression generation model;

[0127] In this step, the expression generation model is updated according to the first driving expression image and the second driving expression image, such as Figure 5 As shown, it may specifically include:

[0128] S51: calculating expressionless loss according to the first expressionless frontal face key point and the second expressionless frontal face key point;

[0129] As mentioned above, since the first expression image sample and the second expression image sample are the same object, theoretically, the first expressionless frontal face key points and the second expressionless frontal face key points should be the same, so by calculating the expressionless loss, the expression generation model can effectively obtain the expressionless frontal face key points of the input expression image.

[0130] S52: Calculate a first eye loss, a first eyebrow loss, and a first mouth loss respectively according to the first facial expression key points and the corresponding eye key point labels, eyebrow key point labels, and mouth key point labels;

[0131] S53: Calculate a second eye loss, a second eyebrow loss, and a second mouth loss respectively according to the second facial expression key points and the corresponding eye key point labels, eyebrow key point labels, and mouth key point labels;

[0132] By calculating the losses of the eyes, eyebrows and mouth respectively, the expression generation model can focus on learning the features of the eyes, eyebrows and mouth, thereby improving the accuracy of generating the eyes, eyebrows and mouth of the expression image.

[0133] S54: calculating a first reconstruction loss according to the first driving expression image and the second expression image sample;

[0134] S55: Calculating a second reconstruction loss according to the second driving expression image and the first expression image sample;

[0135] S56: Obtaining an image reconstruction loss according to the first reconstruction loss and the second reconstruction loss;

[0136] It can be understood that the first reconstruction loss and the second reconstruction loss represent the overall loss of the expression generation model in generating a driving expression image based on the input expression image.

[0137] In this embodiment, the image reconstruction loss can be expressed as:

[0138] Loss reconstruction =‖X′1-Image1‖2‖VGG(X′1)-VGG(Image1)‖2+‖X′2-Image2‖2‖VGG(G(X′2))-VGG(Image2)‖2

[0139] Where, X′1 represents the first driving expression image, X′2 represents the second driving expression image; VGG represents the pre-trained image feature extraction model, VGG(·) represents the image features extracted from the corresponding image; Loss reconstruction Represents the image reconstruction loss; Image1 represents the first expression image sample, and Image2 represents the second expression image sample.

[0140] S57: Update the expression generation model according to the expressionless loss, and / or the first eye loss, and / or the first eyebrow loss, and / or the first mouth loss, and / or the second eye loss, and / or the second eyebrow loss, and / or the second mouth loss, and / or the image reconstruction loss.

[0141] Preferably, the expression generation model is updated according to the expressionless loss, the first eye loss, the first eyebrow loss, the first mouth loss, the second eye loss, the second eyebrow loss, the second mouth loss and the image reconstruction loss, so that the obtained expression generation model is more stable.

[0142] In this embodiment, after completing the expression generation model update, Figure 6 As shown, it also includes:

[0143] A1: Build an expression refinement model;

[0144] In this embodiment, the expression refinement model can be constructed using the existing neural network encoder Unet;

[0145] A2: obtaining a first eye coefficient, a first eyebrow coefficient and a first mouth coefficient according to the first facial expression key points, and obtaining a second eye coefficient, a second eyebrow coefficient and a second mouth coefficient according to the second facial expression key points;

[0146] Specifically, the first eye coefficient and the second eye coefficient are eye closure ratio coefficients, the first eyebrow coefficient and the second eyebrow coefficient are eyebrow movement coefficients, the first mouth coefficient and the second mouth coefficient are eyebrow movement coefficients, and the value range of the eye closure ratio coefficient, the eyebrow movement coefficient and the mouth opening and closing coefficient is 0 to 1; preferably, the eye closure ratio coefficient, the eyebrow movement coefficient and the eyebrow movement coefficient can be obtained by extracting the facial expression coefficients in the corresponding first expression image sample and the second expression image sample.

[0147] A3: Using the expression refinement model, processing the first driving expression image, the first eye coefficient, the first eyebrow coefficient, and the first mouth coefficient to obtain a first expression refinement image;

[0148] A4: Using the expression refinement model, processing the second driving expression image, the second eye coefficient, the second eyebrow coefficient, and the second mouth coefficient to obtain a second expression refinement image;

[0149] The expression refinement model refines the first driving expression image according to the first eye coefficient, the first eyebrow coefficient and the first mouth coefficient, and obtains the refined first driving expression image as the first expression refinement image; and refines the second driving expression image according to the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient, and obtains the refined second driving expression image as the second expression refinement image. The expression refinement model is used to obtain specific expression coefficients of the areas that best reflect facial expressions, such as eyes, eyebrows and mouth, to refine facial expressions, so that the facial expressions finally obtained have more accurate and vivid expressions.

[0150] A5: updating the expression refinement model according to the first expression refinement image and the second expression refinement image to obtain the trained expression refinement model;

[0151] In this step, the expression refinement model is updated according to the first expression refinement image and the second expression refinement image, such as Figure 7 As shown, it may specifically include:

[0152] A51: Calculating a first refinement loss according to the first refined expression image and the second expression image sample;

[0153] A52: Calculating a second refined loss according to the second refined expression image and the first expression image sample;

[0154] A53: Obtaining an image refinement and reconstruction loss according to the first refinement loss and the second refinement loss;

[0155] In this embodiment, the image refinement and reconstruction loss can be expressed as:

[0156] Loss refined =‖X″1-Image1‖2‖VGG(X″1)-VGG(Image1)‖2+‖X″2-Image2‖2‖VGG(G(X″2))-VGG(Image2)‖2

[0157] Wherein, X″1 represents the first refined expression image, X″2 represents the second refined expression image; VGG represents the pre-trained image feature extraction model, VGG(·) represents the image features extracted from the corresponding image; Loss refined Represents the image refinement and reconstruction loss; Image1 represents the first expression image sample, and Image2 represents the second expression image sample.

[0158] A54: Update the expression refinement model according to the image refinement reconstruction loss.

[0159] In a specific implementation of the present embodiment, the expression generation model and the expression refinement model can be integrated into an expression generation and refinement large model. It is understandable that the training and updating of the expression generation and refinement large model includes two stages, the first stage is the expression generation training stage, and the second stage is the expression refinement training stage. During the training process of the first stage, the expression refinement model part in the expression generation and refinement large model is fixed, and the expression generation model part in the expression generation and refinement large model is updated through the collected training expression sample training. The specific training steps can refer to the specific description of the above steps S1-S5; during the training process of the second stage, the expression generation model part is fixed, and the expression refinement model part is updated through the first drive expression image and the second drive expression training generated in the first stage. The specific training steps can refer to the above steps A1-A5.

[0160] S6: collecting a target basic image and a target expression driving image, inputting the target basic image and the target expression driving image into the trained expression generation model, and obtaining a target expression image.

[0161] It is understandable that in this embodiment, after acquiring the target expression image through the expression generation model, the following steps are further included:

[0162] A6: Process the target expression image using the trained expression refinement model to obtain a target refined expression image.

[0163] In a specific implementation of this embodiment, the expression generation model obtains the corresponding target expressionless front face key points and target object facial features according to the target basic image, and obtains the corresponding target expression facial key points according to the target expression driving image; then, according to the target expressionless front face key points and target expression facial key points, the target migration data is obtained; the obtained target migration data and the target object facial features are input into the expression generation model to obtain the target driven expression image as the target expression image. Then, the target eye coefficient, target eyebrow coefficient and target mouth coefficient of the target basic image are obtained, and the target eye coefficient, target eyebrow coefficient and target mouth coefficient and the target expression image are input into the expression refinement model to obtain the refined target expression refined image.

[0164] In this embodiment, by collecting a first expression image sample and a second expression image sample containing the same object and different expressions of the object as training expression samples, the migration data between the first expression image sample and the second expression image sample is obtained through an expression generation model. On the one hand, the first expression image sample and the second expression image sample can be used as training samples at the same time, thereby increasing the number of samples trained by the expression generation model. On the other hand, the first expression image sample and the second expression image sample can be used as references to each other, thereby enabling the expression generation model to fully learn, so that the finally trained expression generation model can effectively and fully utilize the expression information in the image containing the target expression to generate a target expression image of the corresponding expression on the base image.

[0165] At the same time, this embodiment further refines the facial expression by constructing an expression refinement model for the expression coefficients of the eyebrow and mouth. The image refined by the expression refinement model has more accurate and vivid expression details.

[0166] Based on the same inventive concept, this embodiment also provides an expression generating device, such as Figure 8 As shown, the expression generating device may specifically include:

[0167] The sample collection module 11 is used to collect a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions;

[0168] In this embodiment, in this embodiment, the sample collection module 11 can be used to perform Figure 2 As shown in step S1 , for a detailed description of the sample collection module 11 , reference may be made to the description of step S1 .

[0169] A generation model building module 12, used to build an expression generation model;

[0170] In this embodiment, in this embodiment, the generation model building module 12 can be used to perform Figure 2 As shown in step S2, for the specific description of the generation model building module 12, reference may be made to the description of step S2.

[0171] A data processing module 13 is used to obtain migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model;

[0172] In this embodiment, the data processing module 13 can be used to execute Figure 2 Step S3 shown, and Figure 3 As shown in steps S31-S33, for the specific description of the data processing module 13, reference may be made to the description of step S3 and steps S31-S33.

[0173] An expression driving module 14, used for obtaining a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding migration data through the expression generation model;

[0174] In this embodiment, the expression driving module 14 can be used to execute Figure 2 Step S4 shown, and Figure 4 As shown in steps S41-S42, for the detailed description of the expression driving module 14, reference may be made to the description of step S3 and steps S41-S42.

[0175] A generation model updating module 15, used for updating the expression generation model according to the first driving expression image and the second driving expression image to obtain the trained expression generation model;

[0176] In this embodiment, the generation model updating module 15 can be used to perform Figure 2 Step S5 shown, and Figure 5 As shown in steps S51-S57, for the specific description of the generation model updating module 15, reference may be made to the description of step S5 and steps S51-S57.

[0177] A target expression generation module 16 is used to collect a target basic image and a target expression drive image, input the target basic image and the target expression drive image into the trained expression generation model, and obtain a target expression image;

[0178] In this embodiment, the target expression generation module 16 can be used to perform Figure 2 As shown in step S6, for the specific description of the target expression generating module 16, reference may be made to the description of step S6.

[0179] A refinement model building module 17, used to build a refinement model for facial expressions;

[0180] In this embodiment, the refined model building module 17 can be used to perform Figure 6 As shown in step A1, for the detailed description of the refined model building module 17, reference may be made to the description of step A1.

[0181] A sample expression refinement module 18, used for refining the first driven expression image and the second driven expression image obtained according to the training expression sample;

[0182] In this embodiment, the sample expression refinement module 18 can be used to perform Figure 6 For the detailed description of the sample expression refinement module 18 , please refer to the description of the steps A2 - A4 .

[0183] A refinement module update module 19, used to update the expression refinement model;

[0184] In this embodiment, the refinement module update module 19 can be used to execute Figure 6 Step A5 shown, and Figure 7 As shown in steps A51-A54, for the detailed description of the refinement module update module 19, please refer to the description of step A5 and steps A51-A54.

[0185] The target expression refinement module 20 is used to process the target expression image by using the trained expression refinement model to obtain a target refined expression image.

[0186] In this embodiment, the target expression refinement module 20 can be used to perform Figure 6 As shown in step A6, for the detailed description of the sample target expression refinement module 20, reference may be made to the description of step A6.

[0187] This embodiment also provides an electronic device, Fig. 9 The structure diagram of the electronic device of this embodiment is shown, which includes a memory 31 and a processor 32. The memory 31 stores computer-readable instructions, and the processor 32 executes the computer-readable instructions to implement the expression generation method of this embodiment.

[0188] Preferably, the electronic device further includes a bus 33 and a communication interface 34 , and the processor 32 , the communication interface 34 and the memory 31 are connected via the bus 33 .

[0189] The memory 31 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 34 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 33 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus 33 may be divided into an address bus, a data bus, a control bus, etc. (not fully drawn in the figure).

[0190] The processor 32 may be an integrated circuit chip with signal processing capability. In the specific implementation process, each step in the above method embodiment may be completed by the hardware integrated logic circuit or software instructions in the processor 32. The above processor 32 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, which may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor 32 may also be any conventional processor 32, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in a decoding processor. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 31, and the processor 32 reads the information in the memory 31 and completes the steps of the method of the above embodiment in combination with its hardware.

[0191] An embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor 32, the computer-executable instructions prompt the processor 32 to implement the above-mentioned expression generation method. The specific implementation can be found in the embodiment, which will not be repeated here.

[0192] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0193] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for generating an expression, characterized in that: The generation method comprises: Collecting a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions of the object; Construct expression generation model; Obtaining migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model; Obtaining, by means of the expression generation model, a first driven expression image of the first expression image sample according to the corresponding migration data, and obtaining a second driven expression image of the second expression image sample; Update the expression generation model according to the first drive expression image and the second drive expression image to obtain the trained expression generation model; A target basic image and a target expression driving image are collected, and the target basic image and the target expression driving image are input into the trained expression generation model to obtain a target expression image.

2. The expression generation method according to claim 1, characterized in that: The step of obtaining the migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model specifically includes: Using the expression generation model, the first expression image sample corresponding to each training expression sample is processed to obtain the first expressionless front face key points and the first expression facial key points; Processing the second expression image sample corresponding to each training expression sample using the expression generation model to obtain second expressionless front face key points and second expression facial key points; The first expression migration data is obtained according to the second expression facial key points and the first expressionless frontal face key points, and the second expression migration data is obtained according to the first expression facial key points and the second expressionless frontal face key points.

3. The expression generation method according to claim 2, characterized in that: The step of obtaining a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding migration data through the expression generation model specifically includes: Inputting the first expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding first object facial features; Inputting the second expression image sample corresponding to each training expression sample into the expression generation model to obtain the corresponding second object facial features; the expression generation model obtains the first driving expression image according to the first expression migration data and the first object facial features; The expression generation model obtains the second driving expression image according to the second expression migration data and the second object facial feature data.

4. The expression generation method according to any one of claims 1 to 3, characterized in that: After collecting a number of training expression samples, the method further includes adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples; The step of adding key point labels to the first expression image sample and the second expression image sample in each of the training expression samples is specifically as follows: Add eye key point labels, eyebrow key point labels and mouth key point labels to the first expression image sample, and add eye key point labels, eyebrow key point labels and mouth key point labels to the second expression image sample; The step of obtaining the migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples is specifically as follows: Acquire migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples after adding the key point labels.

5. A method for generating facial expressions according to claim 4, characterized in that: The updating of the expression generation model according to the first driving expression image and the second driving expression image specifically includes: Calculating the expressionless loss according to the first expressionless frontal face key point and the second expressionless frontal face key point; According to the first facial expression key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively calculating a first eye loss, a first eyebrow loss and a first mouth loss; According to the second facial expression key points and the corresponding eye key point labels, eyebrow key point labels and mouth key point labels, respectively calculating the second eye loss, the second eyebrow loss and the second mouth loss; Calculate a first reconstruction loss according to the first driving expression image and the second expression image sample; Calculating a second reconstruction loss according to the second driving expression image and the first expression image sample; Obtaining an image reconstruction loss according to the first reconstruction loss and the second reconstruction loss; The expression generation model is updated according to the expressionless loss, and / or the first eye loss, and / or the first eyebrow loss, and / or the first mouth loss, and / or the second eye loss, and / or the second eyebrow loss, and / or the second mouth loss, and / or the image reconstruction loss.

6. The expression generation method according to claim 5, characterized in that: Before acquiring the target basic image and the target expression driving image, the method further includes: Construct expression refinement model; Acquire a first eye coefficient, a first eyebrow coefficient, and a first mouth coefficient according to the first facial expression key points, and acquire a second eye coefficient, a second eyebrow coefficient, and a second mouth coefficient according to the second facial expression key points; The first driving expression image, the first eye coefficient, the first eyebrow coefficient and the first mouth coefficient are processed by using the expression refinement model to obtain a first expression refinement image; The second driving expression image, the second eye coefficient, the second eyebrow coefficient and the second mouth coefficient are processed by using the expression refinement model to obtain a second expression refinement image; Update the expression refinement model according to the first expression refinement image and the second expression refinement image to obtain the trained expression refinement model; After obtaining the target expression image, the method further includes: The target expression image is processed by the trained expression refinement model to obtain a target refined expression image.

7. The expression generation method according to claim 6, characterized in that: The updating of the expression refinement model according to the first expression refinement image and the second expression refinement image specifically includes: Calculate a first refinement loss based on the first refined expression image and the second expression image sample; Calculating a second refinement loss according to the second refined expression image and the first expression image sample; Obtaining an image refinement and reconstruction loss according to the first refinement loss and the second refinement loss; The expression refinement model is updated according to the image refinement reconstruction loss.

8. An expression generating device, characterized in that: The generating device comprises: A sample collection module is used to collect a number of training expression samples, each of which includes a first expression image sample and a second expression image sample; the first expression image sample and the second expression image sample of the same training expression sample have the same object, but different expressions; Generative model building module, used to build expression generation model; A data processing module, used for obtaining migration data between the first expression image sample and the second expression image sample corresponding to each of the training expression samples through the expression generation model; An expression driving module, used for obtaining a first driving expression image of the first expression image sample and a second driving expression image of the second expression image sample according to the corresponding migration data through the expression generation model; A generation model updating module, used for updating the expression generation model according to the first driving expression image and the second driving expression image to obtain the trained expression generation model; The target expression generation module is used to collect a target basic image and a target expression driving image, input the target basic image and the target expression driving image into the trained expression generation model, and obtain a target expression image.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement an expression generation method as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: A computer-readable program is stored thereon, and when the computer-readable program is executed, the expression generation method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Expression driving method and device, electronic equipment and storage medium

    CN113870399A

  • Expression migration method and device, electronic equipment and storage medium

    CN115330980A

  • Method for training expression driving generation model, expression driving method and device

    CN115512014A

  • Method and apparatus to perform facial expression recognition and training

    US20180144185A1