Model Training Method, Image Generation Method, Device, Equipment and Medium

Through joint training of the first network model and the second network model, combined with face key point images and voice information, the image generation model is optimized, and the problem of poor image details generation in the prior art is solved, and more realistic virtual digital human generation is achieved.

CN114724209BActive Publication Date: 2025-07-08HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210247058.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-07-08
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

The existing image generation methods have poor details in detail generation, especially the difficulty in realizing portrait details such as eyebrows and lip shapes, resulting in poor generation effects of virtual digital people.

Method used

By training the first network model, the first target model is obtained, and then the second network model and the first target model are jointly trained to generate an image generation model. The image generation model is optimized using the first training sample and the second training sample, including the combination of face key point images and speech information to improve the details and authenticity of image generation.

Benefits of technology

The effect of the image generation model in detail generation is improved, making the generated digital human images closer to the real image, and improving the generation quality of virtual digital humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724209B_ABST
    Figure CN114724209B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, an image generation method, an apparatus, a device, and a medium; wherein, the method includes: training a first network model based on a first training sample to obtain a trained first target model, the first training sample including a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image of the target image, and face key point data of the target image; training a second network model and the first target model based on a second training sample to obtain a trained image generation model, the second training sample including a second face image generated by the first target model and the target image. By first training the first network model to obtain the first target model and then jointly training the second network model and the first target model to obtain the image generation model, the embodiments of the present disclosure make the image detail generation effect better and are beneficial to improving the generation effect of digital human images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision and image processing, and particularly to a model training method, an image generation method, an apparatus, a device, and a medium. Background Art

[0002] In recent years, with the progress of technology, the generation technology of virtual digital humans has become increasingly mature. There are mainly two presentation methods for virtual digital humans: automatic broadcast type (such as news anchors, etc.) and interactive type (such as emotional digital humans, virtual robots, and various interactive assistants, etc.). In the generation process of virtual digital humans, the image generation of digital humans is a relatively important link.

[0003] The existing image generation methods are basically two types: one is to perform face stretching and transformation based on traditional methods, and this method has a poor processing effect on real human images with multi-point non-rigid shapes; the other is the existing model learning-based method to achieve the generation and transformation between faces, but this method has a poor generation effect on image details, especially it is difficult to achieve a realistic effect for details such as eyebrows and mouth shapes in portrait generation, resulting in a poor effect of the finally generated digital human image. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a model training method, an image generation method, an apparatus, a device, and a medium, which can train a first network model according to a first training sample to obtain a trained first target model, and jointly train a second network model and the first target model according to a second training sample to obtain a trained image generation model, so that the image generated by the image generation model has a better effect in terms of details and can be closer to the real image.

[0005] To achieve the above object, the technical solutions provided by the embodiments of the present disclosure are as follows:

[0006] In a first aspect, the present disclosure provides an image generation model training method, and the method includes:

[0007] Training a first network model based on a first training sample to obtain a trained first target model, where the first training sample includes a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image corresponding to the target image, and face key point data of the target image;

[0008] Training a second network model and the first target model based on a second training sample to obtain a trained image generation model, where the second training sample includes a second face image generated by the first target model and the target image.

[0009] As an optional implementation manner of an embodiment of the present disclosure, the first network model includes a first generative adversarial network model, and the first generative adversarial network model includes a first generator and a first discriminator;

[0010] Correspondingly, training the first network model based on the first training sample to obtain a trained first target model includes:

[0011] Inputting the first face image, the face key point image, and the face key point data of the target image in the first training sample into the first generator to obtain a first predicted image;

[0012] Inputting the first face image and the first predicted image into the first discriminator to obtain a first discrimination result;

[0013] Training the first generator based on a preset loss function according to the first predicted image, the target image, and the first discrimination result to obtain the first target model after training the first generator.

[0014] As an optional implementation manner of an embodiment of the present disclosure, training the first generator based on a preset loss function according to the first predicted image, the target image, and the first discrimination result to obtain the first target model after training the first generator includes:

[0015] Based on a first preset loss function, determining a first loss value corresponding to the first generator according to the first predicted image and the target image;

[0016] Based on a second preset loss function, determining a second loss value corresponding to the first discriminator according to the first discrimination result;

[0017] Adjusting the parameters of the first generator according to the first loss value and the second loss value until the first generator converges to obtain the first target model after training the first generator.

[0018] As an optional implementation manner of an embodiment of the present disclosure, the first network model includes a first generative adversarial network model, the second network model includes a second generative adversarial network model, the second generative adversarial network model includes a second generator and a second discriminator, and the structure of the second generative adversarial network model is different from the structure of the first generative adversarial network model;

[0019] Correspondingly, training the second network model and the first target model based on the second training sample to obtain a trained image generation model includes:

[0020] Input the second face image generated by the first target model in the second training sample into the second generator to obtain a second predicted image;

[0021] Input the second predicted image and the target image into the second discriminator to obtain a second discrimination result;

[0022] Based on a target loss function, train the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result to obtain a trained image generation model.

[0023] As an optional implementation manner of an embodiment of the present disclosure, the training the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result based on a target loss function to obtain a trained image generation model includes:

[0024] Based on a first target loss function, determine a third loss value corresponding to the second generator according to the second predicted image and the target image;

[0025] Based on a second target loss function, determine a fourth loss value corresponding to the second discriminator according to the second discrimination result;

[0026] Based on a third target loss function and the second predicted image, determine a fifth loss value corresponding to the second generator;

[0027] According to the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the second generator until the second generator converges to obtain a second target model after training the second generator;

[0028] Based on the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the first target model to obtain an adjusted first target model;

[0029] Construct the image generation model according to the adjusted first target model and the second target model.

[0030] In a second aspect, the present disclosure provides an image generation method, and the method includes:

[0031] Obtain a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data, where the target face key point data is predicted based on the face image to be predicted and corresponding voice information;

[0032] Input the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into an image generation model to obtain a corresponding target prediction image;

[0033] Wherein, the image generation model is trained based on the method described in any item of the first aspect.

[0034] In a third aspect, the present disclosure provides an image generation model training device, which includes:

[0035] A first target model determination module, configured to train a first network model based on a first training sample to obtain a trained first target model, where the first training sample includes a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image corresponding to the target image, and face key point data of the target image;

[0036] An image generation model determination module, configured to train a second network model and the first target model based on a second training sample to obtain a trained image generation model, where the second training sample includes a second face image generated by the first target model and the target image.

[0037] As an optional implementation manner of an embodiment of the present disclosure, the first network model includes a first generative adversarial network model, and the first generative adversarial network model includes a first generator and a first discriminator;

[0038] Correspondingly, the first target model determination module includes:

[0039] A first prediction unit, configured to input the first face image, the face key point image, and the face key point data of the target image in the first training sample into the first generator to obtain a first prediction image;

[0040] A first discrimination unit, configured to input the first face image and the first prediction image into the first discriminator to obtain a first discrimination result;

[0041] A first model determination unit, configured to train the first generator based on a preset loss function according to the first prediction image, the target image, and the first discrimination result to obtain a first target model after training the first generator.

[0042] As an optional implementation manner of an embodiment of the present disclosure, the first model determination unit is specifically configured to:

[0043] Based on a first preset loss function, determine a first loss value corresponding to the first generator according to the first prediction image and the target image;

[0044] Based on the second preset loss function, determine the second loss value corresponding to the first discriminator according to the first discrimination result;

[0045] According to the first loss value and the second loss value, adjust the parameters of the first generator until the first generator converges, and obtain the first target model after training the first generator.

[0046] As an optional implementation manner of the embodiments of the present disclosure, the first network model includes a first generative adversarial network model, the second network model includes a second generative adversarial network model, the second generative adversarial network model includes a second generator and a second discriminator, and the structure of the second generative adversarial network model is different from the structure of the first generative adversarial network model;

[0047] Correspondingly, the image generation model determination module includes:

[0048] A second prediction unit, configured to input the second face image generated by the first target model in the second training sample into the second generator to obtain a second prediction image;

[0049] A second discrimination unit, configured to input the second prediction image and the target image into the second discriminator to obtain a second discrimination result;

[0050] An image model determination unit, configured to train the second generator and the first target model based on a target loss function according to the second prediction image, the target image, and the second discrimination result, and obtain a trained image generation model.

[0051] As an optional implementation manner of the embodiments of the present disclosure, the image model determination unit is specifically configured to:

[0052] Based on a first target loss function, determine a third loss value corresponding to the second generator according to the second prediction image and the target image;

[0053] Based on a second target loss function, determine a fourth loss value corresponding to the second discriminator according to the second discrimination result;

[0054] Based on a third target loss function and the second prediction image, determine a fifth loss value corresponding to the second generator;

[0055] According to the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the second generator until the second generator converges, and obtain a second target model after training the second generator;

[0056] Based on the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the first target model to obtain an adjusted first target model;

[0057] Construct the image generation model according to the adjusted first target model and the second target model.

[0058] Fourthly, the present disclosure provides an image generation device, which includes:

[0059] An image acquisition module, configured to acquire a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data, where the target face key point data is predicted based on the face image to be predicted and corresponding voice information;

[0060] A predicted image generation module, configured to input the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into an image generation model to obtain a corresponding target predicted image;

[0061] Wherein, the image generation model is trained based on the method according to any one of the first aspect.

[0062] Fifthly, the present disclosure further provides a computer device, including:

[0063] One or more processors;

[0064] A storage device, configured to store one or more programs,

[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation model training method according to any one of the first aspect, or the image generation method according to the second aspect.

[0066] Sixthly, the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image generation model training method according to any one of the first aspect, or the image generation method according to the second aspect.

[0067] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art: First, the present disclosure trains a first network model based on a first training sample to obtain a trained first target model. The first training sample includes a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image corresponding to the target image, and face key point data of the target image. Then, based on a second training sample, the second network model and the first target model are trained to obtain a trained image generation model. The second training sample includes a second face image generated by the first target model and the target image, which solves the problem of poor image detail generation effect in the prior art. By first training the first network model to obtain the first target model, and then jointly training the second network model and the first target model, the obtained image generation model is more accurate, has a better effect in image detail generation, and is beneficial to improving the generation effect of digital human images. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0069] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative efforts.

[0070] Figure 1 It is a schematic diagram of the portrait transformation process in the prior art;

[0071] Figure 2 It is a schematic diagram of the software configuration of a computer device according to one or more embodiments of the present disclosure;

[0072] Figure 3A It is a schematic flowchart of a method for training an image generation model provided by an embodiment of the present disclosure;

[0073] Figure 3B It is a schematic diagram of the principle of a method for training an image generation model provided by an embodiment of the present disclosure;

[0074] Figure 3C It is a schematic diagram of the process of visualizing key points in an image in this embodiment;

[0075] Figure 4A It is a schematic flowchart of another method for training an image generation model provided by an embodiment of the present disclosure;

[0076] Figure 4BSchematic diagram of the principle for training a first network model provided by an embodiment of the present disclosure;

[0077] Figure 4C Another schematic diagram of the principle for training a first network model provided by an embodiment of the present disclosure;

[0078] Figure 4D Another schematic diagram of the principle for training a first network model provided by an embodiment of the present disclosure;

[0079] Figure 4E Schematic diagram of the structure of a generator provided by an embodiment of the present disclosure;

[0080] Figure 5A Schematic flowchart of another method for training an image generation model provided by an embodiment of the present disclosure;

[0081] Figure 5B Schematic diagram of the principle for training a second network model provided by an embodiment of the present disclosure;

[0082] Figure 5C Schematic diagram of the principle for training a second network model and a first target model provided by an embodiment of the present disclosure;

[0083] Figure 6A Schematic flowchart of an image generation method provided by an embodiment of the present disclosure;

[0084] Figure 6B Schematic diagram of the principle of an image generation method provided by an embodiment of the present disclosure;

[0085] Figure 6C Another schematic diagram of the principle of an image generation method provided by an embodiment of the present disclosure;

[0086] Figure 7A Schematic diagram of the structure of an image generation model training device provided by an embodiment of the present disclosure;

[0087] Figure 7B Schematic diagram of the structure of the first target model determination module in the image generation model training device of the embodiment of the present disclosure;

[0088] Figure 7C Schematic diagram of the structure of the image generation model determination module in the image generation model training device of the embodiment of the present disclosure;

[0089] Figure 8 Schematic diagram of the structure of an image generation device provided by an embodiment of the present disclosure;

[0090] Figure 9 Schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0091] In order to more clearly understand the above-mentioned objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0092] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0093] The terms "first" and "second" in the present disclosure are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first network model and the second network model are used to distinguish different network models, rather than to describe a specific order of the network models.

[0094] With the rapid development of intelligent technologies and the increasing popularity of intelligent terminals, multi-modal interactions such as voice, semantics, and images have become increasingly important ways. There are also more and more data such as images and videos, which are the main carriers of people's communication and life. In various forms of human-computer interaction and the experiences of various intelligent terminals, the processing of images, especially human portraits, has attracted particular attention. For example, in the process of generating virtual digital humans, the image generation of digital humans is a relatively important link. Therefore, high-quality human portrait generation has received much attention in the existing related research fields, but there are disadvantages such as poor robustness and insufficient realism in the existing technologies, making it difficult to popularize and experience on various intelligent terminals. Therefore, in-depth research and effect optimization for image generation are very crucial.

[0095] In image generation, especially in human portrait generation, the traditional method is based on the idea of triangulation, and stretching transformations are performed on multiple triangular points in the human portrait to achieve the shape of the target image. The advantage of this method is its fast speed and good effect for processing images such as anime and cartoons. However, when processing real human portraits, especially the face area, since the human face is non-rigid, the overall effect is poor due to the stretching and rendering methods.

[0096] Figure 1 It is a schematic diagram of the human portrait transformation process in the prior art. As Figure 1As shown below, its main implementation process is as shown in the following figure. Figure A is the current state image of a certain user, and Figure B is the next state image of Figure A. First, the face regions in Figure A are detected respectively to obtain the original face key points A. Then, the original face key points A are triangulated to obtain multiple divided planes (a plane is formed by three points). The target face key points B are also triangulated to obtain multiple divided planes. Then, the different planes obtained based on Figure A and the different planes obtained based on the target face key points B are clustered respectively. Furthermore, relevant rendering is performed according to the clustering situation of Figure A combined with the clustering situation of the target face key points B to obtain a new image (Figure B), that is, a portrait transformation is completed.

[0097] The above portrait transformation is mainly generated by Figure A and the target face key points B (landmarksB) through methods such as triangulation, clustering, patching, and rendering.

[0098] Aiming at the shortcomings in the above method, in the embodiments of the present disclosure, the first network model is first trained to obtain the first target model, and then the second network model and the first target model are jointly trained to obtain an image generation model. After obtaining the image generation model, a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data are obtained. The target face key point data is predicted based on the face image to be predicted and the corresponding voice information. The face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data are input into the image generation model to obtain the corresponding target prediction image. Image generation based on this image generation model enables the finally generated target prediction image to have better effects in terms of detail generation. This solution can be applied to multiple image processing fields such as portrait transformation, portrait generation, and portrait refinement, and can also be applied to the field of digital human image generation, which is beneficial to improving the generation effect of digital human images.

[0099] The image generation model training method and image generation method provided by the embodiments of the present disclosure can be implemented based on a computer device, or a functional module or functional entity in the computer device.

[0100] Among them, the computer device can be a personal computer (PC), server, mobile phone, tablet computer, laptop computer, mainframe computer, etc., and the embodiments of the present disclosure do not make specific limitations on this.

[0101] Exemplarily, Figure 2 is a software configuration schematic diagram of a computer device according to one or more embodiments of the present disclosure, as Figure 2As shown in the figure, the system is divided into four layers, from top to bottom are the Application layer (referred to as the "Application layer" for short), the Application Framework layer (referred to as the "Framework layer" for short), the Android runtime and the System Library layer (referred to as the "System Runtime Library layer" for short), and the Kernel layer.

[0102] The image generation model training method and the image generation method provided by the embodiments of the present application can be implemented based on the above computer device.

[0103] To describe the image generation model training solution in more detail, the following will be described in an exemplary manner in combination with Figure 3A It can be understood that Figure 3A The steps involved in may include more steps, or fewer steps, and the order of these steps may also be different, subject to the image generation model training method provided in the embodiments of the present application.

[0104] Figure 3A It is a schematic flowchart of a method for training an image generation model provided by an embodiment of the present disclosure. Figure 3B It is a schematic diagram of the principle of a method for training an image generation model provided by an embodiment of the present disclosure. This embodiment is applicable to the case of training a model to obtain an image generation model. The method of this embodiment can be executed by an image generation model training device, which can be implemented in a hardware / or software manner and can be configured in a computer device.

[0105] As Figure 3A shown, the method specifically includes the following steps:

[0106] S310, training a first network model based on a first training sample to obtain a trained first target model.

[0107] Among them, the first training sample includes a first face image, a face key-point image corresponding to the first face image, a target image, a target face key-point image corresponding to the target image, and face key-point data of the target image. The first training sample can be a training sample determined from a pre-determined training data set. The training data set can be multiple consecutive images intercepted from multimedia data, such as video clips, face key-point data corresponding to the multiple consecutive images, and multiple face key-point images obtained by visualizing each face key-point data, etc. The first face image can be a certain frame of face image determined from the training data set. The face key-point image corresponding to the first face image can be understood as an image obtained by visualizing the face key-point data of the first face image. The target image can be understood as the next frame of face image corresponding to the first face image. The face key-point data of the target image can be understood as the face key-point data extracted from the target image. The target face key-point image corresponding to the target image can be understood as a face key-point image obtained by visualizing the face key-point data of the target image. The face key-point data can be key-point data extracted based on different detailed regions of the face, such as eyebrows, eyes, nose, mouth, and face shape, etc. When extracting the face key-point data, it can be extracted through a face key-point detector, and the present disclosure places no limitations on the extraction method and the number of key points extracted. The first network model can be a generative adversarial network model or a U-net network model, etc., and the present disclosure places no limitations.

[0108] Training the first network model using the first training sample can obtain the trained first target model. Specifically, for the convenience of subsequent training of the second network model and the first target model based on the second training sample, when training the first network model, the first network model can be trained until it just converges, which can not only save time but also facilitate subsequent adjustment of the parameters of the first target model to make the parameters of the first target model reach the optimal.

[0109] S320, training the second network model and the first target model based on the second training sample to obtain the trained image generation model, where the second training sample includes a second face image generated by the first target model and the target image.

[0110] Among them, the second network model can be a generative adversarial network model implemented using lightweight depthwise separable convolutions, or a U-net network model, etc., and the present disclosure places no limitations. The second face image is the output of the first target model.

[0111] After obtaining the first target model, based on the target images in the first sample and the output of the first target model: the second face image, a second training sample is obtained. The second training sample is used to jointly train the second network model and the first target model, and then the trained image generation model is obtained. The image generation model includes a front-end generation model and a back-end fine-tuning model. Among them, the front-end generation model is the model obtained by training the first target model, and the back-end fine-tuning model is the model obtained by training the second network model.

[0112] In this embodiment, first, the first network model is trained based on the first training sample to obtain the trained first target model. Then, the second network model and the first target model are trained based on the second training sample to obtain the trained image generation model. Since the image generation model includes two parts and the face key point image is added to the model input, the extraction of face features is more refined during face feature extraction, solving the problem of poor image detail generation effect in the prior art. By first training the first network model to obtain the first target model, and then jointly training the second network model and the first target model, the obtained image generation model is more accurate, has a better effect in image detail generation, and is beneficial to improving the generation effect of digital human images.

[0113] In some embodiments, when performing face key point detection based on an existing face key point detector (the number of face key points can be 68, 74, 98, 106, or 212, etc.), the key points are generally 2D (x, y two-dimensional coordinates) or 3D (x, y, z three-dimensional coordinates) coordinate point information, and the spatial topological structure is poor. When generating and transforming the key attributes of the face in a portrait, the topological structure of different regions is particularly crucial. For face key points (face-landmarks), the contour information of several key regions in the face (such as eyebrows, eyes, nose, mouth, and face shape, etc.) can be outlined, thereby assisting the training of deep learning models, facilitating the generation and transformation of key regions in the face, and making the results more refined and realistic.

[0114] Specifically, the face key point image can be obtained through the following method:

[0115] Since the face key point image and the face key point data are consistent in the spatial dimension, the spatial coordinate information of the face key points can be drawn on an equal-sized Mask image (i.e., a 0, 1 black and white image), and the key points are outlined according to a predetermined drawing strategy, thereby obtaining the face key point image after image-based transformation of the face key point data.

[0116] The drawing strategy may include: 1) using a fully enclosed polygon or a non-fully enclosed polygon to draw different regions in the facial attributes, and different regions can be drawn with different color lines; 2) adjusting the drawing form and strategy according to the difference between the obtained facial key-point image and the original image. The forms include but are not limited to the following: in the subsequent model training process, according to the local contour information of the face, Mask maps are separately constructed for the eyebrows, eyes, nose, mouth, etc. for fine regression training to ensure that the generation effect of the image generation model is more realistic and delicate.

[0117] In some embodiments, considering that when training the first network model, a convolutional neural network is usually used to extract and optimize the features of the image, so when the facial key-point image obtained by imageizing the facial key points has three-channel information of Red Green Blue (RGB), it can be better combined with the convolutional neural network. In this way, the facial key-point information can be more effectively combined, positively promoting the authenticity and sense of detail of the generation result. At this time, both the first facial image in the first training sample and the facial key-point image corresponding to the first facial image can be converted into RGB-form images and then input into the first network model for training.

[0118] Optionally, Figure 3C is a schematic diagram of the process of imageizing the key points in the image in this embodiment. Figure 3C An exemplary implementation method is given in. The specific process of imageizing the key points has been described and will not be repeated here.

[0119] Figure 4A is a schematic flowchart of another image generation model training method provided by the embodiments of the present disclosure. This embodiment is further extended and optimized on the basis of the above embodiments. Optionally, a possible implementation method of S310 in this embodiment is as follows:

[0120] S3101, input the first facial image, the facial key-point image, and the facial key-point data of the target image in the first training sample into the first generator to obtain a first predicted image.

[0121] When the first network model includes a first generative adversarial network model, the first generative adversarial network model includes a first generator and a first discriminator.

[0122] Among them, the first generator can adopt a U-net network structure, which is designed based on a fully convolutional network. This first generator can support multi-channel and multi-resolution image inputs. The first discriminator can adopt a simple Encoder encoder. The essence of the first generative adversarial network is a process of game, where the first discriminator is used to determine whether the output of the first generator is real, and the two promote each other. The present disclosure does not limit the structures adopted by the first generator and the first discriminator.

[0123] Input the first face image, face key point image, and face key point data of the target image in the first training sample into the first generator, and then the first predicted image can be obtained.

[0124] Exemplarily, if the first generator adopts an Encoder-Decoder structure, the input first face image, face key point image, and face key point data of the target image can be encoded by the encoder and transcribed by the decoder to obtain the first predicted image.

[0125] In some embodiments, the first predicted image can also be obtained in the following way:

[0126] Input the first face image, target face key point image, and face key point data of the target image in the first training sample into the first generator to obtain the first predicted image.

[0127] The specific implementation method is similar to S3101 and will not be elaborated here.

[0128] It should be noted that the input of the first generator can also include: the first face image, face key point image, target face key point image, and face key point data of the target image, mainly for combining the data and images included in the first training sample. The present disclosure does not limit the specific combination method.

[0129] S3102, input the first face image and the first predicted image into the first discriminator to obtain the first discrimination result.

[0130] Input the first face image and the first predicted image into the first discriminator, and the first discrimination result can be obtained. Specifically, if the first discrimination result is 0, it is a fake image; if the first discrimination result is 1, it is a real image. The first discriminator can pay more attention to the generated realism during the training process, enabling the first generator to better fit the mapping relationship.

[0131] S3103, based on a preset loss function, train the first generator according to the first predicted image, target image, and the first discrimination result to obtain the first target model after training the first generator.

[0132] Among them, the preset loss function may include: L1 loss function (L1 loss), multi-scale structural similarity loss function (multi-scale Structural Similarity loss, abbreviated as MS SSIM loss), Style loss, Vgg loss, and generative adversarial GAN_loss, etc. The present disclosure does not make any limitations.

[0133] Through the preset loss function, the first generator is trained according to the first predicted image, the target image, and the first discrimination result. To save time, the first generator can be trained until it just converges, so as to obtain the first target model after training the first generator.

[0134] In this embodiment, obtaining the first target model through the above method is simple and efficient, which is beneficial to accelerating the subsequent model training process.

[0135] In some embodiments, the training of the first generator based on the preset loss function according to the first predicted image, the target image, and the first discrimination result to obtain the first target model after training the first generator includes:

[0136] Based on the first preset loss function, determine the first loss value corresponding to the first generator according to the first predicted image and the target image;

[0137] Based on the second preset loss function, determine the second loss value corresponding to the first discriminator according to the first discrimination result;

[0138] According to the first loss value and the second loss value, adjust the parameters of the first generator until the first generator converges, so as to obtain the first target model after training the first generator.

[0139] Among them, the first preset loss function may adopt at least one of L1 loss, MS SSIM loss, Style loss, Vgg loss, and G_gan_loss included in GAN_loss. The second preset loss function may adopt D_fake_loss and D_real_loss included in GAN_loss. The present disclosure does not make specific limitations on the first preset loss function and the second preset loss function.

[0140] Specifically, based on the first preset loss function, according to the first predicted image and the target image, the first loss value corresponding to the first generator can be obtained; based on the second preset loss function, according to the first discrimination result, the second loss value corresponding to the first discriminator can be obtained; by using the first loss value and the second loss value to adjust the parameters of the first generator until the first generator converges, the first target model after training the first generator is obtained.

[0141] In this embodiment, through the above method, the first generator can reach a converged state, reducing the error of the predicted image of the first generator.

[0142] Exemplarily, Figure 4B is a schematic diagram of the principle for training the first network model provided by an embodiment of the present disclosure. As Figure 4B shown, the training process is as follows:

[0143] The first network model includes a first generator and a first discriminator. The first face image and the face key point image are used as input 1, and the face key point data of the target image is used as input 2, and input into the first generator to train the first generator; the output first predicted image of the first generator and the first face image are input into the first discriminator to train the first discriminator; the parameters of the first generator are adjusted by the first discriminator loss (i.e., the second loss value corresponding to the first discriminator) and the first generator loss (i.e., the first loss value corresponding to the first generator) until the first generator converges, and the first target model after training the first generator is obtained.

[0144] Exemplarily, Figure 4B the input 1 in can also be the first face image and the target face key point image, or other combination methods for the images included in the first training sample, and the present disclosure does not make limitations.

[0145] Exemplarily, Figure 4C is another schematic diagram of the principle for training the first network model provided by an embodiment of the present disclosure. As Figure 4C shown, the training process is as follows:

[0146] The first network model includes a third generator. The third generator can adopt a U-net network structure or other structures, and the present disclosure does not limit this. The first face image and the face key point image are used as input 1, or the first face image and the target face key point image can also be used as input 1. The present disclosure does not limit the combination manner of the images in input 1. The face key point data of the target image is used as input 2 and input into the third generator to train the third generator; through a corresponding loss function, the third generator loss is determined based on the third predicted image output by the third generator and the target image; the parameters of the third generator are adjusted according to the third generator loss until the third generator converges, and the first target model after training the first generator is obtained.

[0147] Exemplarily, Figure 4D is a schematic diagram of the principle of training the first network model provided by an embodiment of the present disclosure. As Figure 4D shown:

[0148] The first face image and the target face key point image are used as input 1 and input into the fourth generator to train the fourth generator; the remaining processes are similar to the above training process and will not be elaborated here.

[0149] Exemplarily, Figure 4E is a schematic diagram of the structure of a generator provided by an embodiment of the present disclosure. As Figure 4E shown, the generator includes: an Encoder and a Decoder. Z can be understood as the output of the encoder. Specifically, in combination with Figure 4B and Figure 4C , input 1 can be input into the Figure 4E Encoder, and input 2 is input into Z for model training.

[0150] It should be noted that: the first generator, the third generator, and the fourth generator can all adopt the Figure 4E generator structure.

[0151] Figure 5A is a schematic flowchart of another image generation model training method provided by an embodiment of the present disclosure. This embodiment is further extended and optimized on the basis of the above embodiment. Optionally, a possible implementation manner of S320 in this embodiment is as follows:

[0152] S3201, input the second face image generated by the first target model in the second training sample into the second generator to obtain a second predicted image.

[0153] The first network model includes a first generative adversarial network model. The second network model includes a second generative adversarial network model. The second generative adversarial network model includes a second generator and a second discriminator. The structure of the second generative adversarial network model is different from that of the first generative adversarial network model.

[0154] Specifically, the differences between the structure of the second generative adversarial network model and that of the first generative adversarial network model are mainly as follows:

[0155] 1. It is implemented using lightweight depthwise separable convolutions;

[0156] 2. The number of short connections in the low-level feature layer is increased, thereby increasing the influence of low-level detailed features in the subsequent model training process.

[0157] The second generator can adopt a U-net network structure. The second discriminator can adopt a simple Encoder encoder. The present disclosure does not limit the structures adopted by the second generator and the second discriminator.

[0158] Inputting the second face image generated by the first target model in the second training sample into the second generator can obtain a second predicted image. This second predicted image has been finely adjusted by the second generator and can be closer to the target image.

[0159] Exemplarily, if the second generator adopts an Encoder-Decoder structure, the input second face image can be encoded by the encoder and transcribed by the decoder to obtain a second predicted image.

[0160] It should be noted that: the second generator can also adopt Figure 4E the generator structure.

[0161] S3202: Input the second predicted image and the target image into the second discriminator to obtain a second discrimination result.

[0162] Inputting the second predicted image and the target image into the second discriminator can obtain a second discrimination result. Specifically, if the second discrimination result is 0, it is a fake image; if the second discrimination result is 1, it is a real image. The second discriminator can pay more attention to the generated realism during the training process, enabling the second generator to better fit the mapping relationship.

[0163] S3203: Based on the target loss function, train the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result to obtain a trained image generation model.

[0164] Among them, the target loss function may include: L1 loss, MS SSIM loss, Style loss, Vggloss, generative adversarial GAN_loss, and an approximate loss of the image domain-based quality evaluation index SSIM, Peak Signal to Noise Ratio (PSNR), etc. The present disclosure places no restrictions.

[0165] By using the target loss function to train the second generator and the first target model based on the second predicted image, the target image, and the second discrimination result, a trained image generation model can be obtained. Since the first target model was previously trained on the first network model based on the first training sample, the parameters of the first target model can be adjusted at this time.

[0166] In this embodiment, by using the above method to obtain the image generation model, the effect of generating image details can be improved, which is beneficial to improving the authenticity and sense of details of the finally predicted image.

[0167] In some embodiments, the training of the second generator and the first target model based on the target loss function according to the second predicted image, the target image, and the second discrimination result to obtain the trained image generation model includes:

[0168] Based on the first target loss function, determine the third loss value corresponding to the second generator according to the second predicted image and the target image;

[0169] Based on the second target loss function, determine the fourth loss value corresponding to the second discriminator according to the second discrimination result;

[0170] Based on the third target loss function and the second predicted image, determine the fifth loss value corresponding to the second generator;

[0171] According to the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the second generator until the second generator converges to obtain the second target model after training the second generator;

[0172] Based on the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the first target model to obtain the adjusted first target model;

[0173] According to the adjusted first target model and the second target model, construct the image generation model.

[0174] Among them, the first target loss function can adopt at least one of L1 loss, MS SSIM loss, Style loss, Vgg loss, and G_gan_loss included in GAN_loss. The second preset loss function can adopt D_fake_loss and D_real_loss included in GAN_loss. The third target loss function can adopt an approximate loss based on SSIM and PSNR. The present disclosure does not specifically limit the first target loss function, the second target loss function, and the third target loss function.

[0175] Specifically, based on the first target loss function, according to the second predicted image and the target image, the third loss value corresponding to the second generator can be obtained; based on the second target loss function, according to the second discrimination result, the fourth loss value corresponding to the second discriminator can be obtained; based on the third target loss function and the second predicted image, the fifth loss value corresponding to the second generator can be obtained; according to the third loss value, the fourth loss value, and the fifth loss value, the parameters of the second generator are adjusted until the second generator converges, and the second target model after training the second generator is obtained; based on the third loss value, the fourth loss value, and the fifth loss value, the parameters of the first target model are adjusted to obtain the adjusted first target model; when the parameters of the adjusted first target model and the second target model both reach the optimal, the two are combined to obtain the image generation model.

[0176] In this embodiment, by jointly training the second network model and the first target model to make both reach the optimal state, the error of the image predicted by the image generation model can be reduced.

[0177] In some embodiments, the training strategies for the first network model and the second network model may include:

[0178] 1) Train in stages (adopt the generator-first strategy if the model includes a generator and a discriminator) and by tasks (the resolution of the image from low to high) to improve the generation quality of the generator relative to the discriminator;

[0179] 2) Optimize the learning rate adjustment according to the actual evaluation indicators of the training samples and the validation samples to improve the generalization ability of the model;

[0180] 3) Use the image data augmentation method to augment the training samples to enhance the data coverage and improve the generalization ability and robustness of the model.

[0181] Exemplarily, Figure 5B is a schematic diagram of the principle of training the second network model provided by the embodiments of the present disclosure. As Figure 5B shown, the training process is as follows:

[0182] The second network model includes a second generator and a second discriminator. The parameters of the second generator are adjusted by the second discriminator loss (i.e., the fourth loss value corresponding to the second discriminator) and the second generator loss (i.e., the third loss value and the fifth loss value corresponding to the second generator) until the second generator converges, and the second target model after training the second generator is obtained.

[0183] Exemplarily, Figure 5C FIG. is a schematic diagram of the principle for training the second network model and the first target model provided by an embodiment of the present disclosure. As Figure 5C shown, the training process is as follows:

[0184] On the Figure 5B basis, the parameters of the second generator are adjusted by the second discriminator loss and the second generator loss until the second generator converges, and the second target model after training the second generator is obtained. The parameters of the first target model are adjusted by the second discriminator loss and the second generator loss to obtain the adjusted first target model.

[0185] To describe the image generation scheme in more detail, the following will be described in an exemplary manner in combination with Figure 6A It can be understood that Figure 6A the steps involved in may include more steps or fewer steps in actual implementation, and the order between these steps may also be different, subject to the image generation method provided in the embodiments of the present application.

[0186] Figure 6A FIG. is a schematic flowchart of an image generation method provided by an embodiment of the present disclosure, Figure 6B FIG. is a schematic diagram of the principle of an image generation method provided by an embodiment of the present disclosure, Figure 6C FIG. is a schematic diagram of the principle of another image generation method provided by an embodiment of the present disclosure. This embodiment is applicable to situations such as image generation and face transformation. The method of this embodiment can be executed by an image generation device, and the device can be implemented in a hardware / or software manner and can be configured in a computer device.

[0187] As Figure 6A shown, the method specifically includes the following steps:

[0188] S610, obtaining a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data, where the target face key point data is predicted based on the face image to be predicted and the corresponding voice information.

[0189] Among them, the face image to be predicted can be the face image in the current state. The face key point image corresponding to the face image to be predicted can be an image obtained by visualizing the face key point data of the face image to be predicted. The target face key point data can be the key point data after the face change in the next state predicted based on the face image to be predicted and the corresponding voice information. Among them, the voice information can be pre-determined voice data.

[0190] In some embodiments, an image corresponding to the target face key point data can also be obtained.

[0191] S620, input the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into the image generation model to obtain the corresponding target prediction image.

[0192] Among them, the image generation model is trained based on the image generation model training method described in any embodiment. The target prediction image is the output result of the image generation model.

[0193] By inputting the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into the image generation model for prediction, the corresponding target prediction image can be obtained.

[0194] In this embodiment, through this image generation method, the generation of the image, especially the face transformation, can be realized quickly and accurately, and the finally generated image has a good effect in terms of detail generation.

[0195] In some embodiments, the target prediction image can also be obtained in the following manner:

[0196] Input the face image to be predicted, the image corresponding to the target face key point data, and the target face key point data into the image generation model for prediction to obtain the corresponding target prediction image.

[0197] Alternatively, input the face image to be predicted, the face key point image corresponding to the face image to be predicted, the image corresponding to the target face key point data, and the target face key point data into the image generation model for prediction to obtain the corresponding target prediction image.

[0198] It should be noted that: The present disclosure does not limit the combination manner of the images input into the image generation model. Exemplarily, as Figure 6B shown: Assume that the face image to be predicted is Figure C, and the face key point image corresponding to the face image to be predicted is Figure C1. Input Figure C, Figure C1, and the target face key point data into the image generation model, and the target prediction image can be obtained.

[0199] Exemplarily, as Figure 6CShown: Assume that the face image to be predicted is Figure C, and assume that the image corresponding to the target face key point data is Figure C2. Inputting Figure C, Figure C2, and the target face key point data into the image generation model can also obtain the target prediction image.

[0200] Figure 7A The following is a schematic structural diagram of an image generation model training device provided by an embodiment of the present disclosure; this device is configured in a computer device and can implement the image generation model training method described in any embodiment of the present application. Specifically, this device includes the following:

[0201] The first target model determination module 701 is used to train the first network model based on the first training sample to obtain the trained first target model. The first training sample includes the first face image, the face key point image corresponding to the first face image, the target image, the target face key point image corresponding to the target image, and the face key point data of the target image.

[0202] The image generation model determination module 702 is used to train the second network model and the first target model based on the second training sample to obtain the trained image generation model. The second training sample includes the second face image generated by the first target model and the target image.

[0203] As an optional implementation manner of an embodiment of the present disclosure, the first network model includes a first generative adversarial network model, and the first generative adversarial network model includes a first generator and a first discriminator.

[0204] Correspondingly, Figure 7B is a schematic structural diagram of the first target model determination module 701 in the image generation model training device of the embodiment of the present disclosure. As Figure 7B shown, the first target model determination module 701 includes:

[0205] The first prediction unit 7011 is used to input the first face image, the face key point image, and the face key point data of the target image in the first training sample into the first generator to obtain the first prediction image.

[0206] The first discrimination unit 7012 is used to input the first face image and the first prediction image into the first discriminator to obtain the first discrimination result.

[0207] The first model determination unit 7013 is used to train the first generator based on a preset loss function according to the first prediction image, the target image, and the first discrimination result to obtain the first target model after training the first generator.

[0208] As an alternative implementation manner of the embodiment of the present disclosure, the first model determination unit 7013 is specifically configured to:

[0209] Based on a first preset loss function, determine a first loss value corresponding to the first generator according to the first predicted image and the target image;

[0210] Based on a second preset loss function, determine a second loss value corresponding to the first discriminator according to the first discrimination result;

[0211] According to the first loss value and the second loss value, adjust the parameters of the first generator until the first generator converges, and obtain a first target model after training the first generator.

[0212] As an alternative implementation manner of the embodiment of the present disclosure, the first network model includes a first generative adversarial network model, the second network model includes a second generative adversarial network model, the second generative adversarial network model includes a second generator and a second discriminator, and the structure of the second generative adversarial network model is different from the structure of the first generative adversarial network model;

[0213] Correspondingly, Figure 7C is a schematic structural diagram of an image generation model determination module 702 in the image generation model training device according to an embodiment of the present disclosure. As Figure 7C shown, the image generation model determination module 702 includes:

[0214] A second prediction unit 7021, configured to input a second face image generated by the first target model in the second training sample into the second generator to obtain a second predicted image;

[0215] A second discrimination unit 7022, configured to input the second predicted image and the target image into the second discriminator to obtain a second discrimination result;

[0216] An image model determination unit 7023, configured to train the second generator and the first target model based on a target loss function according to the second predicted image, the target image, and the second discrimination result, and obtain a trained image generation model.

[0217] As an alternative implementation manner of the embodiment of the present disclosure, the image model determination unit 7023 is specifically configured to:

[0218] Based on a first target loss function, determine a third loss value corresponding to the second generator according to the second predicted image and the target image;

[0219] Based on a second target loss function, determine a fourth loss value corresponding to the second discriminator according to the second discrimination result;

[0220] Determine a fifth loss value corresponding to the second generator based on the third objective loss function and the second predicted image;

[0221] Adjust parameters of the second generator according to the third loss value, the fourth loss value, and the fifth loss value until the second generator converges, and obtain a second target model after training the second generator;

[0222] Adjust parameters of the first target model based on the third loss value, the fourth loss value, and the fifth loss value to obtain an adjusted first target model;

[0223] Construct the image generation model according to the adjusted first target model and the second target model.

[0224] The image generation model training apparatus provided by the embodiments of the present disclosure can execute the image generation model training method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method. To avoid repetition, details are not described herein again.

[0225] Figure 8 FIG. is a schematic structural diagram of an image generation apparatus provided by an embodiment of the present disclosure; the apparatus is configured in a computer device and can implement the image generation method described in any embodiment of the present application. The apparatus specifically includes the following:

[0226] An image acquisition module 801, configured to acquire a to-be-predicted face image, a face key point image corresponding to the to-be-predicted face image, and target face key point data, where the target face key point data is predicted based on the to-be-predicted face image and corresponding voice information;

[0227] A predicted image generation module 802, configured to input the to-be-predicted face image, the face key point image corresponding to the to-be-predicted face image, and the target face key point data into an image generation model to obtain a corresponding target predicted image;

[0228] Wherein, the image generation model is trained based on the image generation model training method described in any embodiment.

[0229] The image generation apparatus provided by the embodiments of the present disclosure can execute the image generation method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method. To avoid repetition, details are not described herein again.

[0230] An embodiment of the present disclosure provides a computer device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement any one of the image generation model training methods or any one of the image generation methods in the embodiments of the present disclosure.

[0231] Figure 9 FIG. 4 is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. As Figure 9 shown, the computer device includes a processor 910 and a storage device 920; the number of processors 910 in the computer device may be one or more, Figure 9 and one processor 910 is taken as an example here; the processor 910 and the storage device 920 in the computer device may be connected through a bus or other means, Figure 9 and the connection through a bus is taken as an example here.

[0232] The storage device 920, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the image generation model training method in the embodiments of the present disclosure; or program instructions / modules corresponding to the image generation method in the embodiments of the present disclosure. The processor 910 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the storage device 920, that is, implements the image generation model training method or the image generation method provided by the embodiments of the present disclosure.

[0233] The computer device provided in this embodiment can be used to execute the method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0234] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the method executed in any of the above embodiments and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0235] Among them, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0236] For the sake of convenience in explanation, the above description has been made in connection with specific embodiments. However, the discussion in some of the above embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific uses.

Claims

1. A method for training an image generation model, characterized in that, The method includes: Training a first network model based on a first training sample to obtain a trained first target model. The first training sample includes a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image corresponding to the target image, and face key point data of the target image. The target image is the next-frame face image corresponding to the first face image. The face key point data of the target image is the face key point data extracted from the target image. The target face key point image corresponding to the target image is a face key point image obtained by imageizing the face key point data of the target image. The face key point image is obtained by the following method: Plotting the spatial coordinate information of the face key points on a Mask map of the same size. The Mask map is a black-and-white map of 0 and 1. The face key points are outlined according to a predetermined plotting strategy to obtain a face key point image after imageizing the face key point data. Wherein, the plotting strategy includes: Using a fully enclosed polygon or a non-fully enclosed polygon to plot different regions in the face attributes, and using different color lines to plot different regions; Training a second network model and the first target model based on a second training sample to obtain a trained image generation model. The second training sample includes a second face image generated by the first target model and the target image; The first network model includes a first generative adversarial network model. The second network model includes a second generative adversarial network model. The second generative adversarial network model includes a second generator and a second discriminator. The structure of the second generative adversarial network model is different from the structure of the first generative adversarial network model; Correspondingly, the training of the second network model and the first target model based on the second training sample to obtain a trained image generation model includes: Inputting the second face image generated by the first target model in the second training sample into the second generator to obtain a second predicted image; Inputting the second predicted image and the target image into the second discriminator to obtain a second discrimination result; Based on a target loss function, training the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result to obtain a trained image generation model.

2. The method according to claim 1, characterized in that, The first network model includes a first generative adversarial network model. The first generative adversarial network model includes a first generator and a first discriminator; Correspondingly, the training of the first network model based on the first training sample to obtain a trained first target model includes: Inputting the first face image, the face key point image, and the face key point data of the target image in the first training sample into the first generator to obtain a first predicted image; Inputting the first face image and the first predicted image into the first discriminator to obtain a first discrimination result; Based on a preset loss function, train the first generator according to the first predicted image, the target image, and the first discrimination result to obtain a first target model after training the first generator.

3. The method according to claim 2, wherein The training of the first generator according to the first predicted image, the target image, and the first discrimination result based on a preset loss function to obtain a first target model after training the first generator includes: Based on a first preset loss function, determine a first loss value corresponding to the first generator according to the first predicted image and the target image; Based on a second preset loss function, determine a second loss value corresponding to the first discriminator according to the first discrimination result; According to the first loss value and the second loss value, adjust the parameters of the first generator until the first generator converges to obtain a first target model after training the first generator.

4. The method according to claim 1, wherein The training of the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result based on a target loss function to obtain a trained image generation model includes: Based on a first target loss function, determine a third loss value corresponding to the second generator according to the second predicted image and the target image; Based on a second target loss function, determine a fourth loss value corresponding to the second discriminator according to the second discrimination result; Based on a third target loss function and the second predicted image, determine a fifth loss value corresponding to the second generator; According to the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the second generator until the second generator converges to obtain a second target model after training the second generator; Based on the third loss value, the fourth loss value, and the fifth loss value, adjust the parameters of the first target model to obtain an adjusted first target model; Construct the image generation model according to the adjusted first target model and the second target model.

5. An image generation method, characterized in that, The method includes: Obtain a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data, where the target face key point data is predicted based on the face image to be predicted and the corresponding voice information; Input the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into the image generation model to obtain a corresponding target predicted image; Wherein, the image generation model is trained based on the method according to any one of claims 1 to 4.

6. An image generation model training device, characterized in that, The device includes: The first target model determination module is configured to train a first network model based on a first training sample to obtain a trained first target model. The first training sample includes a first face image, a face key point image corresponding to the first face image, a target image, a target face key point image corresponding to the target image, and face key point data of the target image. The target image is the next-frame face image corresponding to the first face image. The face key point data of the target image is the face key point data extracted from the target image. The target face key point image corresponding to the target image is a face key point image obtained by visualizing the face key point data of the target image. The face key point image is obtained by the following method: Plotting the spatial coordinate information of the face key points on a Mask graph of the same size; the Mask graph is a black-and-white graph of 0 and 1; the face key points are outlined according to a predetermined plotting strategy to obtain a face key point image after visualizing the face key point data; wherein, the plotting strategy includes: Using a full-enclosing polygon or a non-full-enclosing polygon to plot different regions in the face attributes, and using different color lines to plot different regions; The image generation model determination module is configured to train a second network model and the first target model based on a second training sample to obtain a trained image generation model. The second training sample includes a second face image generated by the first target model and the target image; The first network model includes a first generative adversarial network model. The second network model includes a second generative adversarial network model. The second generative adversarial network model includes a second generator and a second discriminator. The structure of the second generative adversarial network model is different from the structure of the first generative adversarial network model; Correspondingly, the training of the second network model and the first target model based on the second training sample to obtain a trained image generation model includes: Inputting the second face image generated by the first target model in the second training sample into the second generator to obtain a second predicted image; Inputting the second predicted image and the target image into the second discriminator to obtain a second discrimination result; Based on a target loss function, training the second generator and the first target model according to the second predicted image, the target image, and the second discrimination result to obtain a trained image generation model.

7. An image generation device, characterized in that, The device includes: An image acquisition module, configured to acquire a face image to be predicted, a face key point image corresponding to the face image to be predicted, and target face key point data, where the target face key point data is predicted based on the face image to be predicted and the corresponding voice information; A predicted image generation module, configured to input the face image to be predicted, the face key point image corresponding to the face image to be predicted, and the target face key point data into an image generation model to obtain a corresponding target predicted image; Wherein, the image generation model is trained based on the method according to any one of claims 1 to 4.

8. A computer device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Model training method and device, electronic equipment and storage medium

    CN111783948A

  • Face conversion model training method and face image conversion method

    CN113689527A