Image generation method, image generation device and related equipment

By training and parameter adjustment of the StyleGAN model, the target face style image is generated using the target dataset and random noise, the problem of difficulty in obtaining training samples of the style transfer model is solved, and low-cost and efficient sample generation is achieved.

CN115187450BActive Publication Date: 2025-07-22BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210735272.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-07-22
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

In the prior art, the acquisition cost and difficulty of training samples of style transfer models is relatively high, and it is difficult to directly obtain a sufficient number of training image groups from the network.

Method used

The pre-trained StyleGAN model is trained and parameterized using the target dataset. By inputting random noise to the StyleGAN model and the stylized model, images with the target face style are generated, reducing the difficulty of obtaining training samples.

Benefits of technology

By generating a large number of stable sample data, the cost and difficulty of style transfer model training is reduced, and the flexibility and convenience of image generation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187450B_ABST
    Figure CN115187450B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide an image generation method, an image generation device, and related equipment. Among them, the method includes: using a target data set with a target face style to train and adjust the model parameters of a pre-trained StyleGAN model to obtain a stylized model, where the stylized model is used to generate face images with the target face style; inputting random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains a first face image based on the random noise input, and the stylized model obtains a second face image based on the random noise input and the first face image. The second face image is a face image with the target face style obtained after style transfer of the first face image. The method provided by the embodiments of the present invention can reduce the cost and difficulty of training samples used to obtain the style transfer model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to an image generation method, an image generation device and related equipment. Background Art

[0002] Face stylization refers to converting real face images in videos or pictures into face images of different styles. In the process of face stylization, it is usually necessary to input real face images into a style transfer model for style transfer, so as to obtain images of the face in different styles.

[0003] In the process of training the style transfer model, at least several thousand or tens of thousands of image groups are required for training samples, and each image group needs to include real face images and their corresponding stylized face images. Since the number of image groups required for training the style transfer model is large and it is difficult to directly obtain on the network, it usually needs to be manually drawn. Therefore, in the prior art, there are problems of high cost and difficulty in obtaining the training samples used for the style transfer model. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide an image generation method, an image generation device and related equipment to reduce the cost and difficulty of obtaining the training samples used for the style transfer model. The specific technical solutions are as follows:

[0005] In the first aspect implemented by the present invention, first, an image generation method is provided, including:

[0006] Using a target data set with a target face style to train and adjust the model parameters of a pre-trained StyleGAN model to obtain a stylized model, where the stylized model is used to generate face images with the target face style;

[0007] Inputting random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains a first face image based on the random noise input, and the stylized model obtains a second face image based on the random noise input and the first face image. The second face image is a face image with the target face style obtained after style transfer of the first face image.

[0008] Optionally, the StyleGAN model obtaining the first face image based on the random noise input includes:

[0009] The first mapping layer of the StyleGAN model extracts features based on the random noise input to obtain a first hidden feature; the first hidden feature is used to characterize the attributes of the face region in the first face image;

[0010] Perform weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature; the first average feature is determined based on the results output by the first mapping layer during each training of the StyleGAN model;

[0011] The first synthesis module of the StyleGAN model performs image generation based on the input of the first target feature to obtain the first face image.

[0012] Optionally, the performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature includes:

[0013] Perform weighted processing on the first hidden feature and a preset first average feature to obtain a first intermediate feature;

[0014] Adjust the parameters corresponding to the first intermediate feature to obtain the first target feature.

[0015] Optionally, the style model obtaining the second face image based on the random noise input and the first face image includes:

[0016] The second mapping layer of the style model performs feature extraction based on the random noise input to obtain a second hidden feature, and the second hidden feature is used to characterize the attributes of the face region in the second face image;

[0017] Perform weighted processing on the second hidden feature and a preset second average feature to obtain a second target feature; the second average feature is determined based on the results output by the second mapping layer during each training of the style model;

[0018] The second synthesis module of the style model performs image generation based on the input of the second target feature and the input of the first face image to obtain the second face image.

[0019] Optionally, the inputting the random noise into the StyleGAN model and the style model respectively to obtain the first face image and the second face image includes:

[0020] Mix the StyleGAN model and the style model to obtain a mixed model; wherein, the mixed model includes the StyleGAN model and the style model, and the output end of the first synthesis module of the StyleGAN model is connected to the input end of the second synthesis module of the style model;

[0021] Input the random noise into the StyleGAN model and the mixed model for image generation to obtain the first face image and the second face image.

[0022] Optionally, after training the pre-trained StyleGAN model with the target dataset having the target face style and adjusting the model parameters to obtain a stylized model, the method further includes:

[0023] Obtain a third hidden feature of the sample real image, where the third hidden feature is used to characterize the attributes of the face region in the sample real image;

[0024] Perform weighted processing on the third hidden feature and the preset second average feature to obtain a third target feature;

[0025] The first synthesis module of the StyleGAN model generates an image based on the input of the third hidden feature to obtain a first real image;

[0026] The second synthesis module of the stylized model generates an image based on the input of the third target feature and the input of the first real image to obtain a second real image; the second real image is a face image with the target face style after style transfer of the first real image.

[0027] Optionally, after respectively inputting the random noise into the StyleGAN model and the stylized model to obtain a first face image and a second face image, the method further includes:

[0028] Determine the attributes of the background region in the sample real image and determine the attributes of the background region in the second real image;

[0029] Use the attributes of the background region in the sample real image to replace the attributes of the background region in the second real image.

[0030] In a second aspect of the implementation of the present invention, an image generation device is provided, including:

[0031] A processing module, configured to train the pre-trained StyleGAN model with the target dataset having the target face style and adjust the model parameters to obtain a stylized model, where the stylized model is used to generate a face image with the target face style;

[0032] An input module, configured to respectively input the random noise into the StyleGAN model and the stylized model, so that the StyleGAN model obtains a first face image based on the input of the random noise, and the stylized model obtains a second face image based on the input of the random noise and the first face image, and the second face image is a face image with the target face style after style transfer of the first face image.

[0033] In the third aspect of the implementation of the present invention, an electronic device is provided, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0034] The memory is used to store programs;

[0035] The processor is used to implement the method steps described in the first aspect when executing the programs stored on the memory.

[0036] In the third aspect of the implementation of the present invention, a readable storage medium is provided, on which a program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.

[0037] In the embodiments of the present invention, random noise is respectively input into the StyleGAN model and the stylization model to obtain a first face image and a second face image. On the one hand, the StyleGAN model obtains the first face image based on the input of the random noise, and the first face image is a real face image. On the other hand, the stylization model obtains the second face image based on the input of the random noise and the first face image, and the second face image is a stylized image corresponding to the first face image. In this way, through the method provided by the embodiments of the present invention, a large amount of stable sample data for training the style transfer model can be generated, reducing the cost and difficulty of obtaining the training samples used for the style transfer model. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art.

[0039] Figure 1 It is a schematic flowchart of an image generation method in an embodiment of the present invention;

[0040] Figure 2 It is a schematic structural diagram of a StyleGAN2 model in an embodiment of the present invention;

[0041] Figure 3 It is a schematic structural diagram of a stylization model in an embodiment of the present invention;

[0042] Figure 4 It is a schematic structural diagram of a hybrid model in an embodiment of the present invention;

[0043] Figure 5 It is a schematic flowchart of a process for obtaining a hybrid model in an embodiment of the present invention;

[0044] Figure 6It is a schematic flowchart of data processing in an embodiment of the present invention;

[0045] Figure 7 It is a schematic structural diagram of an image generation device in an embodiment of the present invention;

[0046] Figure 8 It is a structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.

[0048] As Figure 1 shown, the embodiments of the present invention provide an image generation method, including the following steps:

[0049] Step 101: Use a target data set with a target face style to train a pre-trained Style Generative Adversarial Network (StyleGAN) model and adjust the model parameters to obtain a stylized model, where the stylized model is used to generate a face image with the target face style;

[0050] Step 102: Input random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains the first face image based on the input of the random noise, and the stylized model obtains the second face image based on the input of the random noise and the first face image. The second face image is a face image with the target face style obtained after style transfer of the first face image.

[0051] It should be understood that the training data set used for pre-training the StyleGAN model can be any data set. In some embodiments, the used training data set is an open-source large-sample data set. For example, in some embodiments, the used training data set can be the Flickr-Faces-High-Quality (FFHQ) data set of high-quality faces in the Yahoo Web Album. In other embodiments, the used training database can be the Public Figures Face Database (PubFig).

[0052] It should be understood that the specific structure of the StyleGAN model is not limited herein. In specific implementation, the StyleGAN model may be a StyleGAN basic model or a model optimized based on the StyleGAN basic model. For example, in some embodiments, the StyleGAN model may also be a StyleGAN2 model.

[0053] For the convenience of description, in the subsequent embodiments, the StyleGAN model is the pre-trained StyleGAN model, and the case where the StyleGAN model is the StyleGAN2 model will be taken as an example for illustration.

[0054] It should be understood that the face style can be understood as the display style of various parts of the face in the image, including different face attributes such as face shape, facial expression, face orientation, hairstyle, skin color of the face, face lighting, shape of hair strands, hair color, and wrinkles.

[0055] That the target data set has the target face style can be understood as that all sample images in the target data set have the target face style. The target data set can be any data set. For example, in some embodiments, the target data set can be a self-built data set, that is, the target data set is established by collecting images with the target face style. In other embodiments, the target data set can also be an open-source data set.

[0056] In specific implementation, since the painting methods of facial features, brushstrokes, color application methods, etc. in the works of the same painter are relatively unified, it can be considered that the works of the same painter usually have the same face style. Therefore, the works of any painter can be collected as the target data set, and the face style of the works of this painter is the target face style of the target data set.

[0057] It should be noted that according to the different target face styles of the used target data set, the target face styles of the second face images generated by the trained stylized model are also different.

[0058] Since the StyleGAN2 model is pre-trained, when training the StyleGAN2 model and adjusting the model parameters, the used target data set can be a large-sample data set or a small-sample data set. For example, in some embodiments, the range of the number of samples included in the target data set is 50 to 2000. In other embodiments, the range of the number of samples included in the target data set is 80 to 300.

[0059] In specific implementation, the difficulty of obtaining a small-sample dataset is usually less than that of obtaining a large-sample dataset, and using a large-sample dataset for model training can usually achieve better results. Therefore, in this embodiment, first, the StyleGAN2 model can be pre-trained using a large-sample dataset to obtain a better training effect. Then, using the target dataset, the StyleGAN2 model is trained and its model parameters are adjusted to obtain a stylized model. Through the above settings, on the one hand, the training effect of the stylized model is improved, and on the other hand, the difficulty of obtaining the target dataset is reduced.

[0060] In some embodiments, since the target dataset can be a small-sample dataset, overfitting may occur during the process of training the StyleGAN2 model and adjusting its model parameters using the target dataset. Therefore, in some embodiments, to improve the training effect, training the StyleGAN2 model can also be understood as training the StyleGAN2 model through the Style Generative Adversarial Network-Adaptive Discriminator Augmentation (StyleGAN-ada) method and the early stop method. Further, in some other embodiments, training the StyleGAN2 model can also be understood as training the StyleGAN2 model through the StyleGAN-ada method, the Freeze the Discriminator (FreezeD) method, and the early stop method.

[0061] Adjusting the model parameters of the StyleGAN2 model can be understood as fine-tuning the StyleGAN2 model. Specifically, after training the StyleGAN2 model to obtain a trained StyleGAN2 model, based on the trained StyleGAN2 model, the parameters of at least one network layer are adjusted to obtain the stylized model.

[0062] In specific implementation, the StyleGAN2 model can be trained multiple times and its model parameters can be adjusted multiple times to obtain different stylized models. Then, the style transfer effects of the obtained multiple stylized models are compared to determine the final stylized model, where the stylized model is used to generate face images with the target face style.

[0063] After inputting the random noise into the StyleGAN2 model, the StyleGAN2 model can generate the first face image based on the input of the random noise. In this embodiment, the first face image can be considered as a face image with a real face style.

[0064] After inputting the random noise into the stylization model, the stylization model can generate the second face image based on the input of the random noise. In this embodiment, the second face image is a face image with the target face style obtained after style transfer of the first face image. Specifically, the face content of the second face image is generated based on the face content of the first face image, but the second face image has the target face style.

[0065] It should be understood that the first face image and the second face image can be considered as an image group, which includes a real face image and its corresponding stylized face image. The first face image can be understood as the generated real face image, and the second face image can be understood as the stylized face image corresponding to the first face image.

[0066] It should be noted that after inputting the random noise into the StyleGAN2 model and the stylization model respectively, the first face image and the second face image can be obtained. Therefore, the random noise input into the StyleGAN2 model and the stylization model is the same.

[0067] In the embodiment of the present invention, by inputting the random noise into the StyleGAN model and the stylization model respectively, the first face image and the second face image are obtained. On the one hand, the StyleGAN model obtains the first face image based on the input of the random noise, and the first face image is a real face image. On the other hand, the stylization model obtains the second face image based on the input of the random noise and the first face image, and the second face image is the stylized image corresponding to the first face image. In this way, through the method provided by the embodiment of the present invention, a large amount of stable sample data for training the style transfer model can be generated, reducing the cost and difficulty of obtaining the training samples used for the style transfer model.

[0068] It should be understood that the specific method for the StyleGAN2 model to obtain the first face image based on the input of the random noise is not limited herein. Optionally, in some embodiments, the StyleGAN2 model obtaining the first face image based on the input of the random noise specifically includes the following steps:

[0069] The first mapping layer of the StyleGAN2 model extracts features based on the random noise input to obtain a first hidden feature; the first hidden feature is used to characterize the attributes of the face region in the first face image;

[0070] The first hidden feature is weighted with a preset first average feature to obtain a first target feature; the first average feature is determined based on the output result of the first mapping layer during each training of the StyleGAN2 model;

[0071] The first synthesis module of the StyleGAN2 model generates an image based on the input of the first target feature to obtain the first face image.

[0072] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a StyleGAN2 model in an embodiment of the present invention. As Figure 2 shown, the StyleGAN2 model includes a first mapping layer and a first synthesis module. In some embodiments, the first mapping layer may also be referred to as a first mapping network (Mapping Network), and the first synthesis module may also be referred to as a first synthesis network (Synthesis Network).

[0073] Specifically, the first mapping layer is used to extract features based on the random noise input to obtain a first hidden feature, where the first hidden feature can also be understood as a latent vector or latent factor (Latent Code) used to characterize the attributes of the face region in the first face image. In specific implementation, different attributes in the face region of the face image are usually correlated with each other and have a high coupling degree, and the Latent Code is a feature obtained after decoupling the above different attributes.

[0074] As can be seen from the above, in this embodiment, by extracting features based on the random noise input through the first mapping layer to obtain a first hidden feature, the representation effect of the attributes of the face region in the first face image can be improved, and at the same time, the operation convenience of adjusting different attributes in the face region of the first face image is also improved.

[0075] After the first hidden feature is obtained by the first mapping layer, the first hidden feature needs to be weighted with a preset first average feature to obtain a first target feature. Among them, the first average feature is determined based on the output result of the first mapping layer during each pre-training of the StyleGAN2 model. In this embodiment, during the process of weighting the first hidden feature with the preset first average feature, the weights of the first hidden feature and the first average feature are not limited herein.

[0076] It should be understood that the specific manner in which the first average feature is determined based on the results output by the first mapping layer during each training of the StyleGAN2 model is not limited herein. For example, in some embodiments, during the pre-training of the StyleGAN2 model, the results output by the first mapping layer during each training process are recorded, and the mean value is determined based on the results output by the first mapping layer during each training process. This mean value is the first average feature. Further, in some other embodiments, the first average feature may also be determined based on the results output by the first mapping layer during other training processes after abnormal data is removed.

[0077] In this embodiment, by performing weighted processing on the first hidden feature and a preset first average feature, and then adjusting the weights of the first hidden feature and the first average feature, the value range of the first target feature can be made more reasonable, thereby improving the generation effect of the first face image and at the same time improving the generation stability of the first face image.

[0078] Optionally, in some embodiments, the step of performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature includes:

[0079] Performing weighted processing on the first hidden feature and a preset first average feature to obtain a first intermediate feature;

[0080] Adjusting the parameters corresponding to the first intermediate feature to obtain the first target feature.

[0081] Specifically, both the first hidden feature and the first average feature can be represented as high-dimensional vectors. Therefore, after performing weighted processing on the first hidden feature and the first average feature, the obtained first intermediate feature can also be represented as a high-dimensional vector.

[0082] It should be understood that each element in the high-dimensional vector corresponding to the first intermediate feature can be understood as a parameter of the first intermediate feature, or as an attribute characterizing the face region in the first face image. Therefore, adjusting the parameters corresponding to the first intermediate feature can also be understood as editing the elements in the high-dimensional vector, thereby achieving the effect of editing the attributes of the face region in the first face image.

[0083] For ease of understanding, an example will be given below. For example, in one embodiment, the attributes of the face region in the first face image include Asian, open mouth, closed eyes, turned head, tilted head up, and tilted head down, and the above attributes can all be edited.

[0084] As can be seen from the above, in this embodiment, the first intermediate feature is a six-dimensional vector, and each element in the six-dimensional vector corresponds to one of the above attributes one by one. After determining the correspondence between the attribute and the element, the attribute can be edited by adjusting the corresponding element.

[0085] For example, in one case, it is desired that the face in the first face image is in an open-mouth pose. At this time, the second element in the six-dimensional vector can be adjusted accordingly so that the attribute represented by the six-dimensional vector is the open-mouth pose. For another example, in another case, it is desired that the face in the first face image is an Asian and in a closed-eye pose. At this time, both the first element and the third element in the six-dimensional vector can be adjusted accordingly so that the attribute represented by the six-dimensional vector is an Asian and in a closed-eye pose.

[0086] In this embodiment, after performing weighted processing on the first hidden feature and a preset first average feature to obtain the first intermediate feature, the parameters corresponding to the first intermediate feature can be adjusted to obtain the first target feature. By adjusting the parameters corresponding to the first intermediate feature, different first target features can be obtained. Depending on the differences in the first target features, the attributes of the face region in the first face image are also different, thereby causing the face poses in the first sample image to be different. Therefore, through the above settings, the first face images with different poses can be obtained, further improving the richness of the data.

[0087] It should be understood that the first synthesis module of the StyleGAN2 model is used to generate an image based on the input of the first target feature to obtain the first face image, and the attributes of the face region in the first face image are determined based on the first target feature.

[0088] Among them, the structure of the first synthesis module is not limited herein. In some embodiments, the first synthesis module generally includes at least two first image processing layers connected in sequence; and the resolution of the result output by the first image processing layer gradually increases from the upper layer to the lower layer. The number and specific structure of the first image processing layer are not limited herein.

[0089] Therefore, the first synthesis module of the StyleGAN2 model generates an image based on the first target feature input, and obtaining the first face image can also be understood as inputting the first target feature into each of the first image processing layers of the first synthesis module. In other words, each of the first image processing layers can perform image processing based on the input of the previous first image processing layer and the first target feature input, and input the processed result into the next first image processing layer. When the first image processing layer is the uppermost first image processing layer, this first image processing layer generates an image based on the first target feature input. When the first image processing layer is the lowermost first image processing layer, the result output by this first image processing layer is the first face image.

[0090] It should be understood that the specific method by which the stylization model obtains the second face image based on the random noise input and the first face image is not limited herein. Optionally, in some embodiments, the stylization model obtaining the second face image based on the random noise input and the first face image specifically includes the following steps:

[0091] The second mapping layer of the stylization model extracts features based on the random noise input to obtain a second hidden feature, and the second hidden feature is used to characterize the attributes of the face region in the second face image;

[0092] The second hidden feature is weighted with a preset second average feature to obtain a second target feature; the second average feature is determined based on the result output by the second mapping layer each time the stylization model is trained;

[0093] The second synthesis module of the stylization model performs image generation based on the second target feature input and the first face image input to obtain the second face image.

[0094] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a stylization model in an embodiment of the present invention. As Figure 3 shown, the stylization model includes a second mapping layer and a second synthesis module. In some embodiments, the second mapping layer can also be referred to as the second Mapping Network, and the second synthesis module can also be referred to as the second Synthesis Network.

[0095] Specifically, the second mapping layer is used to extract features based on the random noise input to obtain a second hidden feature, where the second hidden feature can also be understood as a Latent Code used to characterize the attributes of the face region in the second face image.

[0096] As can be seen from the above, in this embodiment, by performing feature extraction on the random noise input through the second mapping layer to obtain the second hidden feature, the representation effect of the attributes of the face region in the second face image can be improved, and at the same time, the operation convenience of adjusting different attributes of the face region in the second face image is also improved.

[0097] After the second mapping layer obtains the second hidden feature, it is necessary to perform weighted processing on the second hidden feature and a preset second average feature to obtain a second target feature. Among them, the second average feature is determined based on the output result of the second mapping layer during each training of the stylization model. In this embodiment, during the process of performing weighted processing on the second hidden feature and the second average feature, the weights of the second hidden feature and the second average feature are not limited herein.

[0098] It should be understood that the specific manner in which the second average feature is determined based on the output result of the second mapping layer during each training of the stylization model is not limited herein. For example, in some embodiments, during the pre-training of the stylization model, the output results of the second mapping layer during each training process are recorded, and the mean value is determined based on the output results of the second mapping layer during each training process, and this mean value is the second average feature. Further, in some other embodiments, the second average feature can also be determined based on the output results of the second mapping layer during other training processes after abnormal data is removed.

[0099] In this embodiment, by performing weighted processing on the second hidden feature and a preset second average feature and then adjusting the weights of the second hidden feature and the second average feature, the value range of the second target feature can be made more reasonable, thereby improving the generation effect of the second face image and at the same time improving the generation stability of the second face image.

[0100] At the same time, the stylization model is used to perform style transfer on the first face image. Therefore, by adjusting the weights of the second hidden feature and the second average feature, the intensity of the style transfer performed by the stylization model on the first face image can be controlled, or the style intensity. Therefore, in the case where the weights of the second hidden feature and the second average feature are different, according to different style intensities, the target face style can also include multiple target face sub-styles, thereby obtaining a more satisfactory style transfer effect.

[0101] It should be understood that the second synthesis module of the stylization model is used to perform image generation based on the input of the second target feature to obtain the second face image, and the attributes of the face region in the second face image are determined based on the second target feature.

[0102] Among them, the structure of the second synthesis module is not limited herein. In some embodiments, the second synthesis module generally includes at least two sequentially connected second image processing layers; and the resolution of the result output by the second image processing layer gradually increases from the upper layer to the lower layer. Among them, the number and specific structure of the second image processing layer are not limited herein.

[0103] Therefore, the second synthesis module of the stylization model generates an image based on the second target feature input, and obtaining the second face image can also be understood as inputting the second target feature into each of the second image processing layers of the second synthesis module. In other words, each of the second image processing layers can perform image processing based on the input of the previous second image processing layer and the second target feature input, and input the processed result into the next second image processing layer. When the second image processing layer is the topmost second image processing layer, this second image processing layer generates an image based on the second target feature input. When the second image processing layer is the bottommost second image processing layer, the result output by this second image processing layer is the second face image.

[0104] Optionally, in some embodiments, the step of respectively inputting random noise into the StyleGAN model and the stylization model to obtain the first face image and the second face image specifically includes the following steps:

[0105] Mix the StyleGAN model and the stylization model to obtain a hybrid model; wherein, the hybrid model includes the StyleGAN model and the stylization model, and the output end of the first synthesis module of the StyleGAN model is connected to the input end of the second synthesis module of the stylization model;

[0106] Input the random noise into the StyleGAN model and the hybrid model for image generation to obtain the first face image and the second face image.

[0107] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a hybrid model in an embodiment of the present invention. As Figure 4 shown, the hybrid model includes a StyleGAN2 model and the stylization model, and the output end of the first synthesis module of the StyleGAN2 model is connected to the input end of the second synthesis module of the stylization model.

[0108] The connection between the output end of the first synthesis module of the StyleGAN2 model and the input end of the second synthesis module of the stylization model can be understood as that the result output by the first synthesis module of the StyleGAN2 model will be input into the second synthesis module of the stylization model.

[0109] As Figure 4 shown, in one case, the result output by the first synthesis module of the StyleGAN2 model is only input to the second synthesis module of the stylization model. In another case, the result output by the first synthesis module of the StyleGAN2 model is input to the second synthesis module of the stylization model and at the same time output as the first face image.

[0110] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a process for obtaining a hybrid model provided in this embodiment. As Figure 5 shown, first, sample data with the target face style is collected to form the target database, and then the StyleGAN2 model is trained and fine-tuned on the target dataset through the StyleGAN-ada method, the FreezeD method, and the early stop method to obtain the stylization model. The stylization model and the StyleGAN2 model are mixed to obtain the hybrid model.

[0111] In this embodiment, the StyleGAN model and the stylization model are mixed to obtain the hybrid model. The random noise is input to the StyleGAN model and the hybrid model for image generation to obtain the first face image and the second face image. In this way, the first face image and the second face image can still be obtained for training the style transfer model. At the same time, in the case where only the second face image needs to be obtained, the random noise can be input only to the hybrid model for processing to obtain the second face image.

[0112] As can be seen from the above, through the method provided in this embodiment, a StyleGAN2 model as Figure 2 shown, a stylization model as Figure 3 shown, and a hybrid model as Figure 4 shown can be obtained. In the case of needing to generate different face images, different models can be used, improving the flexibility and convenience of image generation.

[0113] Optionally, in some embodiments, after step 101, the method further includes the following steps:

[0114] Obtain a third hidden feature of the sample real image, where the third hidden feature is used to characterize the attributes of the face region in the sample real image;

[0115] Perform weighted processing on the third hidden feature and the preset second average feature to obtain a third target feature;

[0116] The first synthesis module of the StyleGAN model generates an image based on the third hidden feature input to obtain a first real image;

[0117] The second synthesis module of the stylization model generates an image based on the third target feature input and the first real image input to obtain a second real image; the second real image is a face image with the target face style after the style transfer of the first real image.

[0118] It should be noted that the first face image can be considered as a face image with a real face style. And the sample real image can be understood as a real face image obtained on the network or through other means. The difference between the first face image and the sample real image can be understood as that the person corresponding to the face area in the first face image is false, while the person corresponding to the face area in the sample real image is real.

[0119] It should be noted that the weighted processing of the third hidden feature and the preset second average feature to obtain the third target feature can be understood as using the third hidden feature to replace the second hidden feature. Therefore, in this embodiment, the second hidden feature is not generated by the random noise, but is determined according to the sample real image.

[0120] Since the purpose of the stylization model is to transfer the target face style to the first face image to obtain the second face image. And the second average feature is used to control the intensity of face stylization. Therefore, by replacing the second hidden feature with the third hidden feature, the face attributes of the sample real image can be input. At the same time, by weighting the third hidden feature and the preset second average feature, the stylization intensity of the sample real image can still be adjusted.

[0121] It should be noted that the first synthesis module of the StyleGAN2 model generates an image based on the third hidden feature input to obtain a first real image can be understood as using the third hidden feature to replace the first target feature.

[0122] Since the purpose of the StyleGAN2 model is to generate a face image with a real face style, therefore, using the third hidden feature to replace the first target feature can improve the similarity between the generated first face image and the sample real image, and improve the authenticity of the first face image. Of course, in some embodiments, the sample real image can also be directly used to replace the first face image, that is, the sample real image and the second real image are used as an image group.

[0123] In this embodiment, the real image of the sample is collected, and at the same time, the third hidden feature of the real image of the sample is obtained. By using the third hidden feature to replace the second hidden feature and the first average feature, the stability of the output first real image and second real image is improved. At the same time, the source variety of the sample data used for training the style transfer model is increased, making the sample data used for training the style transfer model more stable and real, thereby improving the training effect of the style transfer model.

[0124] Optionally, in some embodiments, after step 102, the method further includes the following steps:

[0125] Determine the attribute of the background area in the real image of the sample, and determine the attribute of the background area in the second real image;

[0126] Use the attribute of the background area in the real image of the sample to replace the attribute of the background area in the second real image.

[0127] It should be understood that the specific method for determining the attribute of the background area in the real image of the sample is not limited herein. Similarly, the specific method for determining the attribute of the background area in the second real image is not limited herein.

[0128] For example, in some embodiments, determining the attribute of the background area in the real image of the sample can be understood as performing image segmentation on the real image of the sample to determine the background area of the real image of the sample. Determining the attribute of the background area in the second real image can be understood as performing image segmentation on the second real image to determine the background area of the second real image.

[0129] In this embodiment, determining the attribute of the background area in the real image of the sample, and determining the attribute of the background area in the second real image are understood as determining the background area in the real image of the sample, and determining the background area in the second real image. Therefore, using the attribute of the background area in the real image of the sample to replace the attribute of the background area in the second real image can be understood as using the background in the real image of the sample to replace the background in the second real image.

[0130] In other embodiments, determining the attribute of the background area in the real image of the sample can be understood as performing image segmentation on the real image of the sample to determine the attribute parameters of the background area of the real image of the sample. Determining the attribute of the background area in the second real image can be understood as performing image segmentation on the second real image to determine the attribute parameters of the background area of the second real image.

[0131] In this embodiment, determining the attributes of the background region in the sample real image and determining the attributes of the background region in the second real image is understood as determining the attribute parameters of the background region in the sample real image and determining the attribute parameters of the background region in the second real image. Therefore, replacing the attributes of the background region in the second real image with the attributes of the background region in the sample real image can be understood as replacing the background parameters in the second real image with the background parameters in the sample real image.

[0132] In the process of performing style transfer on the sample real image, the background region in the sample real image may also be subject to style transfer, making the background region in the second real image different from the background region in the sample real image.

[0133] In this embodiment, it is necessary to determine the attributes of the background region in the sample real image and determine the attributes of the background region in the second real image; replace the attributes of the background region in the second real image with the attributes of the background region in the sample real image. Through the above settings, the background region in the second real image is made the same as the background region in the sample real image, ensuring that the backgrounds of the images in the image group are the same.

[0134] It should be noted that, as Figure 6 shown, in some embodiments, after obtaining the second real image, the sample real image and the second real image can be further aligned by a piecewise method to further improve the effect of the second real image.

[0135] It should be noted that, in some embodiments, after obtaining the second real image, the clothing in the second real image can be replaced with the clothing in the sample real image to further improve the matching degree of the clothing in the second real image and the sample real image and the richness of the clothing in the second real image.

[0136] It should be noted that, in some embodiments, after obtaining the first face image, the second face image, the first real image, and / or the second real image, accessory data can be added to the first face image, the second face image, the first real image, and / or the second real image to further enrich the types and quantities of the sample data for training the style transfer model. Among them, the accessory data can be glasses data, jewelry data, mask data, etc.

[0137] It should be noted that in some embodiments, after obtaining the first face image, the second face image, the first real image, and / or the second real image, data cleaning may also be performed on the first face image, the second face image, the first real image, and / or the second real image to eliminate data with poor effects, thereby improving the quality of the sample data ultimately used for training the style transfer model.

[0138] Please refer to Figure 7 , Figure 7 which is a structural diagram of an image generation device 700 provided by an embodiment of the present invention.

[0139] As Figure 7 shown, this embodiment provides an image generation device 700, including:

[0140] A processing module 701, configured to train a pre-trained StyleGAN model and adjust model parameters by using a target data set with a target face style to obtain a stylized model, where the stylized model is used to generate a face image with the target face style;

[0141] An input module 702, configured to input random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains the first face image based on the random noise input, and the stylized model obtains the second face image based on the random noise input and the first face image, where the second face image is a face image with the target face style obtained after style transfer of the first face image.

[0142] Optionally, the StyleGAN model obtaining the first face image based on the random noise input includes:

[0143] The first mapping layer of the StyleGAN model extracts features based on the random noise input to obtain a first hidden feature; the first hidden feature is used to characterize the attributes of the face region in the first face image;

[0144] Performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature; the first average feature is determined based on the result output by the first mapping layer during each training of the StyleGAN model;

[0145] The first synthesis module of the StyleGAN model generates an image based on the input of the first target feature to obtain the first face image.

[0146] Optionally, the performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature includes:

[0147] Perform weighted processing on the first hidden feature and a preset first average feature to obtain a first intermediate feature;

[0148] Adjust the parameters corresponding to the first intermediate feature to obtain the first target feature.

[0149] Optionally, the stylization model obtains the second face image based on the random noise input and the first face image, including:

[0150] The second mapping layer of the stylization model performs feature extraction based on the random noise input to obtain a second hidden feature, which is used to characterize the attributes of the face region in the second face image;

[0151] Perform weighted processing on the second hidden feature and a preset second average feature to obtain a second target feature; the second average feature is determined based on the results output by the second mapping layer during each training of the stylization model;

[0152] The second synthesis module of the stylization model performs image generation based on the second target feature input and the first face image input to obtain the second face image.

[0153] Optionally, the input module 702 includes:

[0154] A mixing unit, configured to mix the StyleGAN model and the stylization model to obtain a mixed model; wherein, the mixed model includes the StyleGAN model and the stylization model, and the output end of the first synthesis module of the StyleGAN model is connected to the input end of the second synthesis module of the stylization model;

[0155] An image generation unit, configured to input the random noise into the mixed model for image generation to obtain a first face image and a second face image.

[0156] Optionally, the image generation device 700 further includes:

[0157] An acquisition module, configured to acquire a third hidden feature of a sample real image, where the third hidden feature is used to characterize the attributes of the face region in the sample real image;

[0158] A weighted processing module, configured to perform weighted processing on the third hidden feature and the preset second average feature to obtain a third target feature;

[0159] A first image generation module, configured to perform image generation by the first synthesis module of the StyleGAN model based on the third hidden feature input to obtain a first real image;

[0160] A second image generation module, configured to generate an image based on the third target feature input and the first real image input by the second synthesis module of the stylization model, so as to obtain a second real image; the second real image is a face image with the target face style after the first real image is subjected to style transfer.

[0161] Optionally, the image generation device 700 further includes:

[0162] A determination module, configured to determine the attribute of the background region in the sample real image and determine the attribute of the background region in the second real image;

[0163] A replacement module, configured to replace the attribute of the background region in the second real image with the attribute of the background region in the sample real image.

[0164] The image generation device 700 provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments, and for the sake of brevity, details are not described herein again.

[0165] An embodiment of the present invention further provides an electronic device, as Figure 8 shown, including a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804.

[0166] The memory 803 is used for storing programs;

[0167] The processor 801, when executing the programs stored in the memory 803, implements the following steps:

[0168] Training a pre-trained StyleGAN model and adjusting model parameters by using a target data set with a target face style to obtain a stylization model, where the stylization model is used to generate a face image with the target face style;

[0169] Inputting random noise into the StyleGAN model and the stylization model respectively, so that the StyleGAN model obtains the first face image based on the random noise input, and the stylization model obtains the second face image based on the random noise input and the first face image, where the second face image is a face image with the target face style obtained after the first face image is subjected to style transfer.

[0170] Optionally, the StyleGAN model obtaining the first face image based on the random noise input includes:

[0171] The first mapping layer of the StyleGAN model extracts features based on the random noise input to obtain a first hidden feature; the first hidden feature is used to characterize the attributes of the face region in the first face image;

[0172] The first hidden feature is weighted with a preset first average feature to obtain a first target feature; the first average feature is determined based on the output result of the first mapping layer during each training of the StyleGAN model;

[0173] The first synthesis module of the StyleGAN model generates an image based on the input of the first target feature to obtain the first face image.

[0174] Optionally, the step of weighting the first hidden feature with a preset first average feature to obtain a first target feature includes:

[0175] The first hidden feature is weighted with a preset first average feature to obtain a first intermediate feature;

[0176] The parameters corresponding to the first intermediate feature are adjusted to obtain the first target feature.

[0177] Optionally, the step that the stylization model obtains the second face image based on the random noise input and the first face image includes:

[0178] The second mapping layer of the stylization model extracts features based on the random noise input to obtain a second hidden feature, and the second hidden feature is used to characterize the attributes of the face region in the second face image;

[0179] The second hidden feature is weighted with a preset second average feature to obtain a second target feature; the second average feature is determined based on the output result of the second mapping layer during each training of the stylization model;

[0180] The second synthesis module of the stylization model generates an image based on the input of the second target feature and the input of the first face image to obtain the second face image.

[0181] Optionally, the step of respectively inputting random noise into the StyleGAN model and the stylization model to obtain the first face image and the second face image includes:

[0182] Mix the StyleGAN model and the stylization model to obtain a hybrid model; wherein, the hybrid model includes the StyleGAN model and the stylization model, and the output end of the first synthesis module of the StyleGAN model is connected to the input end of the second synthesis module of the stylization model;

[0183] Input the random noise into the hybrid model for image generation to obtain a first face image and a second face image.

[0184] Optionally, after training the pre-trained StyleGAN model using a target dataset with a target face style and adjusting the model parameters to obtain a stylization model, the method further includes:

[0185] Obtain a third hidden feature of the sample real image, where the third hidden feature is used to characterize the attributes of the face region in the sample real image;

[0186] Perform weighted processing on the third hidden feature and the preset second average feature to obtain a third target feature;

[0187] The first synthesis module of the StyleGAN model performs image generation based on the input of the third hidden feature to obtain a first real image;

[0188] The second synthesis module of the stylization model performs image generation based on the input of the third target feature and the input of the first real image to obtain a second real image; the second real image is a face image with the target face style after style transfer of the first real image.

[0189] Optionally, after respectively inputting random noise into the StyleGAN model and the stylization model to obtain a first face image and a second face image, the method further includes:

[0190] Determine the attributes of the background region in the sample real image and determine the attributes of the background region in the second real image;

[0191] Use the attributes of the background region in the sample real image to replace the attributes of the background region in the second real image.

[0192] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0193] The communication interface is used for communication between the above terminal and other devices.

[0194] The memory can include a Random Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0195] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0196] In another embodiment provided by the present invention, a readable storage medium is also provided. Instructions are stored in the readable storage medium, and when they run on a processor, the processor is caused to execute the image generation method described in any one of the above embodiments.

[0197] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, the computer is caused to execute the image generation method described in any one of the above embodiments.

[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0199] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0200] Each embodiment in this specification is described in a related manner. For the same and similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.

[0201] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. An image generation method, characterized in that, Including: Using a target dataset with a target face style to train and adjust the model parameters of a pre-trained style generation adversarial network StyleGAN model to obtain a stylized model, where the stylized model is used to generate face images with the target face style; Inputting random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains a first face image based on the random noise input, and the stylized model obtains a second face image based on the random noise input and the first face image. The second face image is a face image with the target face style obtained after style transfer of the first face image.

2. The method according to claim 1, wherein The StyleGAN model obtaining the first face image based on the random noise input includes: The first mapping layer of the StyleGAN model extracts features based on the random noise input to obtain a first hidden feature; the first hidden feature is used to characterize the attributes of the face region in the first face image; Performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature; the first average feature is determined based on the output result of the first mapping layer during each training of the StyleGAN model; The first synthesis module of the StyleGAN model generates an image based on the first target feature input to obtain the first face image.

3. The method according to claim 2, characterized in that, The performing weighted processing on the first hidden feature and a preset first average feature to obtain a first target feature includes: Performing weighted processing on the first hidden feature and a preset first average feature to obtain a first intermediate feature; Adjusting the parameters corresponding to the first intermediate feature to obtain the first target feature.

4. The method according to claim 2, wherein The stylized model obtaining the second face image based on the random noise input and the first face image includes: The second mapping layer of the stylized model extracts features based on the random noise input to obtain a second hidden feature, and the second hidden feature is used to characterize the attributes of the face region in the second face image; Performing weighted processing on the second hidden feature and a preset second average feature to obtain a second target feature; the second average feature is determined based on the output result of the second mapping layer during each training of the stylized model; The second synthesis module of the stylized model generates an image based on the second target feature input and the first face image input to obtain the second face image.

5. The method according to claim 4, wherein The inputting random noise into the StyleGAN model and the stylized model respectively to obtain a first face image and a second face image includes: Mixing the StyleGAN model and the stylized model to obtain a hybrid model; where the hybrid model includes the StyleGAN model and the stylized model, and the output end of the first synthesis module of the StyleGAN model is connected to the input end of the second synthesis module of the stylized model; Input the random noise into the StyleGAN model and the hybrid model for image generation to obtain a first face image and a second face image.

6. The method according to claim 4, characterized in that After using a target dataset with a target face style to train a pre-trained StyleGAN model and adjust its model parameters to obtain a stylized model, the method further includes: Obtain a third hidden feature of the sample real image, where the third hidden feature is used to characterize the attributes of the face region in the sample real image; Perform weighted processing on the third hidden feature and the preset second average feature to obtain a third target feature; The first synthesis module of the StyleGAN model generates an image based on the input of the third hidden feature to obtain a first real image; The second synthesis module of the stylized model generates an image based on the input of the third target feature and the input of the first real image to obtain a second real image; the second real image is a face image with the target face style after style transfer of the first real image.

7. The method according to claim 6, wherein After inputting the random noise into the StyleGAN model and the stylized model respectively to obtain a first face image and a second face image, the method further includes: Determine the attributes of the background region in the sample real image and determine the attributes of the background region in the second real image; Use the attributes of the background region in the sample real image to replace the attributes of the background region in the second real image.

8. An image generation device, characterized in that, Includes: A processing module, configured to use a target dataset with a target face style to train a pre-trained StyleGAN model and adjust its model parameters to obtain a stylized model, where the stylized model is used to generate a face image with the target face style; An input module, configured to input random noise into the StyleGAN model and the stylized model respectively, so that the StyleGAN model obtains a first face image based on the input of the random noise, and the stylized model obtains a second face image based on the input of the random noise and the first face image, and the second face image is a face image with the target face style after style transfer of the first face image.

9. An electronic device, characterized in that, Includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store programs; The processor, when executing the programs stored on the memory, implements the steps of the method according to any one of claims 1-7.

10. A readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for establishing training set, electronic equipment and medium

    CN111062426A

  • Image generation method and device, electronic equipment and storage medium

    CN113837934A