Image generation model training, image generation method, apparatus, medium, and device

By generating pseudo-labels for supervised model training, the feature distributions of real and virtual faces are learned, solving the problem of low similarity between virtual characters and real faces, and achieving efficient face-shaping effects.

CN114863214BActive Publication Date: 2025-12-12BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210533820.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-12-12
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively improve the similarity between virtual characters and real faces, especially during the character creation process, where the differences between the distribution of features in real and virtual faces are significant, resulting in low efficiency and similarity in character creation.

Method used

By acquiring sample object images, sample virtual object images, and corresponding face-pinching parameters, pseudo-labels are generated for supervised model training. The model learns the feature distribution of real and virtual faces, and uses pseudo-labels containing virtual face features for model training to improve the model's generalization ability.

Benefits of technology

It improves the efficiency and similarity of face creation, reduces the difference between the distribution of real and virtual facial features, saves labor costs, and improves the efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863214B_ABST
    Figure CN114863214B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image generation model training, an image generation method, an apparatus, a medium and an equipment. The training method comprises: obtaining a sample object image, a sample virtual object image and a sample face pinching parameter corresponding to the sample virtual object image; generating a pseudo label of the sample object image according to the sample object image, the sample virtual object image and the sample face pinching parameter; and performing supervised model training by using the pseudo label to obtain an image generation model. In this way, the image generation model can simultaneously learn the real face feature distribution and the feature distribution specific to the virtual face, improve the model generalization ability, reduce the difference between the real face feature distribution and the virtual face feature distribution, and then quickly render a virtual face similar to the user's face according to the face pinching parameter extracted by the image generation model, with high face pinching efficiency and similarity. In addition, the model training by using the pseudo label can avoid manual marking of the corresponding face pinching parameter, and improve the model training efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and in particular, to an image generation model training method, an image generation method, an image generation device, a medium and an image generation equipment. BACKGROUND

[0002] With the development of mobile terminal and computer technology, more and more role-playing games appear. In order to meet the individual customization needs of different players, when creating a virtual role corresponding to a player, the player is usually added some face kneading functions, so that the player can create a role according to his own preferences, for example, a virtual role similar to his real face. How to improve the similarity between the virtual face and the real face of the player becomes the key to the virtual role face kneading. SUMMARY

[0003] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.

[0004] In a first aspect, the present disclosure provides an image generation model training method, comprising: obtaining a sample object image, a sample virtual object image, and a sample face kneading parameter corresponding to the sample virtual object image; generating a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face kneading parameter; and performing supervised model training by using the pseudo label to obtain an image generation model.

[0005] In a second aspect, the present disclosure provides an image generation method, comprising: obtaining a target object image in response to a user request; determining a target face kneading parameter corresponding to the target object image by an image generation model based on the target object image, wherein the image generation model is obtained by training the image generation model training method provided in the first aspect of the present disclosure; and performing face rendering based on the target face kneading parameter to obtain a target virtual object image corresponding to the target object image.

[0006] In a third aspect, the present disclosure provides an image generation model training device, comprising: a first obtaining module configured to obtain a sample object image, a sample virtual object image, and a sample face kneading parameter corresponding to the sample virtual object image; a generating module configured to generate a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face kneading parameter obtained by the first obtaining module; and a training module configured to perform supervised model training by using the pseudo label generated by the generating module to obtain an image generation model.

[0007] In a fourth aspect, the present disclosure provides an image generation apparatus, comprising: a second acquisition module configured to acquire a target object image in response to a user request; a face pinching parameter extraction module configured to determine a target face pinching parameter corresponding to the target object image based on the target object image acquired by the second acquisition module, by using an image generation model, wherein the image generation model is obtained by training the image generation model training method provided in the first aspect of the present disclosure; and a rendering module configured to perform face rendering based on the target face pinching parameter extracted by the face pinching parameter extraction module, to obtain a target virtual object image corresponding to the target object image.

[0008] In a fifth aspect, the present disclosure provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the method provided in the first aspect or the second aspect of the present disclosure.

[0009] In a sixth aspect, the present disclosure provides an electronic device, comprising: a storage device having stored thereon a computer program; and a processing apparatus configured to execute the computer program in the storage device to implement the steps of the method provided in the first aspect or the second aspect of the present disclosure.

[0010] In the above technical solution, the pseudo label of the sample object image is generated according to the acquired sample object image, sample virtual object image and sample face pinching parameter corresponding to the sample virtual object image, so that the pseudo label can learn the feature distribution specific to the virtual face; then, the pseudo label containing the virtual face feature is used for supervised model training, so that the image generation model can not only learn the real face feature distribution, but also learn the feature distribution specific to the virtual face, thereby improving the generalization ability of the model, reducing the difference between the real face feature distribution and the virtual face feature distribution, and further enabling the virtual face similar to the user face to be quickly rendered according to the face pinching parameter extracted by the image generation model, and the efficiency and similarity of face pinching are high. In addition, the pseudo label containing the virtual face feature is used for model training, which can avoid manual marking of the face pinching parameter corresponding to the sample object image, thereby improving the efficiency of model training and saving labor costs.

[0011] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and the actual sizes and elements are not necessarily drawn to scale. In the drawings:

[0013] Figure 1 is a flowchart of a method for training an image generation model according to an example embodiment.

[0014] Figure 2 is a flowchart of a method for generating a pseudo label of a sample object image according to a sample object image, a sample virtual object image, and a sample face-sculpting parameter corresponding to the sample virtual object image according to an example embodiment.

[0015] Figure 3 is a flowchart of a method for generating an image according to an example embodiment.

[0016] Figure 4 is a block diagram of an apparatus for training an image generation model according to an example embodiment.

[0017] Figure 5 is a block diagram of an apparatus for generating an image according to an example embodiment.

[0018] Figure 6 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and embodiments are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0020] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0021] As used herein, the term "includes" and its variants are to be read to be equivalent to "including but not limited to." The term "based on" is to be read as "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions are given throughout the description.

[0022] It should be noted that the terms "first", "second", and the like used in the present disclosure are merely intended to distinguish different devices, modules, or units, and do not imply the order or interdependence of the functions performed by these devices, modules, or units.

[0023] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.

[0024] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0025] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0026] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware that performs the operation of the technical solution of the present disclosure according to the prompt information.

[0027] As an optional but non-limiting implementation, in response to receiving the active request of the user, the user can be sent prompt information in the form of a pop-up window, for example, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0029] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.

[0030] Figure 1 is a flowchart of an image generation model training method according to an example embodiment. As shown in Figure 1 S101-S103.

[0031] In S101, a sample object image, a sample virtual object image, and a sample face pinching parameter corresponding to the sample virtual object image are obtained.

[0032] In the present disclosure, the sample object image is from an authorized public dataset. The sample object image can be an image containing a human face. For example, the sample object image can be a sample human face image.

[0033] In an embodiment, the sample face sculpting parameters include bone point parameters, wherein the bone point parameters are parameters related to face shape and human facial feature shape.

[0034] In another embodiment, the sample face sculpting parameters include makeup parameters, wherein the makeup parameters include lip color, eye shadow color, hairstyle, etc.

[0035] In yet another embodiment, the sample face sculpting parameters include bone point parameters and makeup parameters. In this way, the image generation model can not only learn parameters related to face shape and human facial feature shape, but also learn makeup parameters such as lip color, eye shadow color, and hairstyle, thereby improving the similarity of subsequent face sculpting.

[0036] In S102, a pseudo label of the sample object image is generated according to the sample object image, the sample virtual object image, and the sample face sculpting parameters corresponding to the sample virtual object image.

[0037] In S103, supervised model training is performed using the pseudo label of the sample object image to obtain the image generation model.

[0038] In the above technical solution, the pseudo label of the sample object image is generated according to the obtained sample object image, sample virtual object image, and sample face sculpting parameters corresponding to the sample virtual object image. In this way, the pseudo label can learn the feature distribution specific to the virtual human face. Then, supervised model training is performed using the pseudo label containing the virtual human face features, which can enable the image generation model to not only learn the real human face feature distribution, but also learn the feature distribution specific to the virtual human face, thereby improving the generalization ability of the model, reducing the difference between the real human face feature distribution and the virtual human face feature distribution, and further enabling the image generation model to quickly render a virtual human face similar to the user's face according to the extracted face sculpting parameters. The efficiency and similarity of face sculpting are both high. In addition, model training is performed using the pseudo label containing the virtual human face features, which can avoid manual labeling of the face sculpting parameters corresponding to the sample object image, thereby improving the efficiency of model training and saving labor costs.

[0039] The specific embodiments of obtaining the sample virtual object image and the sample face sculpting parameters corresponding to the sample virtual object image in S101 above will be described in detail below.

[0040] In an implementation, a virtual face like a real face can be pinched out by adjusting preset face pinching parameters, and then an image containing the pinched virtual face is taken as a sample virtual object image, and the adjusted face pinching parameters corresponding to the pinched virtual face are taken as sample face pinching parameters corresponding to the sample virtual object image.

[0041] In another implementation, sample face pinching parameters can be randomly generated, and then face rendering is performed based on the sample face pinching parameters to obtain a sample virtual object image corresponding to the sample face pinching parameters.

[0042] Specifically, the sample virtual object image corresponding to the sample face pinching parameters can be obtained based on the sample face pinching parameters by the following manner: first, face information reconstruction is performed based on the sample face pinching parameters by using a preset rendering engine to obtain reconstructed face information corresponding to the sample face pinching parameters; and then, the sample virtual object image corresponding to the sample face pinching parameters is generated based on the reconstructed face information corresponding to the sample face pinching parameters. In this way, the sample virtual object image can be quickly generated by using the preset rendering engine, which saves time and effort.

[0043] The reconstructed face information corresponds to a three-dimensional face model, which can be determined based on the sample face pinching parameters. Specifically, the three-dimensional face model can be determined based on parameter values corresponding to the sample face pinching parameters, and then the sample virtual object image corresponding to the sample face pinching parameters can be generated through a conversion relationship between a three-dimensional space and a two-dimensional space.

[0044] For example, in the field of game technology, the preset rendering engine can be a game engine.

[0045] The following describes in detail a specific implementation of S102, that is, generating a pseudo label of a sample object image according to a sample object image, a sample virtual object image, and sample face pinching parameters corresponding to the sample virtual object image. Figure 2 S201 and S202 shown in FIG. 2 can be implemented.

[0046] In S201, a first regressor is subjected to model parameter updating according to the sample virtual object image and the sample face pinching parameters.

[0047] In an implementation, the sample virtual object image can be input into the first regressor to obtain first predicted face pinching parameters, and then the model parameters of the first regressor are updated according to the first predicted face pinching parameters and the sample face pinching parameters corresponding to the sample virtual object image.

[0048] In S202, a pseudo label of the sample object image is generated by the first regressor subjected to model parameter updating according to the sample object image.

[0049] In an implementation, the sample object image can be input into the first regressor obtained after the model parameter update to obtain second predicted face-squeezing parameters, and the second predicted face-squeezing parameters are taken as pseudo-labels of the sample object image.

[0050] The specific implementation of updating the model parameters of the first regressor according to the first predicted face-squeezing parameters and the sample face-squeezing parameters corresponding to the sample virtual object image will be described in detail below. Specifically, the following steps (1) and (2) can be used to achieve this:

[0051] (1) Calculate a first loss function according to the first predicted face-squeezing parameters and the sample face-squeezing parameters corresponding to the sample virtual object image.

[0052] For example, the absolute value of the difference between the first predicted face-squeezing parameters and the sample face-squeezing parameters can be determined as the first loss function.

[0053] (2) Update the model parameters of the first regressor in a stochastic gradient descent manner according to the first loss function.

[0054] The specific implementation of performing supervised model training using the pseudo-labels of the sample object image in S103 to obtain the image generation model will be described in detail below. Specifically, the following steps [1] to [4] can be used to achieve this:

[0055] [1] Update the model parameters of the second regressor according to the sample object image and the pseudo-labels.

[0056] In an implementation, the sample object image can be input into the second regressor to obtain third predicted face-squeezing parameters, and then the model parameters of the second regressor are updated according to the third predicted face-squeezing parameters and the pseudo-labels.

[0057] [2] Update the parameters of the first regressor obtained after the model parameter update using the parameters of the second regressor obtained after the model parameter update.

[0058] In an implementation, the first regressor includes a first feature extraction module and a first fully connected module, and the second regressor includes a second feature extraction module and a second fully connected module, where the second feature extraction module has the same structure as the first feature extraction module. At this time, the parameters of the first feature extraction module in the first regressor obtained after the model parameter update can be updated using the parameters of the second feature extraction module in the second regressor obtained after the model parameter update.

[0059] For example, the first feature extraction module and the second feature extraction module each consist of four residual convolutional networks connected in sequence.

[0060] The first full connection module can be composed of one full connection layer or multiple full connection layers connected in sequence, and the second full connection module can be composed of one full connection layer or multiple full connection layers connected in sequence. The number of full connection layers included in the first full connection module and the second full connection module can be set in combination with different training purposes. For example, the first full connection module and the second full connection module can each be composed of three full connection layers connected in sequence.

[0061] It should be noted that the structures of the first full connection module and the second full connection module can be the same or different, and the embodiments of the present disclosure do not make specific limitations.

[0062] [3] Determine whether a preset training stop condition is met.

[0063] In the present disclosure, the preset training stop condition can be that the number of training times reaches a preset number, or the sum of the loss of the first regressor and the loss of the second regressor is less than a preset loss threshold.

[0064] If the preset training stop condition is not met, a new sample object image, a new sample virtual face image, and sample face parameter corresponding to the new sample virtual face image are obtained, and a pseudo label of the new sample object image is generated based on the new sample object image, the new sample virtual face image, and the sample face parameter corresponding to the new sample virtual face image. Then, supervised model training is performed using the pseudo label of the new sample object image, that is, the above S101 is returned, until the preset training stop condition is met. Then, the following step [4] is performed.

[0065] [4] The second regressor is used as an image generation model.

[0066] In the above embodiments, the first regressor for extracting virtual face features is updated using the sample virtual face image and the sample face parameter corresponding to the sample virtual face image, so that the first regressor learns the feature distribution specific to the virtual face. The second regressor for extracting real face features is updated using the pseudo label containing the virtual face features, and the feature distribution specific to the virtual face is transferred to the real face through the method of transfer learning, so that the structural similarity information (including the relative position and proportion of the facial features, etc.) between the virtual face and the real face can be used to assist the second regressor in learning the feature distribution of the real face, thereby enabling the training of the second regressor to converge quickly, shortening the training time, and improving the accuracy of the face parameter extracted by the second regressor, thereby improving the similarity of face shaping.

[0067] In addition, the parameters of the second regressor for extracting real face features are supervised by the pseudo label containing the virtual face features, and then the parameters of the first feature extraction module in the first regressor are updated by the updated model parameters of the second feature extraction module in the second regressor after the model parameter update, so that the first feature extraction module and the second feature extraction module both learn the features common to real faces and virtual faces, the first full connection module learns the features specific to virtual faces, and the second full connection module learns the features specific to real faces. In this way, the second regressor has high accuracy in extracting real face features, and the second regressor retains the features specific to virtual faces, so that the second feature extraction module and the first full connection module have high accuracy in extracting virtual face features.

[0068] The specific implementation of updating the model parameters of the second regressor according to the third predicted face pinching parameters and the pseudo label in step [1] is described in detail below. Specifically, the following steps 1) and 2) can be implemented:

[0069] 1) Calculate the second loss function according to the third predicted face pinching parameters and the pseudo label.

[0070] For example, the absolute value of the difference between the third predicted face pinching parameters and the pseudo label can be determined as the second loss function.

[0071] 2) Update the model parameters of the second regressor in the manner of stochastic gradient descent according to the second loss function.

[0072] At this time, the sum of the loss of the first regressor and the loss of the second regressor is the sum of the first loss function and the second loss function, that is, the preset training stop condition is that the sum of the first loss function and the second loss function is less than the preset loss threshold.

[0073] Figure 3 is a flowchart of an image generation method according to an example embodiment. As shown in Figure 3 , the method includes the following S301-S303.

[0074] In S301, in response to a user request, a target object image is obtained.

[0075] In the present disclosure, the user request can be an image generation request, and the target object image (i.e., a real face image) can be an image containing a user face obtained after user authorization, for example, the target object image can be a target face image. The target object image can be obtained by the user using an image acquisition terminal (for example, a smart phone, a tablet computer, etc.), or can be pre-stored by the user, and the present disclosure does not make specific limitations.

[0076] In S302, based on the target object image, a target face pinching parameter corresponding to the target object image is determined by an image generation model.

[0077] In the present disclosure, the above-mentioned image generation model is obtained by training the above-mentioned image generation model training method provided by the present disclosure.

[0078] In S303, based on the target face pinching parameter, face rendering is performed to obtain a sample virtual object image corresponding to the target object image.

[0079] In the present disclosure, face rendering can be performed by using a preset rendering engine. Specifically, based on the target face pinching parameter, face information reconstruction can be performed by using the preset rendering engine to obtain reconstructed face information corresponding to the target face pinching parameter; and based on the reconstructed face information corresponding to the target face pinching parameter, a target virtual object image corresponding to the target object image is generated.

[0080] In the above technical solution, based on the obtained target object image, an image generation model is used to determine a target face pinching parameter corresponding to the target object image; then, based on the target face pinching parameter, face rendering is performed to obtain a target virtual object image corresponding to the target object image. Since the image generation model can not only learn the real face feature distribution, but also learn the feature distribution specific to the virtual face, the difference between the real face feature distribution and the virtual face feature distribution is reduced, so that a virtual face similar to the user's face can be quickly rendered according to the target face pinching parameter extracted by the image generation model, and the efficiency and similarity of face pinching are high.

[0081] The present disclosure also provides an image generation model training device, as shown in Figure 4 The image generation model training device 400 includes:

[0082] The first acquisition module 401 is configured to acquire a sample object image, a sample virtual object image, and a sample face pinching parameter corresponding to the sample virtual object image.

[0083] The generation module 402 is configured to generate a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face pinching parameter acquired by the first acquisition module 401.

[0084] The training module 403 is configured to perform supervised model training by using the pseudo label generated by the generation module 402 to obtain an image generation model.

[0085] In the technical solution, the pseudo label of the sample object image is generated according to the obtained sample object image, sample virtual object image and sample face pinching parameters corresponding to the sample virtual object image, so that the pseudo label can learn the feature distribution specific to the virtual face. Then, the pseudo label containing the virtual face features is used for supervised model training, so that the image generation model can not only learn the real face feature distribution, but also learn the feature distribution specific to the virtual face, thereby improving the generalization ability of the model, reducing the difference between the real face feature distribution and the virtual face feature distribution, and then quickly rendering a virtual face similar to the user's face according to the face pinching parameters extracted by the image generation model, and the efficiency and similarity of face pinching are high. In addition, the pseudo label containing the virtual face features is used for model training, which can avoid manual labeling of the face pinching parameters corresponding to the sample object image, thereby improving the efficiency of model training and saving labor costs.

[0086] Optionally, the generation module 402 comprises:

[0087] The first updating submodule is configured to perform model parameter updating on the first regressor according to the sample virtual object image and the sample face pinching parameters.

[0088] The first generation submodule is configured to generate a pseudo label of the sample object image by the first regressor after model parameter updating according to the sample object image.

[0089] The training module 403 comprises:

[0090] The second updating submodule is configured to perform model parameter updating on the second regressor according to the sample object image and the pseudo label.

[0091] The third updating submodule is configured to update the parameters of the first regressor after model parameter updating by using the parameters of the second regressor after model parameter updating.

[0092] The determining submodule is configured to determine the second regressor as the image generation model when a preset training stop condition is met.

[0093] Optionally, the first updating submodule comprises:

[0094] The first input submodule is configured to input the sample virtual object image into the first regressor to obtain first predicted face pinching parameters.

[0095] The fourth updating submodule is configured to update the model parameters of the first regressor according to the first predicted face pinching parameters and the sample face pinching parameters.

[0096] Optionally, the fourth updating submodule comprises:

[0097] a first calculation submodule, configured to calculate a first loss function according to the first predicted face-squeezing parameter and the sample face-squeezing parameter;

[0098] a fifth updating submodule, configured to update model parameters of the first regressor in a manner of random gradient descent according to the first loss function.

[0099] Optionally, the first generating submodule is configured to input the sample object image into the first regressor obtained after the model parameters are updated, to obtain a second predicted face-squeezing parameter, and use the second predicted face-squeezing parameter as a pseudo label of the sample object image.

[0100] Optionally, the second updating submodule includes:

[0101] a second input submodule, configured to input the sample object image into the second regressor, to obtain a third predicted face-squeezing parameter;

[0102] a sixth updating submodule, configured to update model parameters of the second regressor according to the third predicted face-squeezing parameter and the pseudo label.

[0103] Optionally, the sixth updating submodule includes:

[0104] a second calculation submodule, configured to calculate a second loss function according to the third predicted face-squeezing parameter and the pseudo label;

[0105] a seventh updating submodule, configured to update model parameters of the second regressor in a manner of random gradient descent according to the second loss function.

[0106] Optionally, the first regressor includes a first feature extraction module and a first full connection module, the second regressor includes a second feature extraction module and a second full connection module, and the second feature extraction module has the same structure as the first feature extraction module.

[0107] The third updating submodule is configured to update parameters of the first feature extraction module in the first regressor obtained after the model parameters are updated, by using parameters of the second feature extraction module in the second regressor obtained after the model parameters are updated.

[0108] Optionally, the first obtaining module 401 includes:

[0109] a second generating submodule, configured to randomly generate a sample face-squeezing parameter;

[0110] a rendering submodule, configured to perform face rendering based on the sample face-squeezing parameter, to obtain a sample virtual object image corresponding to the sample face-squeezing parameter.

[0111] Optionally, the rendering submodule includes:

[0112] a reconstruction submodule, configured to perform face information reconstruction based on the sample face sculpting parameter and by using a preset rendering engine, to obtain reconstructed face information;

[0113] a third generation submodule, configured to generate a sample virtual object image corresponding to the sample face sculpting parameter based on the reconstructed face information.

[0114] Optionally, the sample face sculpting parameter comprises a bone point parameter and / or a makeup parameter.

[0115] Figure 5 is a block diagram of an image generation apparatus according to an exemplary embodiment. As shown in the figure, the image generation apparatus 500 comprises: Figure 5

[0116] a second acquisition module 501 configured to acquire a target object image in response to a user request;

[0117] a face sculpting parameter extraction module 502 configured to determine a target face sculpting parameter corresponding to the target object image by using an image generation model based on the target object image acquired by the second acquisition module 501, wherein the image generation model is obtained by training the image generation model training method provided in the present disclosure;

[0118] a rendering module 503 configured to perform face rendering based on the target face sculpting parameter extracted by the face sculpting parameter extraction module 502, to obtain a target virtual object image corresponding to the target object image.

[0119] In the above technical solution, the target face sculpting parameter corresponding to the target object image is determined by using the image generation model based on the acquired target object image, and then face rendering is performed based on the target face sculpting parameter to obtain a target virtual object image corresponding to the target object image. Since the image generation model can not only learn the real face feature distribution, but also learn the feature distribution specific to the virtual face, the difference between the real face feature distribution and the virtual face feature distribution is reduced, so that a virtual face similar to the user's face can be quickly rendered according to the target face sculpting parameter extracted by the image generation model, and the efficiency and similarity of face sculpting are both high.

[0120] It should be noted that the image generation model training apparatus 300 can be independently arranged from the image generation apparatus 500, or can be integrated in the image generation apparatus 500, which is not specifically limited in the present disclosure.

[0121] The present disclosure also provides a computer readable medium having a computer program stored thereon, wherein the program is executed by a processing apparatus to implement the steps of the image generation method provided in the present disclosure, or the steps of the image generation model training method.​

[0122] The following description refers to Figure 6 , which shows a structural diagram of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and the range of use of the embodiments of the present disclosure.

[0123] As Figure 6 shown, the electronic device 600 can include a processing device (e.g., a central processor, a graphic processor, or the like) 601 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage device 608. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0124] In general, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 608 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or provided. More or fewer devices can alternatively be implemented or provided.

[0125] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program in accordance with embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program comprising program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0126] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device, other than the computer-readable storage medium. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination thereof.

[0127] In some embodiments, the client, server can communicate using any currently known or future developed network protocols, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0128] The computer readable medium described above can be included in the electronic device described above; or can exist separately, without being assembled into the electronic device.

[0129] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a sample object image, a sample virtual object image, and a sample face kneading parameter corresponding to the sample virtual object image; generate a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face kneading parameter; and perform supervised model training using the pseudo label to obtain an image generation model.

[0130] Alternatively, the computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: in response to a user request, obtain a target object image; determine a target face kneading parameter corresponding to the target object image based on the target object image through an image generation model, wherein the image generation model is obtained by training the image generation model training method provided by the present disclosure; and perform face rendering based on the target face kneading parameter to obtain a target virtual object image corresponding to the target object image.

[0131] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0132] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by special-purpose hardware-based systems that perform the specified functions or operations, or combinations of hardware and computer instructions.

[0133] The modules involved in the embodiments of the present disclosure can be implemented in the manner of software or hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the first acquisition module can also be described as a "module for acquiring target user face image".

[0134] The functions described in the foregoing description can be implemented in part or in whole by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0136] According to one or more embodiments of the present disclosure, example 1 provides a method for training an image generation model, comprising: obtaining a sample object image, a sample virtual object image, and a sample face pinching parameter corresponding to the sample virtual object image; generating a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face pinching parameter; and performing supervised model training using the pseudo label to obtain an image generation model.

[0137] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, wherein the generating a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face pinching parameter comprises: performing model parameter updating on a first regressor according to the sample virtual object image and the sample face pinching parameter; and generating the pseudo label of the sample object image by the first regressor after model parameter updating according to the sample object image; and the performing supervised model training using the pseudo label to obtain an image generation model comprises: performing model parameter updating on a second regressor according to the sample object image and the pseudo label; updating parameters of the first regressor after model parameter updating using parameters of the second regressor after model parameter updating; and taking the second regressor as the image generation model when a preset training stop condition is met.

[0138] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, wherein the performing model parameter updating on a first regressor according to the sample virtual object image and the sample face pinching parameter comprises: inputting the sample virtual object image into the first regressor to obtain a first predicted face pinching parameter; and updating model parameters of the first regressor according to the first predicted face pinching parameter and the sample face pinching parameter.

[0139] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, wherein the updating the model parameters of the first regressor according to the first predicted face-squeezing parameters and the sample face-squeezing parameters comprises: calculating a first loss function according to the first predicted face-squeezing parameters and the sample face-squeezing parameters; and updating the model parameters of the first regressor in a manner of stochastic gradient descent according to the first loss function.

[0140] According to one or more embodiments of the present disclosure, example 5 provides the method of example 2, wherein the generating the pseudo label of the sample object image by the first regressor after the model parameter updating comprises: inputting the sample object image into the first regressor after the model parameter updating to obtain second predicted face-squeezing parameters, and taking the second predicted face-squeezing parameters as the pseudo label of the sample object image.

[0141] According to one or more embodiments of the present disclosure, example 6 provides the method of example 2, wherein the model parameter updating of the second regressor according to the sample object image and the pseudo label comprises: inputting the sample object image into the second regressor to obtain third predicted face-squeezing parameters; and updating the model parameters of the second regressor according to the third predicted face-squeezing parameters and the pseudo label.

[0142] According to one or more embodiments of the present disclosure, example 7 provides the method of example 6, wherein the updating the model parameters of the second regressor according to the third predicted face-squeezing parameters and the pseudo label comprises: calculating a second loss function according to the third predicted face-squeezing parameters and the pseudo label; and updating the model parameters of the second regressor in a manner of stochastic gradient descent according to the second loss function.

[0143] According to one or more embodiments of the present disclosure, example 8 provides the method of example 2, wherein the first regressor comprises a first feature extraction module and a first full connection module, the second regressor comprises a second feature extraction module and a second full connection module, the second feature extraction module has the same structure as the first feature extraction module; and the updating the parameters of the first regressor after the model parameter updating by the parameters of the second regressor after the model parameter updating comprises: updating the parameters of the first feature extraction module in the first regressor after the model parameter updating by the parameters of the second feature extraction module in the second regressor after the model parameter updating.

[0144] According to one or more embodiments of the present disclosure, example 9 provides the method of example 1, wherein the obtaining the sample virtual object image and the sample face shaping parameter corresponding to the sample virtual object image comprises: randomly generating the sample face shaping parameter; and performing face rendering based on the sample face shaping parameter to obtain the sample virtual object image corresponding to the sample face shaping parameter.

[0145] According to one or more embodiments of the present disclosure, example 10 provides the method of example 9, wherein the performing face rendering based on the sample face shaping parameter to obtain the sample virtual object image corresponding to the sample face shaping parameter comprises: performing face information reconstruction based on the sample face shaping parameter by using a preset rendering engine to obtain reconstructed face information; and generating the sample virtual object image corresponding to the sample face shaping parameter based on the reconstructed face information.

[0146] According to one or more embodiments of the present disclosure, example 11 provides the method of any one of examples 1-10, wherein the sample face shaping parameter comprises a bone point parameter and / or a makeup parameter.

[0147] According to one or more embodiments of the present disclosure, example 12 provides an image generation method, comprising: obtaining a target object image in response to a user request; determining a target face shaping parameter corresponding to the target object image by an image generation model based on the target object image, wherein the image generation model is obtained by training the image generation model training method of any one of examples 1-11; and performing face rendering based on the target face shaping parameter to obtain a target virtual object image corresponding to the target object image.

[0148] According to one or more embodiments of the present disclosure, example 13 provides an image generation model training apparatus, comprising: a first obtaining module configured to obtain a sample object image, a sample virtual object image, and a sample face shaping parameter corresponding to the sample virtual object image; a generating module configured to generate a pseudo label of the sample object image according to the sample object image, the sample virtual object image, and the sample face shaping parameter obtained by the first obtaining module; and a training module configured to perform supervised model training by using the pseudo label generated by the generating module to obtain an image generation model.

[0149] According to one or more embodiments of the present disclosure, example 14 provides an image generation apparatus, comprising: a second acquisition module configured to acquire a target object image in response to a user request; a face pinching parameter extraction module configured to determine a target face pinching parameter corresponding to the target object image based on the target object image acquired by the second acquisition module, by an image generation model, wherein the image generation model is obtained by training the image generation model training method of any one of examples 1-11; and a rendering module configured to perform face rendering based on the target face pinching parameter extracted by the face pinching parameter extraction module to obtain a target virtual object image corresponding to the target object image.

[0150] According to one or more embodiments of the present disclosure, example 15 provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-12.

[0151] According to one or more embodiments of the present disclosure, example 16 provides an electronic device, comprising: a storage device having stored thereon a computer program; and a processing apparatus configured to execute the computer program in the storage device to implement the steps of the method of any one of examples 1-12.

[0152] The above description is merely exemplary of the disclosure and the application of the principles thereof and it is not intended to limit the scope of the disclosure to the specific forms set forth. The disclosure is susceptible to numerous modifications and variations once the scope of the disclosure is understood. For example, the specific sequences of operations described can be modified in various ways. The recitation of one feature does not exclude the presence of another feature. The specific features recited can be combined in any combination. The disclosure is not limited to the specific forms set forth and it is understood that the disclosure includes all new and useful alternatives, modifications, improvements and equivalents thereto.

[0153] Furthermore, while operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order depicted. Rather, many of the operations can be performed in parallel or in a different order than the order in which they are depicted. In addition, while a number of specific implementation details are discussed herein, these should not be understood as limiting the scope of the disclosure. Rather, certain features discussed in the context of separate embodiments can be combined in a single embodiment. Conversely, various features discussed in the context of a single embodiment can be combined in a plurality of embodiments. The scope of the disclosure is defined by the appended claims and their equivalents.

[0154] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the operations are performed by the various modules has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. A method for training an image generation model, characterized in that, The method comprises the following steps: obtaining a sample object image, a sample virtual object image, and sample face parameter corresponding to the sample virtual object image; updating model parameters of a first regressor according to the sample virtual object image and the sample face parameter; generating pseudo labels of the sample object image by the first regressor after model parameter updating according to the sample object image; updating model parameters of a second regressor according to the sample object image and the pseudo labels; updating parameters of the first regressor after model parameter updating by parameters of the second regressor after model parameter updating; when a preset training stop condition is met, taking the second regressor as an image generation model.

2. The method of claim 1, wherein, The step of updating model parameters of the first regressor according to the sample virtual object image and the sample face parameter comprises the following steps: inputting the sample virtual object image into the first regressor to obtain first predicted face parameters; updating model parameters of the first regressor according to the first predicted face parameters and the sample face parameters.

3. The method of claim 2, wherein, The step of updating model parameters of the first regressor according to the first predicted face parameters and the sample face parameters comprises the following steps: calculating a first loss function according to the first predicted face parameters and the sample face parameters; updating model parameters of the first regressor by a stochastic gradient descent method according to the first loss function.

4. The method of claim 1, wherein, The step of generating pseudo labels of the sample object image by the first regressor after model parameter updating according to the sample object image comprises the following steps: inputting the sample object image into the first regressor after model parameter updating to obtain second predicted face parameters, and taking the second predicted face parameters as the pseudo labels of the sample object image.

5. The method of claim 1, wherein, The step of updating model parameters of the second regressor according to the sample object image and the pseudo labels comprises the following steps: inputting the sample object image into the second regressor to obtain third predicted face parameters; updating model parameters of the second regressor according to the third predicted face parameters and the pseudo labels.

6. The method of claim 5, wherein, The step of updating model parameters of the second regressor according to the third predicted face parameters and the pseudo labels comprises the following steps: calculating a second loss function according to the third predicted face parameters and the pseudo labels; updating model parameters of the second regressor by a stochastic gradient descent method according to the second loss function.

7. The method of claim 1, wherein, The first regressor comprises a first feature extraction module and a first full connection module, and the second regressor comprises a second feature extraction module and a second full connection module, wherein the second feature extraction module has the same structure as the first feature extraction module. The step of updating parameters of the first regressor after model parameter updating by parameters of the second regressor after model parameter updating comprises the following steps: updating parameters of the first feature extraction module in the first regressor after model parameter updating by parameters of the second feature extraction module in the second regressor after model parameter updating.

8. The method of claim 1, wherein, The step of obtaining the sample virtual object image and the sample face parameter corresponding to the sample virtual object image comprises the following steps: Randomly generate sample face pinching parameters; Perform face rendering based on the sample face pinching parameters to obtain a sample virtual object image corresponding to the sample face pinching parameters.

9. The method of claim 8, wherein, The face rendering based on the sample face pinching parameters to obtain a sample virtual object image corresponding to the sample face pinching parameters comprises: Perform face information reconstruction based on the sample face pinching parameters by using a preset rendering engine to obtain reconstructed face information; Generate a sample virtual object image corresponding to the sample face pinching parameters based on the reconstructed face information.

10. The method according to any one of claims 1-9, characterized in that, The sample face pinching parameters comprise bone point parameters and / or makeup parameters.

11. An image generation method characterized by, Comprise: In response to a user request, obtain a target object image; Based on the target object image, determine target face pinching parameters corresponding to the target object image by an image generation model, wherein the image generation model is obtained by training the image generation model training method in any one of claims 1-10; Perform face rendering based on the target face pinching parameters to obtain a target virtual object image corresponding to the target object image.

12. An image generation model training apparatus characterized by comprising: Comprise: A first obtaining module for obtaining a sample object image, a sample virtual object image, and sample face pinching parameters corresponding to the sample virtual object image; A generating module for generating pseudo-labels of the sample object image according to the sample object image, the sample virtual object image, and the sample face pinching parameters obtained by the first obtaining module; A training module for performing supervised model training by using the pseudo-labels generated by the generating module to obtain an image generation model; The generating module comprises: A first updating submodule for updating model parameters of a first regressor according to the sample virtual object image and the sample face pinching parameters; A first generating submodule for generating pseudo-labels of the sample object image by the first regressor after model parameter updating according to the sample object image; The training module comprises: A second updating submodule for updating model parameters of a second regressor according to the sample object image and the pseudo-labels; A third updating submodule for updating parameters of the first regressor after model parameter updating by using parameters of the second regressor after model parameter updating; A determining submodule for taking the second regressor as the image generation model when a preset training stop condition is met.

13. An image generation apparatus characterized by comprising: Comprise: A second obtaining module for obtaining a target object image in response to a user request; A face pinching parameter extraction module for determining target face pinching parameters corresponding to the target object image by an image generation model based on the target object image obtained by the second obtaining module, wherein the image generation model is obtained by training the image generation model training method in any one of claims 1-10; A rendering module for performing face rendering based on the target face pinching parameters extracted by the face pinching parameter extraction module to obtain a target virtual object image corresponding to the target object image.

14. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by a processing device to implement the steps of the method in any one of claims 1-11.

15. An electronic device, comprising: Comprise: A storage device having a computer program stored thereon; processing means for executing the computer program in said storage means to implement the steps of the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Model training method, image processing method and device, equipment and medium

    CN109902767A

  • Model training method and device, information output method and device, equipment and storage medium

    CN113052962A