Model training method and system, face image generation method and system, equipment and medium

By extracting the 3DMM coefficients in the face image and using discrete model encoding, the problems of data redundancy, privacy exposure and feature extraction complexity in traditional image encoding methods are solved, and efficient data storage and transmission are achieved while reducing privacy risks.

CN120107720APending Publication Date: 2025-06-06HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188017.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional image encoding methods in the prior art have problems such as data redundancy, privacy exposure and feature extraction.

Method used

By extracting the 3DMM coefficients in the face image, the target face image generation model is obtained by using preset discrete model training, so that the face shape and texture coefficients are encoded into discrete codewords.

Benefits of technology

It improves the data storage and transmission efficiency of facial images, reduces the use of computing resources, and reduces the risk of user privacy and sensitive information exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107720A_ABST
    Figure CN120107720A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and system, a face image generation method and system, equipment and a medium. The training method of the face image generation model comprises the following steps: acquiring a sample face image; extracting a target 3DMM coefficient of the sample face image; wherein the target 3DMM coefficient comprises a face shape coefficient and a texture coefficient in the sample face image; and inputting the face shape coefficient and the texture coefficient into a preset discretization model, outputting discrete code words corresponding to the sample face image, and training to obtain a target face image generation model. According to the method, coefficients in a sample face image are extracted, a face shape coefficient and a texture coefficient in the coefficients are taken as input, and a target face image generation model is obtained based on preset discretization model training, so that the face shape coefficient and the texture coefficient are coded into discrete code words, and the data storage and transmission efficiency of the face image is improved; the use of computing resources is reduced; and the discretized code is more abstract, and the risk of user privacy and sensitive information exposure is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image model training, and in particular to a method, system, device and medium for model training and generation of facial images. Background Art

[0002] With the development of digital technology and the Internet, face images are increasingly used in a wide range of applications, including but not limited to identity authentication, security monitoring, social networking, etc. In order to meet the needs of these applications, efficient representation and transmission of face images has become crucial.

[0003] Traditional image coding methods mainly focus on the pixel level, such as JPEG (an image format), PNG (an image grid) and other formats. Although these methods can effectively compress image data, they have limitations in specific scenarios: for example, traditional methods retain a large amount of unnecessary information, resulting in increased storage and transmission costs; or, traditional methods directly operate on pixel-level data, leading to the risk of exposing personal privacy, especially in scenarios where sensitive information is shared or transmitted; or, when used for analysis or recognition, traditional methods often require additional preprocessing steps to extract useful features, which increases the complexity of feature extraction and leads to inefficient image processing. Summary of the invention

[0004] The technical problem to be solved by the present disclosure is to overcome the defects of data redundancy, privacy exposure and complex feature extraction in the traditional image encoding method in the prior art, and to provide a model training, face image generation method, system, device, medium and program product

[0005] The present invention solves the above technical problems through the following technical solutions:

[0006] In a first aspect, a training method for a face image generation model is provided, the training method comprising:

[0007] Get a sample face image;

[0008] Extracting target 3DMM (3D Morphable Model) coefficients of the sample face image;

[0009] Wherein, the target 3DMM coefficients include the face shape coefficients and texture coefficients in the sample face image;

[0010] The face shape coefficient and the texture coefficient are input into a preset discretization model, a discrete codeword corresponding to the sample face image is output, and a target face image generation model is obtained through training.

[0011] Optionally, the preset discretization model includes a coding layer, a vector layer and a decoding layer, and the step of inputting the face shape coefficient and the texture coefficient into the preset discretization model, outputting a discrete codeword corresponding to the sample face image, and training to obtain a target face image generation model includes:

[0012] Inputting the face shape coefficient and the texture coefficient into a coding layer to obtain corresponding continuous features;

[0013] Inputting the continuous features into the vector layer to obtain corresponding quantized features;

[0014] Inputting the quantized features into the decoding layer to obtain corresponding reconstructed features, and using the reconstructed features as outputs of the preset discretization model;

[0015] Obtaining a loss function of the preset discretization model based on the face shape coefficient, the texture coefficient and the reconstruction feature;

[0016] The preset discretization model is adjusted based on the loss function to obtain the target face image generation model.

[0017] Optionally, the step of inputting the face shape coefficient and the texture coefficient into a preset discretization model, outputting a discrete codeword corresponding to the sample face image, and training to obtain a target face image generation model further includes:

[0018] During the training process, at least one of minimizing the reconstruction error, the regularization term, and introducing a discretization loss term is performed; and / or,

[0019] Adjusting hyperparameters of the preset discretization model;

[0020] Wherein, the hyperparameter includes at least one of a learning rate, a batch size, and a hidden layer dimension;

[0021] Configuring the preset discretization model based on the hyperparameters to obtain an initial image generation model;

[0022] The target face image generation model is trained using the initial image generation model.

[0023] Optionally, the step of extracting target 3DMM coefficients of the sample face image includes:

[0024] Extracting the original 3DMM system number of the sample face image;

[0025] The original 3DMM coefficients are standardized or normalized to obtain the target 3DMM coefficients.

[0026] A second party provides a method for generating a face image, the method comprising:

[0027] Get the original face image;

[0028] Inputting the original face image into a target face image generation model to generate a discretized target face image;

[0029] Wherein, the target facial image generation model is obtained based on the training method of the facial image generation model described in the first aspect.

[0030] In a third aspect, a training system for a face image generation model is provided, wherein the training system comprises a sample acquisition module, a coefficient extraction module and a model training module;

[0031] The sample acquisition module is used to acquire sample face images;

[0032] The coefficient extraction module is used to extract the target 3DMM coefficients of the sample face image;

[0033] Wherein, the target 3DMM coefficients include the face shape coefficient and the texture coefficient in the sample face image;

[0034] The model training module is used to input the face shape coefficient and the texture coefficient into a preset discretization model, output the discrete codeword corresponding to the sample face image, and train to obtain a target face image generation model.

[0035] In a fourth aspect, a system for generating a face image is provided, the system comprising an original image acquisition module and a discrete image generation module;

[0036] The original image acquisition module is used to acquire the original face image;

[0037] The discrete image generation module is used to input the original face image into the target face image generation model to generate a discretized target face image;

[0038] Among them, the target facial image generation model is obtained based on the training system of the facial image generation model described in the third aspect.

[0039] In a fifth aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, wherein when the processor executes the computer program, the training method for the facial image generation model as described in the first aspect or the method for generating a facial image as described in the second aspect is implemented.

[0040] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the face image generation model described in the first aspect or the method for generating a face image described in the second aspect is implemented.

[0041] In a seventh aspect, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the training method for a facial image generation model as described in the first aspect or the method for generating a facial image as described in the second aspect is implemented.

[0042] On the basis of conforming to the common sense in the art, the above optional conditions can be arbitrarily combined to obtain the optional embodiments of the present disclosure.

[0043] The positive progressive effect of the present disclosure is that by extracting the 3DMM coefficients in the sample face image, taking the face shape coefficients and texture coefficients in the 3DMM coefficients as input, and generating a model for the target face image obtained by training based on a preset discretization model, the face shape and texture coefficients are encoded into discrete codewords, thereby improving the data storage and transmission efficiency of the face image and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of a method for training a facial image generation model provided by an exemplary embodiment of the present disclosure;

[0045] Figure 2 A flowchart of step S103 in a method for training a facial image generation model provided by an exemplary embodiment of the present disclosure;

[0046] Figure 3 A flowchart of a method for generating a face image provided by an exemplary embodiment of the present disclosure;

[0047] Figure 4 A schematic diagram of modules of a training system for a facial image generation model provided by an exemplary embodiment of the present disclosure;

[0048] Figure 5 A schematic diagram of modules of a system for generating a facial image provided by an exemplary embodiment of the present disclosure;

[0049] Figure 6 A schematic diagram of the hardware structure of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0050] The present disclosure is further described below by way of examples, but the present disclosure is not limited to the scope of the examples.

[0051] Prefixes such as "first" and "second" are used in the embodiments of the present disclosure only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and no unnecessary limitation should be constituted due to the use of such prefixes. In addition, in the description of the present embodiment, unless otherwise specified, the meaning of "plurality" is two or more.

[0052] Example 1

[0053] This embodiment provides a training method for a facial image generation model. Figure 1 As shown, the training method comprises:

[0054] S101, obtaining a sample face image;

[0055] S102, extracting target 3DMM coefficients of the sample face image;

[0056] Wherein, the target 3DMM coefficients include the face shape coefficient and the texture coefficient in the sample face image;

[0057] S103, inputting the face shape coefficient and the texture coefficient into a preset discretization model, outputting a discrete codeword corresponding to the sample face image, and training to obtain a target face image generation model.

[0058] In step S101, the sample face images may be preprocessed, including performing face detection, feature point location, cropping and / or normalization on each sample face image to keep the size and format of the sample face image consistent. The sample face images may adopt a public face dataset, such as LFW (Labeled Faces in the Wild, a dataset of labeled face images), CelebA (a dataset of labeled face images), etc.

[0059] Among them, face detection is used to detect the face area from the sample face image to effectively remove the background noise in the image, so that the sample face image input to the model only processes the face area, reducing unnecessary calculations, ensuring that the model focuses on the face part, and improving the accuracy of training and testing.

[0060] Feature point positioning is used to accurately align key feature points in the face (such as eyes, nose or mouth), eliminate the effects of rotation, scaling, etc. in the image, help process faces with different postures and expressions, and improve the robustness of the model.

[0061] Cropping is used to crop the sample face image according to the face area obtained by face detection, ensuring the area standardization of the input image and making the size of the cropped sample face image consistent for subsequent processing and model input.

[0062] Normalization is used to normalize the image pixel values ​​in the sample face images, including color normalization, to reduce the impact of illumination changes and color changes on the model to ensure regional standardization of the input image, which helps to speed up the training speed and convergence of the model.

[0063] For the preprocessing method of the sample face image, a combined processing method of face detection, cropping and normalization can be adopted to ensure that the input image only contains the face area and the image pixel values ​​are standardized. This combined processing method is simple and efficient, with a fast processing speed, and can significantly improve the accuracy of the model; alternatively, a combined processing method of face detection, feature point positioning, alignment, cropping and normalization can be adopted to achieve a more refined processing effect on the sample face image, which can process faces with different postures and expressions and improve the robustness and generalization ability of the model.

[0064] In step S102, a publicly trained 3DMM coefficient extraction model (such as Deep3DFace) can be used to extract 3DMM coefficients from the sample face image. Deep3DFace is a technology based on a deep convolutional neural network (DCNN). It automatically learns the feature representation of face images through a multi-layer convolution structure and maps the features to the shape and texture space of the 3DMM.

[0065] Among them, the 3DMM coefficients include the shape coefficient, texture coefficient, expression coefficient and other coefficients of the face. In this embodiment, the 3DMM coefficients include at least the shape coefficient and texture coefficient of the face; the face shape coefficient represents the geometric shape of the face, specifically the 3D coordinates of the vertices representing the face shape in the image; the face texture coefficient represents the color texture of the face, specifically the RGB value at the vertex in the standard image; other coefficients include lighting and / or projection coefficients.

[0066] In step S103, a target face image generation model is generated by training a preset discretization model using face shape coefficients and texture coefficients as inputs, so that the face shape and texture coefficients are encoded as discrete codewords, thereby improving the data storage and transmission efficiency of face images and reducing the use of computing resources; and the discretized coding is more abstract, reducing the risk of user privacy and sensitive information exposure; the discretized coding is more suitable for some special application scenarios, such as federated learning, encrypted communication, etc.

[0067] In this scheme, by extracting 3DMM coefficients from sample face images, taking face shape coefficients and texture coefficients in the 3DMM coefficients as input, a target face image generation model is generated based on a preset discretization model training, so that the face shape and texture coefficients are encoded as discrete codewords, thereby improving the data storage and transmission efficiency of face images and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure.

[0068] As an implementable manner, the preset discretization model includes a coding layer, a vector layer and a decoding layer, such as Figure 2 As shown, step S103 includes:

[0069] S1031, inputting the face shape coefficient and the texture coefficient into a coding layer to obtain corresponding continuous features;

[0070] S1032, inputting the continuous features into the vector layer to obtain corresponding quantized features;

[0071] S1033, inputting the quantized features into the decoding layer to obtain corresponding reconstructed features, and using the reconstructed features as outputs of the preset discretization model;

[0072] S1034, obtaining a loss function of the preset discretization model based on the face shape coefficient, the texture coefficient and the reconstruction feature;

[0073] S1035. Adjust the preset discretization model based on the loss function to obtain the target facial image generation model.

[0074] In this solution, the vector layer uses VQ-VAE (Vector Quantized Variational Autoencoder). The vector layer quantizes the input continuous features and maps them to corresponding discrete codewords. Each codeword represents a discrete coding unit. In order to construct an effective codeword table, an optimal codeword set can be pre-trained from a large number of 3DMM coefficient samples. The vector layer can also use algorithms such as RVQ (Residual Vector Quantization) and FSQ (Finite Scalar Quantization). By using VQ-VAE to discretize continuous features into a finite number of codewords, the robustness of the features is enhanced, the influence of noise and small disturbances is reduced, and the discrete codewords can be represented by integer indexes, which significantly reduces storage space and computing costs, and improves storage and transmission efficiency. Using the encoder-vector quantization layer-decoder structure, the generative model can efficiently process discrete representations, improve training efficiency and convergence speed, and the decoder can recover high-quality original data from discrete representations, retain detail information, and achieve accurate reconstruction.

[0075] As an achievable manner, the step of inputting the face shape coefficient and the texture coefficient into a preset discretization model, outputting a discrete codeword corresponding to the sample face image, and training to obtain a target face image generation model further includes:

[0076] During the training process, at least one of minimizing the reconstruction error, the regularization term, and introducing the discretization loss term is performed.

[0077] In this scheme, the reconstruction error is minimized in the decoding layer. The discrete codewords are reconstructed back to the data space by the decoder. At this time, the reconstruction error is directly calculated and minimized. In the encoding layer and vector layer, the reconstruction error is not directly calculated, but the output quality and quantization operation of the encoder will directly affect the subsequent reconstruction error, which is used to evaluate the reconstruction ability of the model. Regularization term processing In the encoding stage, the regularization term is introduced to ensure that the distance between the continuous features output by the encoder and the discrete codewords in the codebook is as small as possible, and in the quantization stage, the commitment loss ensures that the codebook vector does not deviate from its corresponding coding features. By minimizing the reconstruction error and regularization term loss, the model parameters are optimized to improve the reconstruction quality and stability of the model. The discretization loss term is introduced to ensure that the distance between the continuous features output by the encoder and the selected discrete codewords is as small as possible, while preventing the codebook vector from deviating excessively from its corresponding coding features.

[0078] As an implementable manner, step S103 further includes:

[0079] Adjusting hyperparameters of the preset discretization model;

[0080] Wherein, the hyperparameter includes at least one of a learning rate, a batch size, and a hidden layer dimension;

[0081] Configuring the preset discretization model based on the hyperparameters to obtain an initial image generation model;

[0082] The target face image generation model is trained using the initial image generation model.

[0083] In this solution, hyperparameters are adjusted before formal training of the preset discrete model, and different hyperparameter combinations are explored through experiments to find the optimal configuration. Hyperparameter adjustment methods include grid search, random search or cross-validation. By adjusting the learning rate, the step size of coefficient update is controlled to avoid non-convergence caused by too large coefficients or too slow convergence caused by too small coefficients; by adjusting the batch size, the number of samples used in each training is controlled. Too small a number of samples will lead to unstable gradient estimation, and a large number of samples will require more memory; by adjusting the hidden layer latitude, the scale of the model is controlled. A model that is too small will lead to underfitting, while a model that is too large will lead to overfitting.

[0084] In one embodiment, the hyperparameters also include a codebook size, which is used to control the number of codewords for discrete representation when the VQ-VAE algorithm is used in the vector layer. The preferred codebook size is 1024.

[0085] As an achievable manner, the step of extracting the target 3DMM coefficients of the sample face image includes:

[0086] Extracting the original 3DMM system number of the sample face image;

[0087] The original 3DMM coefficients are standardized or normalized to obtain the target 3DMM coefficients.

[0088] In this solution, the extracted 3DMM coefficients are Z-score normalized to ensure that all coefficients have zero mean and unit variance, which is convenient for subsequent VQ operations. The normalization formula is:

[0089] ;

[0090] Where x is the original 3DMM coefficient, is the mean, is the standard deviation, are the target 3DMM coefficients obtained after standardization.

[0091] In addition, normalization is to scale the 3DMM coefficients to the range of [0, 1] or [-1, 1], depending on the requirements of the application scenario. For example, normalization compresses the data to a specified range, which is suitable for algorithms that are sensitive to value ranges (such as activation functions in deep learning). By limiting the value range, large values ​​can be prevented from causing numerical instability in the model, so that the normalized target 3DMM coefficients can be stored or transmitted in a more compact format.

[0092] The training method for the face generation model provided in this embodiment extracts 3DMM coefficients from a sample face image, takes the face shape coefficients and texture coefficients in the 3DMM coefficients as input, and generates a target face image generation model obtained by training based on a preset discretization model, so that the face shape and texture coefficients are encoded into discrete codewords, thereby improving the data storage and transmission efficiency of face images and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure.

[0093] Example 2

[0094] This embodiment provides a method for generating a face image. Figure 3 As shown, the generating method comprises:

[0095] S201, obtaining an original face image;

[0096] S202, inputting the original face image into a target face image generation model to generate a discretized target face image;

[0097] Wherein, the target facial image generation model is obtained based on the training method of the facial image generation model described in Example 1.

[0098] The facial image generation method provided in this embodiment uses the target facial image generation model obtained by training in Example 1, so that the facial shape and texture coefficients in the original facial image are encoded as discrete codewords, thereby improving the data storage and transmission efficiency of the facial image and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure.

[0099] Example 3

[0100] This embodiment provides a training system 100 for a facial image generation model, such as Figure 4 As shown, the training system includes a sample acquisition module 101, a coefficient extraction module 102 and a model training module 103;

[0101] The sample acquisition module 101 is used to acquire a sample face image;

[0102] The coefficient extraction module 102 is used to extract the target 3DMM coefficients of the sample face image;

[0103] Wherein, the target 3DMM coefficients include the face shape coefficients and texture coefficients in the sample face image;

[0104] The model training module 103 is used to input the face shape coefficient and the texture coefficient into a preset discretization model, output a discrete codeword corresponding to the sample face image, and train to obtain a target face image generation model.

[0105] As an implementable manner, the preset discretization model includes a coding layer, a vector layer and a decoding layer, and the model training module 103 includes a continuous feature coding unit, a vector quantization unit, a decoding reconstruction unit, a loss function establishment unit and a model adjustment unit:

[0106] A continuous feature encoding unit, used for inputting the face shape coefficient and the texture coefficient into a coding layer to obtain corresponding continuous features;

[0107] A vector quantization unit, used for inputting the continuous features into the vector layer to obtain corresponding quantized features;

[0108] A decoding and reconstruction unit, used for inputting the quantized features into the decoding layer to obtain corresponding reconstructed features, and using the reconstructed features as outputs of the preset discretization model;

[0109] A loss function establishing unit, used for obtaining a loss function of the preset discretization model based on the face shape coefficient, the texture coefficient and the reconstruction feature;

[0110] A model adjustment unit is used to adjust the preset discretization model based on the loss function to obtain the target face image generation model.

[0111] As an implementable manner, the model training module 103 further includes an optimization processing unit:

[0112] an optimization processing unit, used to perform at least one of minimizing a reconstruction error, a regularization term, and introducing a discretization loss term during the training process; and / or,

[0113] The model training module 103 also includes a hyperparameter adjustment unit and an initial model configuration unit;

[0114] A hyperparameter adjustment unit, used to adjust the hyperparameters of the preset discretization model;

[0115] Wherein, the hyperparameter includes at least one of a learning rate, a batch size, and a hidden layer dimension;

[0116] An initial model configuration unit is used to configure the preset discretization model based on the hyperparameters to obtain an initial image generation model; and to train the initial image generation model to obtain the target face image generation model.

[0117] As an implementable manner, the coefficient extraction module 102 includes an original coefficient extraction unit and a coefficient processing unit:

[0118] An original coefficient extraction unit, used for extracting the original 3DMM system coefficients of the sample face image;

[0119] The coefficient processing unit is used to perform standardization or normalization processing on the original 3DMM coefficients to obtain the target 3DMM coefficients.

[0120] As for the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The system embodiment described above is only illustrative, in which the units described as separate components may or may not be physically separated, and the components as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution.

[0121] The training system for the face generation model provided in this embodiment extracts 3DMM coefficients from sample face images, takes face shape coefficients and texture coefficients in the 3DMM coefficients as input, and generates a target face image generation model obtained by training based on a preset discretization model, so that the face shape and texture coefficients are encoded into discrete codewords, thereby improving the data storage and transmission efficiency of face images and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure.

[0122] Example 4

[0123] This embodiment provides a system 200 for generating a face image. Figure 5 As shown, the generation system 200 includes an original image acquisition module 201 and a discrete image generation module 202;

[0124] The original image acquisition module 201 is used to acquire the original face image;

[0125] The discrete image generation module 202 is used to input the original face image into the target face image generation model to generate a discretized target face image;

[0126] Among them, the target facial image generation model is obtained based on the training system of the facial image generation model described in Example 3.

[0127] The facial image generation system provided in this embodiment uses the target facial image generation model obtained by training in Example 3, so that the facial shape and texture coefficients in the original facial image are encoded as discrete codewords, thereby improving the data storage and transmission efficiency of the facial image and reducing the use of computing resources; and the discretized encoding is more abstract, reducing the risk of user privacy and sensitive information exposure.

[0128] Example 5

[0129] Figure 6 This is a schematic diagram of the structure of an electronic device shown in an example embodiment of the present disclosure, the electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor, and when the processor executes the computer program, it implements the training method of the face image generation model or the generation method of the face image of any of the above embodiments. Figure 6 The electronic device 90 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0130] like Figure 6 As shown, the electronic device 90 may be in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 90 may include, but are not limited to: at least one processor 91, at least one memory 92, and a bus 93 connecting different system components (including the memory 92 and the processor 91).

[0131] The bus 93 includes a data bus, an address bus, and a control bus.

[0132] The memory 92 may include a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read only memory (ROM) 923 .

[0133] The memory 92 may also include a program tool 925 (or utility) having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0134] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the training method of the face image generation model or the face image generation method provided in any of the above embodiments.

[0135] The electronic device 90 may also communicate with one or more external devices 94 (e.g., keyboards, pointing devices, etc.). Such communication may be performed via an input / output (I / O) interface 95. Furthermore, the electronic device 90 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 96. As shown, the network adapter 96 communicates with other modules of the electronic device 90 via a bus 93. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0136] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.

[0137] Example 6

[0138] The embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for training a facial image generation model or the method for generating a facial image provided in any of the above embodiments is implemented.

[0139] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.

[0140] Example 7

[0141] The embodiments of the present disclosure also provide a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for training a facial image generation model or methods for generating facial images.

[0142] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed completely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or completely on the remote device.

[0143] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A training method for a facial image generation model, characterized in that: The training method comprises: Get a sample face image; Extracting target 3DMM coefficients of the sample face image; Wherein, the target 3DMM coefficients include the face shape coefficients and texture coefficients in the sample face image; The face shape coefficient and the texture coefficient are input into a preset discretization model, a discrete codeword corresponding to the sample face image is output, and a target face image generation model is obtained through training.

2. The training method for a facial image generation model according to claim 1, characterized in that: The preset discretization model includes a coding layer, a vector layer and a decoding layer. The steps of inputting the face shape coefficient and the texture coefficient into the preset discretization model, outputting the discrete codeword corresponding to the sample face image, and training to obtain the target face image generation model include: Inputting the face shape coefficient and the texture coefficient into a coding layer to obtain corresponding continuous features; Inputting the continuous features into the vector layer to obtain corresponding quantized features; Inputting the quantized features into the decoding layer to obtain corresponding reconstructed features, and using the reconstructed features as outputs of the preset discretization model; Obtaining a loss function of the preset discretization model based on the face shape coefficient, the texture coefficient and the reconstruction feature; The preset discretization model is adjusted based on the loss function to obtain the target face image generation model.

3. The training method for a facial image generation model according to claim 2, characterized in that: The step of inputting the face shape coefficient and the texture coefficient into a preset discretization model, outputting a discrete codeword corresponding to the sample face image, and training to obtain a target face image generation model also includes: During the training process, at least one of minimizing the reconstruction error, the regularization term, and introducing a discretization loss term is performed; and / or, Adjusting hyperparameters of the preset discretization model; Wherein, the hyperparameter includes at least one of a learning rate, a batch size, and a hidden layer dimension; Configuring the preset discretization model based on the hyperparameters to obtain an initial image generation model; The target face image generation model is trained using the initial image generation model.

4. The training method for a facial image generation model according to any one of claims 1 to 3, characterized in that: The step of extracting the target 3DMM coefficients of the sample face image comprises: Extracting the original 3DMM system number of the sample face image; The original 3DMM coefficients are standardized or normalized to obtain the target 3DMM coefficients.

5. A method for generating a face image, characterized in that: The generation method comprises: Get the original face image; Inputting the original face image into a target face image generation model to generate a discretized target face image; Wherein, the target facial image generation model is obtained based on the training method of the facial image generation model described in any one of claims 1 to 4.

6. A training system for a facial image generation model, characterized in that: The training system includes a sample acquisition module, a coefficient extraction module and a model training module; The sample acquisition module is used to acquire sample face images; The coefficient extraction module is used to extract the target 3DMM coefficients of the sample face image; Wherein, the target 3DMM coefficients include the face shape coefficients and texture coefficients in the sample face image; The model training module is used to input the face shape coefficient and the texture coefficient into a preset discretization model, output the discrete codeword corresponding to the sample face image, and train to obtain a target face image generation model.

7. A system for generating a face image, characterized in that: The generation system includes an original image acquisition module and a discrete image generation module; The original image acquisition module is used to acquire the original face image; The discrete image generation module is used to input the original face image into the target face image generation model to generate a discretized target face image; Wherein, the target facial image generation model is obtained based on the training system of the facial image generation model described in claim 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the method for training a facial image generation model as described in any one of claims 1 to 4 or the method for generating a facial image as described in claim 5 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for training a facial image generation model as described in any one of claims 1 to 4 or the method for generating a facial image as described in claim 5 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for training a facial image generation model as described in any one of claims 1 to 4 or the method for generating a facial image as described in claim 5 is implemented.