An image processing method and apparatus
The method aligns real and game face images using a transfer model and generation techniques to ensure consistent style, addressing the challenge of optimizing game character face parameters across different environments.
Patent Information
- Application Number
- CN202210499957.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-09
AI Technical Summary
In the prior art, due to the style differences between real faces and game faces, the optimization process of face pinching parameters is affected, and the game face parameter generation effect is not good.
By obtaining the sample dataset, including real face images and stylized face images, the pre-trained migration model is trained using the migration model to generate stylized face images, and optimize the face pinching parameters based on the stylized face image to ensure that the generated face pinching parameters are optimized under the same style.
It effectively avoids the impact of the optimization process of face pinching parameters caused by style differences, and the generated face pinching parameters are more consistent, ensuring that the generated game face image and the face image with optimized face pinching parameters are both game style, which improves the generation effect of face pinching parameters.
Smart Images

Figure CN115063513B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network technology, and specifically to an image processing method. This application also relates to an image processing apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the existing image processing technology, when converting an input real face image into a game face image in a virtual scene, the existing method of estimating game face parameters (also known as face sculpting parameters) using a neural network has achieved good results. The specific process is as follows: By comparing the real face image with the game face image generated by the face sculpting parameters to be optimized, and adjusting the face sculpting parameters based on the comparison result to optimize the face sculpting parameters. However, the real face image and the game face image are images of different styles, and there are significant differences in their shapes, textures, and facial feature distributions. Therefore, if directly comparing the real face with the game face when estimating the game face parameters, for example, using a discriminator trained with real faces to measure the difference between the cross-style real face and the game face, and optimizing the face sculpting parameters based on this difference, it will affect the optimization process of the face sculpting parameters and the generation effect of the game face parameters. For example, the game face generated based on these face sculpting parameters may not be similar to the original real face. Summary of the Invention
[0003] Embodiments of this application provide an image processing method, apparatus, electronic device, and computer-readable storage medium to solve the problem in the prior art that due to the style difference between the real face and the game face, the optimization process of the game face parameters is affected, and the generation effect of the game face parameters is affected.
[0004] Embodiments of this application provide an image processing method, including:
[0005] Obtain a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and the corresponding stylized face images are images generated based on the same input vector;
[0006] Obtain a first real face image from the sample data set, input the first real face image into a migration model, and output a second stylized face image through the migration model, where in the sample data set, the first real face image corresponds to a first stylized face image;
[0007] Train the migration model based on the first stylized face image and the second stylized face image to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain stylized face images.
[0008] Optionally, the method further includes:
[0009] Obtaining a stylized face image;
[0010] Training a pre-trained first generation model based on the stylized face image to obtain a second generation model, wherein the pre-trained first generation model is trained based on real face images;
[0011] Based on the input vector, generating a real face image and a stylized face image through the first generation model and the second generation model respectively;
[0012] Based on the real face image and the stylized face image, obtaining a sample data set.
[0013] Optionally, the method further includes:
[0014] Inputting a target real face image into the pre-trained transfer model, and obtaining a transferred image through the pre-trained transfer model, where the transferred image is a stylized face image;
[0015] Generating an initial image according to the first image parameter;
[0016] Updating the first image parameter according to the transferred image and the initial image to obtain a target image parameter.
[0017] Optionally, the updating the first image parameter according to the transferred image and the initial image includes:
[0018] Extracting a first image feature from the transferred image and extracting a second image feature from the initial image;
[0019] Comparing the first image feature and the second image feature, and adjusting the first image parameter based on the comparison result to obtain an updated first image parameter.
[0020] Optionally, the adjusting the first image parameter based on the comparison result to obtain an updated first image parameter includes:
[0021] In response to the similarity between the first image feature and the second image feature reaching a target similarity threshold, determining the first image parameter as the target parameter;
[0022] Alternatively, in response to the similarity between the first image feature and the second image feature not reaching the target similarity threshold, adjust the first image parameter to obtain a second image parameter; generate a second face image based on the second image parameter, extract a third image feature from the second face image, and obtain the similarity between the first image feature and the third image feature; in response to the similarity between the first image feature and the third image feature reaching the target similarity threshold, or the number of times of adjusting the first image parameter reaching the target number of times, obtain the updated first image parameter.
[0023] Optionally, the method further includes:
[0024] Generate a target image based on the target image parameter.
[0025] Optionally, the training of the pre-trained first generation model based on the stylized face image to obtain a second generation model includes:
[0026] Train the pre-trained first generation model based on the stylized face image to obtain a pre-trained third generation model;
[0027] Mix the convolutional layer of the third generation model with the convolutional layer of the first generation model to obtain a second generation model.
[0028] Optionally, the mixing of the convolutional layer of the third generation model with the convolutional layer of the first generation model includes:
[0029] Based on the selected number of mixing layers, mix the convolutional layer of the third generation model with the convolutional layer of the first generation model, where the number of mixing layers is used to characterize the stylization degree of the stylized face image to be generated by the second generation model.
[0030] An embodiment of the present application further provides an image processing apparatus, and the apparatus includes:
[0031] A sample data set acquisition unit, configured to acquire a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and the corresponding stylized face images are images generated based on the same input vector;
[0032] A second stylized face image output unit, configured to obtain a first real face image from the sample data set, input the first real face image into a migration model, and output a second stylized face image through the migration model, where in the sample data set, the first real face image corresponds to a first stylized face image;
[0033] A migration model training unit is used to train the migration model based on the first stylized face image and the second stylized face image to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain a stylized face image.
[0034] An embodiment of the present application further provides an electronic device, including a processor and a memory; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the above method.
[0035] An embodiment of the present application further provides a computer-readable storage medium, on which one or more computer instructions are stored, and the instructions are executed by a processor to implement the above method.
[0036] Compared with the prior art, the embodiments of the present application have the following advantages:
[0037] The image processing method provided by the embodiments of the present application first obtains a sample data set, which includes real face images and corresponding stylized face images, and the real face images and corresponding stylized face images are images generated based on the same input vector; obtains a first real face image from the sample data set, inputs the first real face image into the migration model, and outputs a second stylized face image through the migration model, where the first real face image corresponds to a first stylized face image in the sample data set; trains the migration model based on the first stylized face image and the second stylized face image to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain a stylized face image. By using the above migration model, a stylized face image corresponding to a real face image can be obtained. When generating face pinching parameters, compared with directly using a real face image to optimize the face pinching parameters, using the stylized face image (inputting a target real face image into the migration model to obtain a migration image corresponding to the target real face graph output by the migration model) obtained by training the migration model using the above method to optimize the face pinching parameters can ensure that the face image generated by the face pinching parameters to be optimized and the face image used to optimize the face pinching parameters are both stylized face images of the same style, for example, both are face images in a game style. Therefore, it is possible to avoid the problems that the optimization process of the face pinching parameters is affected and the effect of the generated face pinching parameters is affected because the face image used to optimize the face pinching parameters is a real face image and the face image generated by the face pinching parameters to be optimized is a stylized face image, thereby obtaining more effective face pinching parameters. Description of the Drawings
[0038] Figure 1 It is a hardware structure block diagram of a mobile terminal that can implement the image processing method provided by the embodiments of the present application;
[0039] Figure 2 It is a flowchart of an image processing method of the method provided by an embodiment of the present application;
[0040] Figure 2-A It is a schematic diagram of the generation result of the hybrid model provided by an embodiment of the present application;
[0041] Figure 2-B It is a schematic diagram of the parameter optimization process provided by an embodiment of the present application;
[0042] Figure 3 It is a block diagram of the units of an image processing device provided by an embodiment of the present application;
[0043] Figure 4 It is a schematic diagram of the logical structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0044] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0045] It should be noted that the terms "first", "second", "third", etc. in each part of the embodiments of the present application and the drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. Such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described herein. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0046] Face pinching is a gameplay in the game scene. When creating a virtual character, the face pinching system can be used to personalize the virtual character to meet the personalized needs of game players. The face pinching system can provide players with rich control points to adjust the shapes of various positions on the face. By setting the target parameters corresponding to each control point, images of game characters with different expressions can be obtained. During the generation process of face pinching parameters, in order to avoid affecting the optimization process of face pinching parameters due to the inconsistency between the real face image and the game character style, and to avoid affecting the effect of the generated face pinching parameters, this application provides an image processing method, an image processing device corresponding to this method, an electronic device capable of implementing this method, and a computer-readable storage medium. Embodiments are provided below to describe the above method, device, electronic device, and computer-readable storage medium in detail.
[0047] The first embodiment of this application provides an image processing method. The application subject of this method can be a computing device application for generating face pinching parameters. This computing device application can run on a game platform server, a mobile terminal, a computer terminal, or a similar computing device. That is, the method provided in this embodiment can be executed on a game platform server, a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is the hardware structure block diagram of the mobile terminal capable of implementing the image processing method provided by the embodiment of the present invention. As Figure 1 shown, the mobile terminal 10 may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA) and a memory 104 for storing data. Optionally, the above mobile terminal 10 may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above mobile terminal 10. For example, the mobile terminal 10 may further include more or fewer components than Figure 1 shown in the figure, or have a structure different from Figure 1The different configurations shown. The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the image processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. The memory 104 can further include memories remotely set relative to the processor 102, and these remote memories can be connected to the mobile terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network can include the wireless network provided by the communication provider of the mobile terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0048] Figure 2 It is a flowchart of the image processing method provided by an embodiment of the present application. The following will be combined with Figure 2 to describe the method provided in this embodiment in detail. The embodiments involved in the following description are used to explain the principle of the method and are not limitations for actual use.
[0049] As Figure 2 shown, the image processing method provided in this embodiment includes the following steps:
[0050] S201, obtain a sample data set.
[0051] This step is used to obtain a sample data set, which is used to train a migration model. The sample data set includes real face images and corresponding stylized face images. The real face images and the corresponding stylized face images are images generated based on the same input vector. The above migration model is used to output a stylized face image corresponding to the real face image in the input real scene based on the input real face image in the real scene. The stylized face image can be the face image corresponding to the above real face image in a virtual scene. The virtual scene can be a game scene, and the above stylized face image is a game-style face image.
[0052] In this embodiment, the above sample data set is obtained in the following manner:
[0053] First, obtain a stylized face image. For example, randomly adjust the facial parameters (such as eye size, eye distance, eye height, nose height, etc.) of a game character (any game character), generate a face photo of the game character from the game engine, form a face dataset of the game character, and this face dataset of the game character is the stylized face image;
[0054] Secondly, train a pre-trained first generation model based on the stylized face image to obtain a second generation model. Among them, the pre-trained first generation model is trained based on real face images and can generate real face images based on randomly generated input vectors;
[0055] Then, based on the input vector, generate a real face image and a stylized face image through the first generation model and the second generation model respectively;
[0056] Finally, obtain a sample dataset based on the real face image and the stylized face image.
[0057] The above-mentioned training of the pre-trained first generation model based on the stylized face image to obtain a second generation model specifically refers to: training the pre-trained first generation model based on the above-mentioned stylized face image to obtain a pre-trained third generation model, and this third generation model can generate realistic stylized face images; mixing the convolutional layers of the third generation model with the convolutional layers of the first generation model to obtain a second generation model. For example, use the game face image dataset to fine-tune the pre-trained first generation model of the real scene (such as the StyleGAN model framework, which can output generated facial images of the real scene based on input random numbers), that is, continue to train the already trained first generation model StyleGAN of the real scene on the game face image dataset for a period of time, so that the third generation model obtained after fine-tuning can generate realistic game face pictures;
[0058] The above-mentioned mixing of the convolutional layers of the third generation model and the convolutional layers of the first generation model specifically means: based on the selected number of mixing layers, mixing the convolutional layers of the third generation model and the convolutional layers of the first generation model, where the number of mixing layers is used to characterize the degree of stylization of the stylized face image to be generated by the second generation model. Taking a game scenario as an example, the above-mentioned process of mixing the convolutional layers of the third generation model and the convolutional layers of the first generation model can specifically be: using the Model Blending (a model blending technology that fuses two different generation models into one model, and the fused model can generate intermediate results of the two models) technology to blend the fine-tuned third generation model and the above-mentioned first generation model based on Layer Swapping. The second generation model after mixing can generate pictures between the real face style and the game face style. Layer Swapping is a model blending strategy, that is, a method of combining different layers of two models with the same structure. For example, if the model structure has a total of six layers, the weights of the first two layers are from one model, and the weights of the last four layers are from another model. It achieves the generation of intermediate results by replacing network layers. For example, for a network with 6 layers, for two models A and B trained based on different training sets, the first and second layers of model A and the third, fourth, fifth, and sixth layers of model B can be used to form a new model, and this new model can generate intermediate results between the two models.
[0059] In this embodiment, based on the predetermined degree of stylization of the stylized face image to be generated by the second generation model to be measured, the number of mixing layers can be selected, and based on this number of mixing layers, the convolutional layers of the first generation model and the convolutional layers of the third generation model are mixed. As Figure 2-A shown, if the starting layers of mixing of the first generation model and the third generation model are different, the degree of mixing of the two models is different. Therefore, according to the different starting layers of mixing, the degree of model mixing can be controlled. Figure 2-A The output results of the second generation model after mixing are given. The first column on the left gives the generated real face pictures, and the subsequent columns give the results after mixing with different numbers of layers. From Figure 2-A this, it can be seen that if the number of mixing layers is more, the generated result is closer to the game-style face. If the number of mixing layers is less, that is, the starting layer of mixing is higher (starting to mix from the x-th layer, the larger x is, the fewer the number of mixing layers), the generated result is closer to the real face. In this embodiment, it can be selected to start mixing from the third layer to obtain a result between the two styles of the real scenario and the game scenario.
[0060] After obtaining the second facial image generation model after the above mixing, a random input vector (i.e., random parameter) is generated, and after inputting this vector into the first generation model, a real face image sample of the real scene output by the first generation model can be obtained. After inputting this vector into the mixed second generation model, a stylized face image sample of the virtual scene output by the second generation model can be obtained. Repeating this process multiple times, the real face image samples and the stylized face image samples constitute an image comparison set of multiple different real face images and corresponding multiple different stylized face images.
[0061] S202, obtain a first real face image from the sample dataset, input the first real face image into the migration model, and output a second stylized face image through the migration model, where in the sample dataset, the first real face image corresponds to a first stylized face image.
[0062] After obtaining the sample dataset containing real face images and corresponding stylized face images in the above steps, and after constructing the style transfer model, this step is used to input the first real face image in the sample dataset into the constructed migration model to obtain the second stylized face image output by the migration model.
[0063] S203, based on the first stylized face image and the second stylized face image, train the migration model to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain stylized face images.
[0064] After obtaining the second stylized face image output by the migration model in the above steps, and after determining the first stylized face image corresponding to the first real face image from the sample dataset, this step is based on the above first stylized face image and the second stylized face image to train the migration model to obtain a pre-trained migration model. The process of model training is to calculate the loss function, use the gradient descent of the loss function to update the model parameters of the migration model, and save the model parameters after reaching the maximum number of iterations. In this embodiment, a joint loss function can be used to supervise the training process of the above migration module, and this joint loss function is composed of the following three parts:
[0065] 1. Reconstruction loss function L rec : In order to directly constrain the generation result of the migration model, in this embodiment, a pre-trained VGG network is used to calculate a series of feature maps for the first stylized face image I d (ground truth map) in the sample dataset and the second stylized face image (generated map) output by the migration model based on the first real face image respectively, and calculate the L1 error between the two. Specifically, first, for the ground truth map I d , generated map Perform image pyramid sampling for images with resolutions of 256×256, 128×128, and 64×64. The image pyramid performs multi-scale sampling on the images, which can measure the similarity of two facial images at multiple scales. Then, each sampled result is sent into the pre-trained VGG network respectively to obtain a series of feature maps. Finally, calculate the L1 error of the corresponding feature maps at the corresponding resolutions and sum up all terms. The calculation process is as follows:
[0066]
[0067] Among them, F i (·) represents the function for extracting the i-th feature map, L is the number of feature maps, and P represents the number of image pyramid samplings.
[0068] 2. Feature matching loss function L FM : To make the training process more stable, this embodiment introduces a feature matching loss. Send the ground truth image I d , and the generated image of the transfer model into the discriminator respectively, extract feature maps for each layer, calculate the L1 error of the corresponding feature maps, and finally sum up all error terms to obtain the feature matching loss. The feature matching loss function can be expressed in the following form:
[0069]
[0070] 3. Adversarial loss function L adv : To make the stylized face image generated by the transfer model more realistic, this embodiment adds the WGAN-GP adversarial loss function. The mathematical form of this loss function is as follows:
[0071]
[0072] Among them, D is the discriminator, is the image after linearly uniformly sampling I d , .
[0073] Finally, the joint loss function proposed in this embodiment can be expressed in the following form:
[0074]
[0075] Among them, λ is an adjustable weight.
[0076] In this embodiment, after obtaining the pre-trained transfer model through the above process, it is also necessary to use this pre-trained transfer model to obtain the target image parameters in the following way, and the target image parameters are the face pinching parameters:
[0077] Input the target real face image into the above-mentioned pre-trained transfer model, and obtain a transferred image through the pre-trained transfer model. The transferred image is a stylized face image. In this embodiment, the target real face image can be a face image captured in a real scene or a face image in a real scene generated by a facial generation model. Moreover, an initial image is generated according to first image parameters. The first image parameters can be random game face-molding parameters. Input the random game face-molding parameters into the generator of the game engine to obtain the output initial image. As Figure 2-B shown, input the game face-molding parameters (first image parameters) into the generator to generate an initial image Input the target real face image I s into the pre-trained transfer model to obtain the output transferred image I t .
[0078] Then, update the first image parameters according to the transferred image and the initial image to obtain target image parameters. The target image parameters can be used to generate a stylized face image in a virtual scene. The target image parameters can determine the attributes of the stylized face image in the virtual scene. For example, the target image parameters are a series of parameters for determining the size, shape, style, etc. of parts such as hair, eyebrows, eyes, nose, mouth, etc. In this embodiment, the target image parameters can be multi-dimensional face-molding parameters, which can be used to control the game face model. By adjusting the values of the target image parameters, a personalized face image of a game character can be made. The similarity between the stylized face image in the virtual scene generated based on the target image parameters and the above-mentioned target real image reaches or exceeds the target similarity threshold. The adjustment of the first image parameters according to the transferred image and the initial image is essentially a process of optimizing the parameters using the transferred image output by the transfer model, so that the stylized face image of the virtual scene generated by the optimized target image parameters can be closer to the above-mentioned transferred image. The specific process is as follows:
[0079] Extract first image features from the transferred image, and extract second image features from the initial image; compare the first image features and the second image features, and adjust the first image parameters based on the comparison result to obtain target image parameters.
[0080] For example, in response to the similarity between the first image features and the second image features reaching the target similarity threshold, determine the first image parameters as the target parameters;
[0081] Alternatively, in response to the similarity between the first image feature and the second image feature not reaching the target similarity threshold, adjust the first image parameter to obtain a second image parameter; generate a second face image based on the second image parameter, extract a third image feature from the second face image, and obtain the similarity between the first image feature and the third image feature; in response to the similarity between the first image feature and the third image feature reaching the target similarity threshold, or the number of times of adjusting the first image parameter reaching the target number of times, obtain the updated first image parameter, and determine the updated first image parameter as the target image parameter.
[0082] In this embodiment, the similarity between the first image feature and the second image feature can be used to indicate the similarity degree of the migrated image and the initial image in terms of image content and image style, and can be characterized by a loss function. The loss function is the objective function in the parameter optimization process, that is, first calculate the loss function, calculate the gradient of the parameter with respect to the loss, update the first image parameter according to the gradient, stop updating after reaching the maximum number of iterations, and output the face pinching parameter. Specifically, as Figure 2-B shown, the L1 error can be used to measure the similarity degree of the migrated image and the initial image in terms of content. When the migrated image and the initial image are closer, for example, the game face image and the real face image are closer, the L1 error is smaller. In this embodiment, the loss function between the first image feature and the second image feature can be represented by the following formula (face segmentation network):
[0083]
[0084] This loss function is a face segmentation loss function, which can measure the gap between the facial features' shapes and distributions of the migrated image and the initial image, where represents the function of the face segmentation network to extract the i-th feature map, L represents the number of feature maps, and P represents the number of image pyramid samplings.
[0085] In this embodiment, the loss function between the first image feature and the second image feature can also be represented by the following formula (face recognition network):
[0086]
[0087] This loss function is a face distance loss function, which can measure the face identity difference between the migrated image and the initial image. represents the function of the face recognition network to extract the i-th feature map, L is the number of feature maps, and P represents the number of image pyramid samplings.
[0088] After obtaining the target image parameters based on the first facial image in the real scenario, it is also possible to perform image rendering in the virtual scenario based on the target image parameters to obtain a target image, which is a stylized face image. For example, the target image parameters are the face sculpting parameters used in a game scenario. After inputting the target image parameters (face sculpting parameters) into the generator of the game engine, the corresponding game face picture can be output.
[0089] In this embodiment, as an alternative implementation, the above-mentioned initial image, transfer image, and target image can all be three-dimensional facial images, thereby enhancing the authenticity of migrating a real face in the real scenario to a stylized face image in the virtual scenario and improving the user experience.
[0090] The image processing method provided in this embodiment first obtains a sample data set, which includes real face images and corresponding stylized face images. The real face images and the corresponding stylized face images are images generated based on the same input vector. Obtain a first real face image from the sample data set, input the first real face image into the transfer model, and output a second stylized face image through the transfer model, where the first real face image in the sample data set corresponds to a first stylized face image. Based on the first stylized face image and the second stylized face image, train the transfer model to obtain a pre-trained transfer model, and the pre-trained transfer model is used to obtain stylized face images. By using this pre-trained transfer model, a stylized face image corresponding to a real face image can be obtained, realizing the migration of a real face image to a stylized face image. When generating face sculpting parameters, compared with directly using a real face image to optimize the face sculpting parameters, using the stylized face image obtained by training the transfer model with the above method to optimize the face sculpting parameters can ensure that the face image generated by the face sculpting parameters to be optimized and the face image used to optimize the face sculpting parameters are both stylized face images, for example, both are game-style face images. Therefore, it is possible to avoid the problems that the optimization process of the face sculpting parameters is affected and the effect of the generated face sculpting parameters is affected because the face image used to optimize the face sculpting parameters is a real face image and the face image generated by the face sculpting parameters to be optimized is a stylized face image, thereby obtaining more effective face sculpting parameters.
[0091] For example, in the prior art, the difference between a real face image and a generated stylized face image is measured based on a discriminative neural network (discriminator) to optimize the face pinching parameters. The discriminative neural network is trained based on face images of the same style (real face style or game face style), and cannot well measure the style difference between the real face image and the stylized face image, resulting in an impact on the optimization process of the face pinching parameters and the effect of the generated face pinching parameters. In the embodiment of the present application, through a pre-trained transfer model, the target real face image is first transferred to the game style, and then the difference between the transferred stylized face image in the game style and the game face image generated by the first image parameters is compared. Based on the comparison result, the game face pinching parameters are optimized, which can ensure that both images required for comparison during the optimization process of the game face pinching parameters are game face images, and can avoid the problem that the discriminator cannot well measure the style difference between the cross-style real face image and the stylized face image, resulting in an impact on the optimization process of the face pinching parameters and the effect of the generated face pinching parameters. By using the image processing method provided in the embodiment of the present application, the problem of difficult parameter optimization caused by the inconsistency between the real face image and the game character style can be solved, and the game face pinching parameters can be estimated more effectively.
[0092] The above embodiment provides an image processing method. Correspondingly, another embodiment of the present application also provides an image processing device. The image processing device can be applied to the server of the game platform in a software or hardware manner. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the details of the relevant technical features, please refer to the corresponding description of the method embodiment provided above. The following description of the device embodiment is only illustrative.
[0093] Please refer to Figure 3 to understand this embodiment. Figure 3 is a unit block diagram of the image processing device provided in an embodiment of the present application. As Figure 3 shown, the image processing device provided in this embodiment includes:
[0094] A sample data set acquisition unit 301, configured to acquire a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and the corresponding stylized face images are images generated based on the same input vector;
[0095] A second stylized face image output unit 302, configured to acquire a first real face image from the sample data set, input the first real face image into a transfer model, and output a second stylized face image through the transfer model, where in the sample data set, the first real face image corresponds to a first stylized face image;
[0096] A migration model training unit 303, configured to train the migration model based on the first stylized face image and the second stylized face image to obtain a pre-trained migration model, where the pre-trained migration model is used to obtain a stylized face image.
[0097] The apparatus further includes:
[0098] A stylized face image acquisition unit, configured to acquire a stylized face image;
[0099] A second generation model obtaining unit, configured to train a pre-trained first generation model based on the stylized face image to obtain a second generation model, where the pre-trained first generation model is obtained by training based on real face images;
[0100] A vector input unit, configured to respectively generate a real face image and a stylized face image based on an input vector through the first generation model and the second generation model;
[0101] A data set obtaining unit, configured to obtain a sample data set based on the real face image and the stylized face image.
[0102] The apparatus further includes:
[0103] A migrated image obtaining unit, configured to input a target real face image into the pre-trained migration model to obtain a migrated image through the pre-trained migration model, where the migrated image is a stylized face image;
[0104] An initial image generation unit, configured to generate an initial image according to first image parameters;
[0105] A target image parameter obtaining unit, configured to update the first image parameters according to the migrated image and the initial image to obtain target image parameters.
[0106] The updating the first image parameters according to the migrated image and the initial image includes:
[0107] Extracting a first image feature from the migrated image and extracting a second image feature from the initial image;
[0108] Comparing the first image feature and the second image feature, and adjusting the first image parameters based on the comparison result to obtain target image parameters.
[0109] The adjusting the first image parameters based on the comparison result to obtain target image parameters includes:
[0110] In response to the similarity between the first image feature and the second image feature reaching a target similarity threshold, determine the first image parameter as the target parameter;
[0111] Alternatively, in response to the similarity between the first image feature and the second image feature not reaching the target similarity threshold, adjust the first image parameter to obtain a second image parameter; generate a second face image based on the second image parameter, extract a third image feature from the second face image, and obtain the similarity between the first image feature and the third image feature; in response to the similarity between the first image feature and the third image feature reaching the target similarity threshold, or the number of times of adjusting the first image parameter reaching a target number of times, obtain an updated first image parameter, and determine the updated first image parameter as the target image parameter.
[0112] The apparatus further includes: a target image generation unit, configured to generate a target image based on the target image parameter.
[0113] Training the pre-trained first generation model based on the stylized face image to obtain a second generation model includes: training the pre-trained first generation model based on the stylized face image to obtain a pre-trained third generation model; mixing the convolutional layer of the third generation model with the convolutional layer of the first generation model to obtain a second generation model.
[0114] Mixing the convolutional layer of the third generation model with the convolutional layer of the first generation model includes: based on the selected number of mixed layers, mixing the convolutional layer of the third generation model with the convolutional layer of the first generation model, where the number of mixed layers is used to characterize the stylization degree of the stylized face image to be generated by the second generation model.
[0115] By using the transfer model trained by the image processing apparatus provided in the embodiments of the present application, a stylized face image corresponding to the input target real face image can be obtained. Optimizing the face pinching parameters based on this stylized face image can ensure that the face image generated by the face pinching parameters to be optimized and the face image used to optimize the face pinching parameters are both stylized face images, for example, both game-style face images. Therefore, it is possible to avoid the problems that the optimization process of the face pinching parameters is affected and the effect of the generated face pinching parameters is affected because the face image used to optimize the face pinching parameters is a real face image and the face image generated by the face pinching parameters to be optimized is a stylized face image, thereby obtaining more effective face pinching parameters.
[0116] In the above embodiments, an image processing method and an image processing apparatus are provided. In addition, another embodiment of the present application further provides an electronic device. Since the embodiment of the electronic device is basically similar to the method embodiment, the description is relatively simple. For the details of the relevant technical features, please refer to the corresponding description of the method embodiment provided above. The following description of the embodiment of the electronic device is only illustrative. The embodiment of the electronic device is as follows:
[0117] Please refer to Figure 4 to understand this embodiment, Figure 4 which is a schematic logical structure diagram of an electronic device provided by an embodiment of the present application.
[0118] As Figure 4 shown, the electronic device provided in this embodiment includes: a processor 401 and a memory 402;
[0119] The memory 402 is used to store computer instructions for executing the image processing method provided in the above embodiment. When the computer instructions are read and executed by the processor 401, the following operations are performed:
[0120] Obtain a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and the corresponding stylized face images are images generated based on the same input vector;
[0121] Obtain a first real face image from the sample data set, input the first real face image into a migration model, and output a second stylized face image through the migration model. Among them, in the sample data set, the first real face image corresponds to a first stylized face image;
[0122] Based on the first stylized face image and the second stylized face image, train the migration model to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain stylized face images.
[0123] It further includes: obtaining a stylized face image; training a pre-trained first generation model based on the stylized face image to obtain a second generation model, where the pre-trained first generation model is obtained by training based on real face images; generating a real face image and a stylized face image respectively through the first generation model and the second generation model based on an input vector; and obtaining a sample data set based on the real face image and the stylized face image.
[0124] It further includes: inputting the target real face image into the pre-trained transfer model to obtain a transferred image through the pre-trained transfer model, where the transferred image is a stylized face image; generating an initial image according to the first image parameter; and updating the first image parameter according to the transferred image and the initial image to obtain a target image parameter.
[0125] The updating the first image parameter according to the transferred image and the initial image includes: extracting a first image feature from the transferred image and a second image feature from the initial image; comparing the first image feature and the second image feature, and adjusting the first image parameter based on the comparison result to obtain a target image parameter.
[0126] The adjusting the first image parameter based on the comparison result to obtain a target image parameter includes: in response to the similarity between the first image feature and the second image feature reaching a target similarity threshold, determining the first image parameter as the target parameter;
[0127] Or, in response to the similarity between the first image feature and the second image feature not reaching the target similarity threshold, adjusting the first image parameter to obtain a second image parameter; generating a second face image based on the second image parameter, extracting a third image feature from the second face image, and obtaining the similarity between the first image feature and the third image feature; in response to the similarity between the first image feature and the third image feature reaching the target similarity threshold, or the number of times of adjusting the first image parameter reaching a target number of times, obtaining the updated first image parameter, and determining the updated first image parameter as the target image parameter.
[0128] It further includes: generating a target image based on the target image parameter.
[0129] The training the pre-trained first generation model based on the stylized face image to obtain a second generation model includes: training the pre-trained first generation model based on the stylized face image to obtain a pre-trained third generation model; mixing the convolutional layer of the third generation model and the convolutional layer of the first generation model to obtain a second generation model.
[0130] The mixing the convolutional layer of the third generation model and the convolutional layer of the first generation model includes: mixing the convolutional layer of the third generation model and the convolutional layer of the first generation model based on the selected number of mixing layers, where the number of mixing layers is used to represent the stylization degree of the stylized face image to be generated by the second generation model.
[0131] By using the electronic device provided in this embodiment, a pre-trained transfer model for obtaining stylized face images can be obtained. By using this pre-trained transfer model, a stylized face image corresponding to a real face image can be obtained, realizing the transfer from a real face image to a stylized face image. When generating face pinching parameters, compared with directly using a real face image to optimize the face pinching parameters, using the stylized face image obtained by the transfer model trained by the above method to optimize the face pinching parameters can ensure that the face images generated by the face pinching parameters to be optimized and the face images used to optimize the face pinching parameters are both stylized face images, for example, both are face images in a game style. Therefore, it is possible to avoid the problems that the optimization process of the face pinching parameters is affected and the effect of the generated face pinching parameters is affected because the face image used to optimize the face pinching parameters is a real face image and the face image generated by the face pinching parameters to be optimized is a stylized face image, thereby obtaining more effective face pinching parameters.
[0132] In the above embodiment, an image processing method, an image processing device, and an electronic device are provided. In addition, another embodiment of the present application also provides a computer-readable storage medium for implementing the above image processing method. The embodiment of the computer-readable storage medium provided in the present application is described relatively simply. For the relevant parts, please refer to the corresponding description in the above method embodiment. The following described embodiments are only illustrative.
[0133] The computer-readable storage medium provided in this embodiment stores computer instructions, and when the instructions are executed by a processor, the following steps are implemented:
[0134] Obtain a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and the corresponding stylized face images are images generated based on the same input vector;
[0135] Obtain a first real face image from the sample data set, input the first real face image into the transfer model, and output a second stylized face image through the transfer model, where in the sample data set, the first real face image corresponds to a first stylized face image;
[0136] Based on the first stylized face image and the second stylized face image, train the transfer model to obtain a pre-trained transfer model, and the pre-trained transfer model is used to obtain stylized face images.
[0137] It further includes: obtaining a stylized face image; training a pre-trained first generation model based on the stylized face image to obtain a second generation model, wherein the pre-trained first generation model is obtained by training based on real face images; generating a real face image and a stylized face image respectively through the first generation model and the second generation model based on an input vector; and obtaining a sample data set based on the real face image and the stylized face image.
[0138] It further includes: inputting a target real face image into the pre-trained transfer model to obtain a transferred image through the pre-trained transfer model, where the transferred image is a stylized face image; generating an initial image according to first image parameters; and updating the first image parameters according to the transferred image and the initial image to obtain target image parameters.
[0139] The updating the first image parameters according to the transferred image and the initial image includes: extracting first image features from the transferred image and second image features from the initial image; comparing the first image features and the second image features, and adjusting the first image parameters based on the comparison result to obtain target image parameters.
[0140] The adjusting the first image parameters based on the comparison result to obtain target image parameters includes: in response to the similarity between the first image features and the second image features reaching a target similarity threshold, determining the first image parameters as the target parameters;
[0141] Or, in response to the similarity between the first image features and the second image features not reaching the target similarity threshold, adjusting the first image parameters to obtain second image parameters; generating a second face image based on the second image parameters, extracting third image features from the second face image, and obtaining the similarity between the first image features and the third image features; in response to the similarity between the first image features and the third image features reaching the target similarity threshold, or the number of times of adjusting the first image parameters reaching a target number of times, obtaining the updated first image parameters and determining the updated first image parameters as the target image parameters.
[0142] It further includes: generating a target image based on the target image parameters.
[0143] Training the pre-trained first generation model based on the stylized face image to obtain a second generation model includes: training the pre-trained first generation model based on the stylized face image to obtain a pre-trained third generation model; mixing the convolutional layers of the third generation model with the convolutional layers of the first generation model to obtain a second generation model.
[0144] The mixing of the convolutional layers of the third generation model with the convolutional layers of the first generation model includes: mixing the convolutional layers of the third generation model with the convolutional layers of the first generation model based on the selected number of mixing layers, where the number of mixing layers is used to characterize the stylization degree of the stylized face image to be generated by the second generation model.
[0145] By executing the computer instructions stored on the computer-readable storage medium provided in this embodiment, a pre-trained transfer model for obtaining stylized face images can be obtained. By using this pre-trained transfer model, a stylized face image corresponding to a real face image can be obtained, realizing the transfer from a real face image to a stylized face image. When generating face pinching parameters, compared with directly using a real face image to optimize the face pinching parameters, using the stylized face image obtained by the transfer model trained by the above method to optimize the face pinching parameters can ensure that the face images generated by the face pinching parameters to be optimized and the face images used to optimize the face pinching parameters are both stylized face images, such as face images in a game style. Therefore, it is possible to avoid the problems that the optimization process of the face pinching parameters is affected and the effect of the generated face pinching parameters is affected because the face image used to optimize the face pinching parameters is a real face image and the face image generated by the face pinching parameters to be optimized is a stylized face image, thereby obtaining more effective face pinching parameters.
[0146] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0147] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0148] 1. A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.
[0149] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0150] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined in the claims of the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain a sample data set, where the sample data set includes real face images and corresponding stylized face images, and the real face images and corresponding stylized face images are images generated based on the same input vector; Obtain a first real face image from the sample data set, input the first real face image into a migration model, and output a second stylized face image through the migration model, where in the sample data set, the first real face image corresponds to a first stylized face image; Based on the first stylized face image and the second stylized face image, train the migration model to obtain a pre-trained migration model, and the pre-trained migration model is used to obtain stylized face images; Input a target real face image into the pre-trained migration model, and obtain a migrated image through the pre-trained migration model, where the migrated image is a stylized face image; Generate an initial image according to first image parameters; Update the first image parameters according to the migrated image and the initial image to obtain target image parameters; Generate a target image based on the target image parameters.
2. The method according to claim 1, wherein The method further includes: Obtain stylized face images; Train a pre-trained first generation model based on the stylized face images to obtain a second generation model, where the pre-trained first generation model is trained based on real face images; Based on an input vector, generate a real face image and a stylized face image through the first generation model and the second generation model respectively; Obtain a sample data set based on the real face image and the stylized face image.
3. The method according to claim 1, wherein The updating the first image parameters according to the migrated image and the initial image includes: Extract a first image feature from the migrated image, and extract a second image feature from the initial image; Compare the first image feature and the second image feature, and adjust the first image parameters based on the comparison result to obtain target image parameters.
4. The method according to claim 3, characterized in that, The adjusting the first image parameters based on the comparison result to obtain target image parameters includes: In response to the similarity between the first image feature and the second image feature reaching a target similarity threshold, determine the first image parameters as the target image parameters; Or, in response to the similarity between the first image feature and the second image feature not reaching the target similarity threshold, adjust the first image parameters to obtain second image parameters; generate a second face image based on the second image parameters, extract a third image feature from the second face image, and obtain the similarity between the first image feature and the third image feature; in response to the similarity between the first image feature and the third image feature reaching the target similarity threshold, or the number of times of adjusting the first image parameters reaching a target number of times, obtain updated first image parameters, and determine the updated first image parameters as the target image parameters.
5. The method according to claim 2, wherein Training the pre-trained first generation model based on the stylized face image to obtain a second generation model, including: Training the pre-trained first generation model based on the stylized face image to obtain a pre-trained third generation model; Mixing the convolutional layer of the third generation model with the convolutional layer of the first generation model to obtain a second generation model.
6. The method according to claim 5, characterized in that, The mixing of the convolutional layer of the third generation model with the convolutional layer of the first generation model includes: Based on the selected number of mixing layers, mixing the convolutional layer of the third generation model with the convolutional layer of the first generation model, where the number of mixing layers is used to characterize the stylization degree of the stylized face image to be generated by the second generation model.
7. An image processing apparatus, characterized in that, The device includes: A sample dataset acquisition unit for acquiring a sample dataset, the sample dataset including real face images and corresponding stylized face images, the real face images and the corresponding stylized face images being images generated based on the same input vector; A second stylized face image output unit for obtaining a first real face image from the sample dataset, inputting the first real face image into the migration model, and outputting a second stylized face image through the migration model, where in the sample dataset the first real face image corresponds to a first stylized face image; A migration model training unit for training the migration model based on the first stylized face image and the second stylized face image to obtain a pre-trained migration model, the pre-trained migration model being used to obtain stylized face images; A migration image acquisition unit for inputting a target real face image into the pre-trained migration model and obtaining a migration image through the pre-trained migration model, the migration image being a stylized face image; An initial image generation unit for generating an initial image according to the first image parameters; A target image parameter acquisition unit for updating the first image parameters according to the migration image and the initial image to obtain target image parameters; A target image generation unit for generating a target image based on the target image parameters.
8. An electronic device, characterized in that, Including a processor and a memory; wherein, The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having one or more computer instructions stored thereon, characterized in that, The instruction is executed by the processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, storage medium and electronic device
CN109636886A
Image style migration method and device, electronic device and storage medium
CN109859096A
Model training method and device, information output method and device, equipment and storage medium
CN113052962A
Image generation method and device, electronic equipment and storage medium
CN113837934A