Image generation method, training method of style transfer model, and related devices

By obtaining the RGB image and feature images of the face image, input it into the style transfer model for style transfer, and iterative training using the loss function, the problem of inaccurate positioning of the style transfer model in the face area is solved, and the stability and naturalness of face style image generation is improved.

CN115019372BActive Publication Date: 2025-07-11BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210740473.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-07-11
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

In the prior art, when the style transfer model processes video in real time, the positioning of the face area is inaccurate, resulting in poor stability of the face style transfer.

Method used

By obtaining the RGB image of the face image and the feature images of the target area (such as eye, mouth, nose features), inputting them to the pre-trained style transfer model for style transfer, using the loss function to iterate the model to improve the accuracy of face area positioning.

Benefits of technology

The positioning accuracy of the style transfer model on the face area is improved, and the stability and nature of face style image generation are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019372B_ABST
    Figure CN115019372B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide an image generation method, a training method for a style transfer model, and related devices. Among them, the image generation method includes: obtaining an RGB image of a first face image; obtaining a feature image corresponding to a target area in the first face image; wherein the target area is used to represent the area where the target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature; inputting the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, and the second face image is a face image obtained by performing style transfer on the first face image. The method provided by the embodiments of the present invention can improve the stability of face style transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image generation method, a training method for a style transfer model, and related devices. Background Art

[0002] Portrait stylization is an advanced application of artificial intelligence (AI) shooting props. With the increasing usage of portrait stylization, higher requirements are put forward for the stylization stability during real-time video processing.

[0003] Currently, when performing real-time processing on a video, usually an image in the video is intercepted, and the corresponding Red Green Blue (RGB) image of the image is used as the input of the model, and is input into the style transfer model through 3 channels for portrait style transfer.

[0004] Since the input data is an RGB image, the accuracy of the style transfer model in locating the face region in the image during the style transfer process is poor, resulting in poor stability of face style transfer. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide an image generation method, a training method for a style transfer model, and related devices to improve the stability of face style transfer. The specific technical solutions are as follows:

[0006] In the first aspect implemented by the present invention, first, an image generation method is provided, including:

[0007] Obtain the Red Green Blue (RGB) image of the first face image;

[0008] Obtain the feature image corresponding to the target region in the first face image; wherein, the target region is used to represent the region where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features;

[0009] Input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is the face image obtained by performing style transfer on the first face image.

[0010] In the second aspect implemented by the present invention, a training method for a style transfer model is provided, including:

[0011] Obtain the RGB sample image of the sample face image;

[0012] Obtain a feature sample image corresponding to a target region in the sample face image; wherein, the target region is used to represent the region where the target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature;

[0013] Iteratively train the style transfer model to be trained by using the RGB sample image and the feature sample image;

[0014] Detect the result output by the style transfer model to be trained by using a loss function to determine a loss value;

[0015] If the change in the loss value is less than a preset value, determine the style transfer model to be trained currently as the style transfer model.

[0016] In a third aspect of the implementation of the present invention, an image generation device is provided, including:

[0017] A first acquisition module, configured to acquire a red, green, and blue (RGB) image of a first face image;

[0018] A second acquisition module, configured to acquire a feature image corresponding to a target region in the first face image; wherein, the target region is used to represent the region where the target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature;

[0019] An input module, configured to input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is a face image obtained by performing style transfer on the first face image.

[0020] In a fourth aspect of the implementation of the present invention, a training device for a style transfer model is provided, including:

[0021] A third acquisition module, configured to acquire an RGB sample image of a sample face image;

[0022] A fourth acquisition module, configured to acquire a feature sample image corresponding to a target region in the sample face image; wherein, the target region is used to represent the region where the target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature;

[0023] An iterative training module, configured to iteratively train the style transfer model to be trained by using the RGB sample image and the feature sample image;

[0024] A first determination module, configured to detect the result output by the style transfer model to be trained by using a loss function to determine a loss value;

[0025] A second determination module, configured to determine the to-be-trained style transfer model currently being trained as the style transfer model if the change in the loss value is less than a preset value.

[0026] In a fifth aspect of the embodiments of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0027] The memory is used to store programs;

[0028] The processor, when executing the programs stored on the memory, implements the method steps described in the first aspect or the second aspect.

[0029] In a sixth aspect of the embodiments of the present invention, a readable storage medium is provided, on which a program is stored. When the program is executed by a processor, the method described in the first aspect or the second aspect is implemented.

[0030] In the embodiments of the present invention, the image generation method specifically includes the following steps: obtaining a Red Green Blue (RGB) image of a first face image; and obtaining a feature image corresponding to a target region in the first face image, where the target region is used to represent the region where the target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature; inputting the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image. In this embodiment, the feature images corresponding to the obtained face feature region and other facial feature regions are input into the style transfer model, so that the style transfer model can locate the face region according to the feature images, improving the accuracy of the style transfer model in locating the face region, and further improving the stability of face style image generation. Description of the Drawings

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0032] Figure 1 One of the flow diagrams of an image generation method in the embodiments of the present invention;

[0033] Figure 2 Another flow diagram of an image generation method in the embodiments of the present invention;

[0034] Figure 3 A structural diagram of a style transfer model in the embodiments of the present invention;

[0035] Figure 4Schematic flowchart of a method for training a style transfer model in an embodiment of the present invention;

[0036] Figure 5 Schematic structural diagram of an image generation device in an embodiment of the present invention;

[0037] Figure 6 Schematic structural diagram of a training device for a style transfer model in an embodiment of the present invention;

[0038] Figure 7 It is a structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.

[0040] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an image generation method in an embodiment of the present invention. As Figure 1 shown, the embodiment of the present invention provides an image generation method, and the method specifically includes the following steps:

[0041] Step 101, obtain a Red Green Blue (RGB) image of the first face image.

[0042] Step 102, obtain a feature image corresponding to a target area in the first face image; wherein, the target area is used to represent the area where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features.

[0043] Step 103, input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is a face image obtained by performing style transfer on the first face image.

[0044] It should be understood that the RGB color model is an industrial color standard, and various colors are obtained by changing the three color channels of red (R), green (G), and blue (B) and their superposition with each other. In this embodiment, obtaining the RGB image of the first face image can be understood as obtaining the image in the RGB color model corresponding to the first face image.

[0045] It should be understood that the target region is used to represent the region where the target feature is located. For example, when the target feature includes a face feature, the target region is used to represent the region where the face feature is located, and the target region can also be understood as the face region. When the target feature includes a mouth feature, the target region is used to represent the region where the mouth feature is located, and the target region can also be understood as the mouth region.

[0046] It should be understood that each target region is used to represent the region where a target feature is located, and each target region corresponds to a feature image. In this embodiment, since the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature, the number of the feature images is at least two.

[0047] It should be noted that in specific implementation, the execution order of step 101 and step 102 is not limited herein. For example, in some embodiments, step 101 and step 102 can be executed simultaneously. In other embodiments, step 101 can be executed first, and then step 102. In still other embodiments, step 102 can be executed first, and then step 101.

[0048] It should be understood that the specific method for obtaining the feature image corresponding to the target region in the first face image is not limited herein. Optionally, in some embodiments, obtaining the feature image corresponding to the target region in the first face image specifically includes the following steps:

[0049] Obtain the key point set in the first face image, where the key point set includes coordinate points for representing the contour information of the target feature;

[0050] Generate the feature image corresponding to the target region based on the key point set of the target region.

[0051] It should be understood that the coordinate points for representing the contour information of the target feature can also be understood as face key points. The key point set includes a plurality of the face key points, and the plurality of face key points are used to represent the contour information of the target feature.

[0052] It should be understood that the number of the target regions is determined based on the number of the target features, the number of the key point sets is determined based on the number of the target regions, and the number of the feature images is determined based on the number of the key point sets, that is, each feature image corresponds to a target feature.

[0053] It should be understood that the specific method for obtaining the key point set in the first face image is not limited herein. For example, in some embodiments, multiple face key points can be extracted from the first face image through a face key point extraction algorithm. In specific implementation, the specific content of the face key point extraction algorithm is not limited herein.

[0054] It should be understood that the specific method for generating the feature image corresponding to the target area based on the key point set of the target area is not limited herein. For example, in some embodiments, after obtaining the coordinate points representing the contour information of the target feature, a polygon graph is obtained by connecting multiple coordinate points in a preset order, and the polygon graph is filled to obtain the feature image corresponding to the target area.

[0055] In the embodiment of the present invention, obtaining the feature image corresponding to the target area in the first face image specifically includes the following steps: obtaining the key point set in the first face image, where the key point set includes the coordinate points representing the contour information of the target feature; generating the feature image corresponding to the target area based on the key point set of the target area. Through the above settings, the feature image corresponding to the target area can be generated based on the face key points in the first face image, improving the accuracy of representing the target feature by the feature image corresponding to the target area, and further improving the accuracy of positioning the target feature.

[0056] In the embodiment of the present invention, the image generation method specifically includes the following steps: obtaining the Red Green Blue (RGB) image of the first face image; and obtaining the feature image corresponding to the target area in the first face image, where the target area is used to represent the area where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features; inputting the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image. In this embodiment, the feature images corresponding to the obtained face feature area and other facial feature areas are input into the style transfer model, so that the style transfer model can locate the face area based on the feature images, improving the accuracy of the style transfer model in locating the face area, and further improving the stability of face style image generation.

[0057] Optionally, as Figure 2 shown, in some embodiments, the image generation method further includes the following steps:

[0058] Performing face segmentation on the first face image to determine a first area and a second area in the first face image, where the first area is used to represent the area where the portrait feature is located, and the second area is used to represent the area where the background feature is located;

[0059] Perform color adjustment on the first region to obtain a third face image, where the third face image is a face image obtained by performing color adjustment on the first region of the first face image;

[0060] Composite and render the third face image and the second face image to obtain a fourth face image.

[0061] It should be understood that by performing face segmentation on the first face image, the first region and the second region in the first face image can be determined, and in subsequent operations, the first region or the second region can be operated on separately.

[0062] It should be understood that the specific method for performing face segmentation on the first face image is not limited herein. For example, in some embodiments, performing face segmentation on the first face image can obtain a mask image. In the mask image, the parameters corresponding to the first region and the second region are different, and through the parameters corresponding to the first region and the second region in the mask image, the first region or the second region can be operated on separately.

[0063] It should be understood that the first region for representing the region where the portrait features are located can be understood as the first region for representing the face features and the body features, where the body features can be understood as the neck feature and at least one of the following: shoulder feature, arm feature, and torso feature.

[0064] It should be understood that performing color adjustment on the first region can be understood as adjusting the overall color of the first region so that the color style of the first region matches the style of the second face image. The first region of the third face image is the first region after color adjustment, and the second region of the third face image is the second region that has not been adjusted.

[0065] It should be understood that composite and rendering the third face image and the second face image to obtain a fourth face image can be understood as performing color adjustment on the portrait parts of the third face image and the second face image, face anti-correction, and fusion of the original images.

[0066] For ease of understanding, an example will be given below. In this embodiment, performing face segmentation on the first face image obtains a mask image, where the obtained mask image is a black and white image, the first region is white, the parameter is 1, the second region is black, and the parameter is 0. Perform color adjustment on the region with parameter 1, and keep the region with parameter 0 unchanged to obtain a third face image.

[0067] Since face style transfer mainly targets face images, the color difference between the face image obtained after style transfer and the body parts such as the neck in the first face image is usually large, making the face frame more obvious, and thus resulting in poor naturalness of the second face image.

[0068] In an embodiment of the present invention, the image generation method further includes the following steps: performing face segmentation on the first face image to determine a first region and a second region in the first face image, performing color adjustment on the first region to obtain a third face image, and synthesizing and rendering the third face image and the second face image to obtain a fourth face image. Through the above settings, by adjusting the color of the first region, the skin color of the person's neck and other positions can be matched with the skin color of the face in the second face image, thereby making the second face image more natural.

[0069] Optionally, in some embodiments, the style transfer model includes:

[0070] A downsampling module, which is used to downsample and extract features from the RGB image and the feature image to obtain a first feature map;

[0071] A first residual module, which is used to extract features and perform a first normalization process on the first feature map to obtain a second feature map;

[0072] A second residual module, which is used to perform a stylization process and a second normalization process on the second feature map to obtain a third feature map;

[0073] An upsampling module, which is used to upsample and extract features from the third feature map, and obtain the second face image based on the result obtained by upsampling and extracting features from the third feature map.

[0074] It should be noted that in this embodiment, the first feature map output by the downsampling module is usually input to the first residual module through an activation function (Activation Function). The activation function can be understood as a function that runs on the neurons of a neural network and is responsible for mapping the input of the neuron to the output end. Therefore, in some embodiments, it can also be considered that there is an activation function layer between the downsampling module and the first residual module.

[0075] Similarly, the second feature map output by the first residual module is usually input to the second residual module through an activation function, and the third feature map output by the second residual module is usually input to the upsampling module through an activation function. Among them, the activation functions set between every two modules can be the same or different, and the specific type of activation function is not limited here.

[0076] It should be understood that the downsampling module is used to downsample and extract features from the RGB image and the feature image to obtain a first feature map, where the specific structure of the downsampling module is not limited herein.

[0077] For example, in some embodiments, the downsampling module may include a convolutional layer and a normalization layer connected in sequence. In other embodiments, the downsampling module may include a plurality of convolutional layers and a normalization layer connected in sequence.

[0078] Optionally, in some embodiments, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer, and the downsampling sub-module includes:

[0079] A first convolutional layer, which is used to downsample and extract features from the input first image to obtain a first sub-feature map;

[0080] A first normalization layer, which is used to normalize the first sub-feature map to obtain an output first result;

[0081] Wherein, the first image corresponding to the first downsampling sub-module is the RGB image and the feature image, and the corresponding first result is a second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is a third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

[0082] It should be understood that the first convolutional layer is used to downsample and extract features from the input first image and output a first sub-feature map. Wherein, the specific method for the first convolutional layer to downsample and extract features from the input first image is not limited herein. In some embodiments, downsampling can also be understood as subsampling.

[0083] For example, in some embodiments, the first convolutional layer includes a downsampling sub-layer and a convolutional sub-layer. The first downsampling sub-layer is used to downsample the input first image to obtain a downsampled feature map, and the convolutional sub-layer is used to extract features from the downsampled feature map to obtain a first sub-feature map.

[0084] In some other embodiments, the value of the stride of the first convolutional layer is a positive integer greater than 1. Herein, the stride can be understood as the number of grids by which the feature map slides after one convolution calculation. Since the value of the stride is a positive integer greater than 1, the first convolutional layer performs downsampling while performing convolution calculation on the input first image.

[0085] It should be understood that the first normalization layer is used to perform normalization processing on the first sub-feature map to obtain the output first result. Herein, the specific method for the first normalization layer to perform normalization processing on the first sub-feature map is not limited herein.

[0086] For example, in some embodiments, the normalization processing may be instance normalization (IN). In some other embodiments, the normalization processing may be batch normalization (BN) or layer normalization (LN).

[0087] It should be understood that in some embodiments, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. In this embodiment, the number of downsampling sub-modules included in the downsampling module is N + 2, that is, the number of downsampling sub-modules is at least 3.

[0088] Of course, in specific implementation, the downsampling module may also include one downsampling sub-module, or two downsampling sub-modules. The specific structure thereof may refer to the description of this embodiment and will not be elaborated herein. For example, in the case where the downsampling module includes one downsampling sub-module, the input of the downsampling sub-module is the RGB image and the feature image, and the output is the first feature map.

[0089] It should be understood that each downsampling sub-module includes a first convolutional layer and a first normalization layer. For different downsampling sub-modules, the first image corresponding to the input of its first convolutional layer is different, and the first result corresponding to the output of its second convolutional layer is different.

[0090] For the first downsampling sub-module, the first image input to this downsampling sub-module is the RGB image and the feature image, and the first result output by this downsampling sub-module is the second sub-feature map.

[0091] For any one of the intermediate downsampling sub-modules, the first image input to the downsampling sub-module is the first result output by the previous downsampling sub-module, and the first result output by the downsampling sub-module is the third sub-feature map.

[0092] It should be understood that the third sub-feature map is used to generally refer to the result output by the first normalization layer of the intermediate downsampling sub-module. In specific implementation, the content of the third sub-feature map corresponding to each intermediate downsampling sub-module is different.

[0093] For the last downsampling sub-module, the first image input to the downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the first result output by the downsampling sub-module is the first feature map.

[0094] In the embodiments of the present invention, the feature maps output by each downsampling sub-module are all different. Through the setting of the first downsampling sub-module, N intermediate downsampling sub-modules, and the last downsampling sub-module connected in sequence, the first face image is downsampled and feature-extracted multiple times. On the one hand, it improves the depth of feature extraction of the first face image, and on the other hand, it also reduces the size of the first face image, improving the speed and efficiency of model calculation.

[0095] It should be understood that the first residual module is used to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map. Wherein, the specific structure of the first residual module is not limited herein.

[0096] For example, in some embodiments, the first residual module includes a convolutional layer and a normalization layer. The convolutional layer is used to perform feature extraction on the first feature map and input the output result to the normalization layer. The normalization layer is used to perform normalization processing on the result output by the convolutional layer to obtain a calculation result, and add the calculation result to the input of the convolutional layer and then output it to the second residual module.

[0097] It should be understood that the specific structure of the normalization layer corresponding to the first residual module is not limited herein, that is, the specific method of the first normalization processing performed by the first residual module is not limited herein. For example, in some embodiments, the first normalization processing is IN.

[0098] It should be understood that the number of the first residual modules is not limited herein. For example, in some embodiments, the number of the first residual modules is at least one, and at least one of the first residual modules are connected in sequence. For example, in some embodiments, the number of the first residual modules is four, and the four first residual modules are connected in sequence.

[0099] It should be understood that the second residual module is used to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map. The specific structure of the second residual module is not limited herein.

[0100] For example, in some embodiments, the second residual module includes a convolutional layer and a normalization layer. The convolutional layer is used to extract features from the third feature map and input the output result into the normalization layer. The normalization layer is used to perform normalization processing on the result output by the convolutional layer to obtain a calculation result, and add the calculation result to the input of the convolutional layer and then output it to the upsampling module.

[0101] It should be understood that the specific structure of the normalization layer corresponding to the second residual module is not limited herein, that is, the specific method of the second normalization processing performed by the second residual module is not limited herein. For example, in some embodiments, the second normalization processing is Adaptive-Instance-Layer-Norm.

[0102] It should be understood that the number of the second residual modules is not limited herein. For example, in some embodiments, the number of the second residual modules is at least one, and at least one of the second residual modules are connected in sequence. For example, in some embodiments, the number of the second residual modules is four, and the four second residual modules are connected in sequence.

[0103] It should be understood that the upsampling module is used to perform upsampling and feature extraction on the third feature map, and obtain the second face image based on the result obtained by performing upsampling and feature extraction on the third feature map. The specific structure of the upsampling module is not limited herein.

[0104] For example, in some embodiments, the upsampling module may include a convolutional layer and a normalization layer connected in sequence. In other embodiments, the upsampling module may include a plurality of convolutional layers and a normalization layer connected in sequence.

[0105] It should be understood that the upsampling module obtains the second face image based on the result obtained by performing upsampling and feature extraction on the third feature map can be understood as that, in some embodiments, the upsampling module may output the result obtained by performing upsampling and feature extraction on the third feature map as the second face image.

[0106] Optionally, in some embodiments, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. The upsampling sub-module includes:

[0107] An upsampling layer for upsampling the input second image to obtain a fourth sub-feature map;

[0108] A second convolutional layer for extracting features from the fourth sub-feature map to obtain a fifth sub-feature map;

[0109] A second normalization layer for normalizing the fifth sub-feature map to obtain an output second result;

[0110] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the second face image.

[0111] It should be understood that the upsampling layer is used to upsample the input second image to obtain a fourth sub-feature map. In some embodiments, upsampling can also be understood as resampling (upsample). The second convolutional layer is used to extract features from the fourth sub-feature map to obtain a fifth sub-feature map.

[0112] In this embodiment, the first convolutional layer and the second convolutional layer are only used to distinguish different components. In specific implementation, the specific structures of the first convolutional layer, the second convolutional layer, and the convolutional layer mentioned in the embodiments of the present invention may be the same or different.

[0113] It should be understood that the second normalization layer is used to normalize the fifth sub-feature map to obtain an output second result, where the specific method for the second normalization layer to normalize the second sub-feature map is not limited herein.

[0114] For example, in some embodiments, the normalization process can be IN. In other embodiments, the normalization process can be BN or group normalization (Group Normalization, GN).

[0115] It should be understood that in some embodiments, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. In this embodiment, the number of upsampling sub-modules included in the upsampling module is M + 2, that is, the number of upsampling sub-modules is at least 3.

[0116] Of course, in specific implementation, the upsampling module may also include one upsampling sub-module or two upsampling sub-modules. Its specific structure may refer to the description of this embodiment and will not be elaborated here. For example, when the upsampling module includes one upsampling sub-module, the input of the upsampling sub-module is the first feature map and the output is the second face image.

[0117] It should be understood that downsampling and upsampling of the feature map can be regarded as opposite operations. In specific implementation, the number of upsampling sub-modules is usually the same as that of the downsampling sub-modules, and the size of the second face image output by the upsampling module is the same as that of the first face image input to the downsampling module.

[0118] It should be understood that each upsampling sub-module includes a second convolutional layer and a second normalization layer. For different upsampling sub-modules, the second image corresponding to the input of its second convolutional layer is different, and the second result corresponding to the output of its second convolutional layer is different.

[0119] For the first upsampling sub-module, the second image input to this upsampling sub-module is the third feature map, and the second result output by this upsampling sub-module is the sixth sub-feature map.

[0120] For any intermediate upsampling sub-module, the second image input to this upsampling sub-module is the second result output by the previous upsampling sub-module, and the second result output by this upsampling sub-module is the seventh sub-feature map.

[0121] It should be understood that the seventh sub-feature map is used to generally refer to the result output by the second normalization layer of the intermediate upsampling sub-module. In specific implementation, the content of the seventh sub-feature map corresponding to each intermediate upsampling sub-module is different.

[0122] For the last upsampling sub-module, the second image input to this upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the second result output by this upsampling sub-module is the second face image.

[0123] In the embodiment of the present invention, the feature maps output by each upsampling sub-module are different. Through the setting of the first upsampling sub-module, N intermediate upsampling sub-modules and the last upsampling sub-module connected in sequence, multiple upsamplings and feature extractions are performed on the second face image. On the one hand, the depth of feature extraction of the second face image is increased, and on the other hand, the influence of style transfer on the image size is reduced, and the quality of the second face image is improved.

[0124] Optionally, in some embodiments, the style transfer model further includes a branch convolutional layer; the branch convolutional layer is configured to downsample and extract features from the RGB image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0125] The upsampling module includes a first sub-module and a second sub-module. The first sub-module is configured to upsample and extract features from the third feature map, and the second sub-module is configured to add the output result of the first sub-module to the fourth feature map to obtain the second face image.

[0126] It should be understood that the style transfer model further includes a branch convolutional layer, and the input of the branch convolutional layer is the RGB image. The branch convolutional layer extracts features from the RGB image, and can obtain the features of the first face image at the overall level.

[0127] It should be understood that, in some embodiments, the input of the branch convolutional layer can also be the RGB image and the feature image.

[0128] It should be understood that the downsampling of the RGB image by the branch convolutional layer can be understood as that the branch convolutional layer downsamples the RGB image so that the size of the fourth feature map is the same as the size of the output result of the first sub-module.

[0129] In the embodiments of the present invention, the style transfer model further includes a branch convolutional layer. The branch convolutional layer is configured to downsample and extract features from the RGB image to obtain a fourth feature map. The upsampling module adds the result of upsampling and extracting features from the third feature map to the fourth feature map to obtain the second face image. The branch convolutional layer can extract the overall features of the first face image, thereby improving the clarity of the second face image and the perfection of the detailed features.

[0130] Optionally, in some embodiments, the style transfer model further includes a parameter calculation layer. The parameter calculation layer is configured to determine target parameters based on the first feature map, and input the target parameters into the second residual module;

[0131] Wherein, the second normalization process includes performing normalization processing based on the target parameters.

[0132] First of all, it should be noted that during the normalization process, affine parameters are usually required for calculation. Specifically, the affine parameters for normalization processing can be understood to include a first parameter and a second parameter. In some embodiments, the first parameter can be understood as a gamma parameter, and the second parameter can be understood as a beta parameter.

[0133] For IN, during the normalization process, the normalization layer usually needs to learn the first parameter and the second parameter through backpropagation. In this embodiment, the target parameters include the first parameter and the second parameter. Therefore, the second normalization process does not need to learn the first parameter and the second parameter and can be adaptive.

[0134] In the present embodiment of the invention, the style transfer model further includes a parameter calculation layer for determining target parameters based on the first feature map, and the second normalization process includes performing normalization processing based on the target parameters. Through the second normalization process, style transfer can be achieved by changing the data distribution of features at the feature map level, with both the computational cost and the storage cost being relatively small and easy to implement.

[0135] For ease of understanding, the following will use the style transfer model as shown in Figure 2 to illustrate the structure of the style transfer model provided by the embodiments of the present invention.

[0136] Please refer to Figure 3 , the style transfer model includes a downsampling module 301, a first residual module 302, a second residual module 303, and an upsampling module 304 connected in sequence. Among them, the input of the downsampling module 301 is the RGB image and the feature image corresponding to the first face image, and the output of the upsampling module 304 is the second face image.

[0137] Among them, the downsampling module 301 includes a first downsampling sub-module, an intermediate downsampling sub-module, and a last downsampling sub-module. The first residual module 302 includes a first first residual sub-module, two intermediate first residual sub-modules, and a last first residual sub-module. The second residual module 303 includes a first second residual sub-module, two intermediate second residual sub-modules, and a last second residual sub-module. The upsampling module 304 includes a first upsampling sub-module, an intermediate upsampling sub-module, and a last upsampling sub-module.

[0138] The style transfer model further includes a branch convolution layer 305. The input of the branch convolution layer 305 is the RGB image corresponding to the first face image, and the output result of the branch convolution layer 305 is input into the intermediate upsampling sub-module. It should be understood that in specific implementation, the output result of the branch convolution layer 305 can be input into any upsampling sub-module.

[0139] In this embodiment, inputting the output result of the branch convolution layer 305 into the intermediate upsampling sub-module can not only ensure adding certain detail features to the output second face image but also reduce the impact on the style transfer effect.

[0140] The style transfer model further includes a parameter calculation layer 306. The input of the parameter calculation layer 306 is the first feature map output by the last downsampling sub-module, and the output result of the parameter calculation layer 306 is input into each second residual sub-module.

[0141] It should be understood that the specific functions of each network layer in this embodiment can be referred to the description of the foregoing embodiments, and will not be elaborated herein.

[0142] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a training method for a style transfer model in an embodiment of the present invention. As Figure 4 shown, the embodiment of the present invention provides a training method for a style transfer model, and the method specifically includes the following steps:

[0143] Step 401, obtaining an RGB sample image of a sample face image.

[0144] Step 402, obtaining a feature sample image corresponding to a target area in the sample face image; wherein, the target area is used to represent an area where a target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature.

[0145] Step 403, iteratively training the style transfer model to be trained by using the RGB sample image and the feature sample image.

[0146] Step 404, detecting the result output by the style transfer model to be trained by using a loss function to determine a loss value.

[0147] Step 405, if the change in the loss value is less than a preset value, determining the style transfer model to be trained currently as the style transfer model.

[0148] It should be noted that in specific implementation, the execution order of the step 401 and the step 402 is not limited herein. For example, in some embodiments, the step 401 and the step 402 can be executed simultaneously. In other embodiments, the step 401 can be executed first, and then the step 402. In other embodiments, the step 402 can be executed first, and then the step 401.

[0149] Optionally, the style transfer model to be trained includes:

[0150] A downsampling module, which is used to perform downsampling and feature extraction on the RGB sample image and the feature sample image to obtain a first feature map;

[0151] The first residual module is configured to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map;

[0152] The second residual module is configured to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map;

[0153] An upsampling module; the upsampling module is configured to perform upsampling and feature extraction on the third feature map, and obtain a stylized face image based on the result obtained by performing upsampling and feature extraction on the third feature map, where the stylized face image is a face image obtained by performing style transfer on the sample face image.

[0154] It should be understood that in this embodiment, obtaining the RGB sample image of the sample face image can be understood as obtaining the image in the RGB color mode corresponding to the sample face image. The description of step 401 can refer to the description of step 101 in the foregoing embodiment. To avoid repetition, it will not be elaborated here. The description of step 402 can refer to the description of step 102 in the foregoing embodiment. To avoid repetition, it will not be elaborated here.

[0155] It should be understood that in this embodiment, the loss function may include at least one loss function. Using the loss function to detect the result output by the style transfer model to be trained and determine the loss value can be understood as using at least one loss function to determine at least one loss sub-value, and then determining the loss value based on the at least one loss sub-value.

[0156] It should be understood that in some embodiments, the loss function includes at least one of the following: mean squared error (MSELoss), perceptual loss (Learned Perceptual Image Patch Similarity, LPIPS Loss), mocoLoss, and loss sensitive generative adversarial network (loss sensitive Generative Adversarial Network, LS-GAN Loss).

[0157] In this embodiment, MSE Loss is a basic regression loss, LPIPS Loss and moco Loss are a kind of perceptual loss, which can be used to improve the clarity of details. LS-GAN Loss can further capture the differences in details in the generated image, thereby improving the detail effect of the generated image.

[0158] It should be understood that the specific method for determining the loss value based on at least one loss sub-value is not limited herein. For example, in some embodiments, the average value of the at least one loss sub-value may be determined as the loss value. In other embodiments, the variance of the at least one loss sub-value may be determined as the loss value.

[0159] It should be understood that the value of the preset value is not limited herein. In specific implementation, the value of the preset value can be set and adjusted according to actual requirements.

[0160] It should be understood that in some embodiments, multiple data augmentation methods may be used to preprocess the sample face image. Among them, the data augmentation method is not limited herein. For example, in some embodiments, the data augmentation method may include at least one of the following: image transformation, color transformation, and other random processing.

[0161] In this embodiment, the image transformation may include at least one of the following: random flipping of the image, image translation, image cropping, and image scaling. The color transformation may include at least one of the following: image brightness adjustment, image saturation adjustment, and image contrast adjustment. The random processing may include adding random noise and adding random background filling for the black edges generated by data transformation.

[0162] In this embodiment, multiple data augmentation methods are used to preprocess the sample face image. By adding random noise and random black edge background filling, the generalization ability of the model for reconstructing the background part can be further improved, and problems such as white fog generated outside the face during real-time processing can be reduced.

[0163] Optionally, the style transfer model to be trained further includes a branch convolutional layer; the branch convolutional layer is used to downsample and extract features from the RGB sample image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0164] The upsampling module includes a first sub-module and a second sub-module. The first sub-module is used to upsample and extract features from the third feature map, and the second sub-module is used to add the output result of the first sub-module to the fourth feature map to obtain the stylized face image.

[0165] Optionally, the style transfer model to be trained further includes a parameter calculation layer, and the parameter calculation layer is used to determine target parameters based on the first feature map and input the target parameters into the second residual module;

[0166] Among them, the second normalization process includes normalization based on the target parameters.

[0167] Optionally, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. Among them, the downsampling sub-module includes:

[0168] A first convolutional layer, which is used to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map;

[0169] A first normalization layer, which is used to perform normalization processing on the first sub-feature map to obtain an output first result;

[0170] Among them, the first image corresponding to the first downsampling sub-module is the RGB sample image and the feature sample image, and the corresponding first result is the second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is the third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

[0171] Optionally, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes:

[0172] An upsampling layer, which is used to perform upsampling on the input second image to obtain a fourth sub-feature map;

[0173] A second convolutional layer, which is used to perform feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map;

[0174] A second normalization layer, which is used to perform normalization processing on the fifth sub-feature map to obtain an output second result;

[0175] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the stylized face image.

[0176] It should be noted that the style transfer model trained by the training method of the style transfer model provided in the embodiments of the present invention is applied to the image generation method provided in the above embodiments. The structure of the style transfer model to be trained can refer to the relevant description in the above embodiments. To avoid repetition, it will not be elaborated here.

[0177] It should be understood that the training samples used for iterative training of the style transfer model to be trained in the embodiments of the present invention are the RGB sample images and the feature sample images. Therefore, the specific process of the style transfer model to be trained for processing the RGB sample images and the feature sample images can refer to the relevant process of the style transfer model for processing RGB images and feature images in the above embodiments. To avoid repetition, it will not be elaborated here.

[0178] The embodiments of the present invention also provide an image generation device. Refer to Figure 5 , Figure 5 which is the structure diagram of the image generation device 500 provided in the embodiments of the present invention. Since the principle of the image generation device 500 for solving problems is similar to that of the Figure 1 image generation method shown in the embodiments of the present invention, the implementation of the image generation device 500 can refer to the implementation of the method. The repeated parts will not be elaborated.

[0179] The embodiments of the present invention also provide an image generation device 500, including:

[0180] A first acquisition module 501, configured to acquire a red, green, and blue (RGB) image of a first face image;

[0181] A second acquisition module 502, configured to acquire a feature image corresponding to a target region in the first face image; wherein, the target region is used to represent a region where target features are located, and the target features include face features and at least one of the following: eye features, mouth features, and nose features;

[0182] An input module 503, configured to input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is a face image obtained by performing style transfer on the first face image.

[0183] Optionally, the acquiring the feature image corresponding to the target region in the first face image includes:

[0184] acquiring a key point set in the first face image, where the key point set includes coordinate points for representing the contour information of the target features;

[0185] generating the feature image corresponding to the target region based on the key point set of the target region.

[0186] Optionally, the image generation device 500 further includes:

[0187] A third determination module, configured to perform face segmentation on the first face image, and determine a first region and a second region in the first face image, where the first region is used to represent the region where portrait features are located, and the second region is used to represent the region where background features are located;

[0188] A color adjustment module, configured to perform color adjustment on the first region to obtain a third face image, where the third face image is a face image obtained by performing color adjustment on the first region of the first face image;

[0189] A synthesis and rendering module, configured to synthesize and render the third face image and the second face image to obtain a fourth face image.

[0190] Optionally, the style transfer model includes:

[0191] A downsampling module, configured to perform downsampling and feature extraction on the RGB image and the feature image to obtain a first feature map;

[0192] A first residual module, configured to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map;

[0193] A second residual module, configured to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map;

[0194] An upsampling module, configured to perform upsampling and feature extraction on the third feature map, and obtain the second face image based on the result obtained by performing upsampling and feature extraction on the third feature map.

[0195] Optionally, the style transfer model further includes a branch convolutional layer; the branch convolutional layer is configured to perform downsampling and feature extraction on the RGB image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0196] The upsampling module includes a first sub-module and a second sub-module, where the first sub-module is configured to perform upsampling and feature extraction on the third feature map, and the second sub-module is configured to add the output result of the first sub-module and the fourth feature map to obtain the second face image.

[0197] Optionally, the style transfer model further includes a parameter calculation layer, configured to determine target parameters based on the first feature map, and input the target parameters into the second residual module;

[0198] Among them, the second normalization process includes performing a normalization process based on the target parameter.

[0199] Optionally, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. Among them, the downsampling sub-module includes:

[0200] A first convolutional layer, which is used to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map;

[0201] A first normalization layer, which is used to perform a normalization process on the first sub-feature map to obtain an output first result;

[0202] Among them, the first image corresponding to the first downsampling sub-module is the RGB image and the feature image, and the corresponding first result is the second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is the third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

[0203] Optionally, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes:

[0204] An upsampling layer, which is used to perform upsampling on the input second image to obtain a fourth sub-feature map;

[0205] A second convolutional layer, which is used to perform feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map;

[0206] A second normalization layer, which is used to perform a normalization process on the fifth sub-feature map to obtain an output second result;

[0207] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the second face image.

[0208] The image generation device 500 provided in the embodiments of the present invention can implement Figure 1 each process implemented by the method embodiments shown in

[0209] and can achieve the same beneficial effects. To avoid repetition, details are not described herein again. Figure 6 , Figure 6 FIG. is a structural diagram of a training device 600 for a style transfer model provided in the embodiments of the present invention. Since the principle of the training device 600 for the style transfer model to solve problems is similar to that of the Figure 4 style transfer model training method shown in the embodiments of the present invention, the implementation of the training device 600 for the style transfer model can refer to the implementation of the method, and repeated parts are not described again.

[0210] The embodiments of the present invention further provide a training device 600 for a style transfer model, including:

[0211] A third acquisition module 601, configured to acquire an RGB sample image of a sample face image;

[0212] A fourth acquisition module 602, configured to acquire a feature sample image corresponding to a target region in the sample face image; wherein, the target region is used to represent a region where a target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature;

[0213] An iterative training module 603, configured to iteratively train a style transfer model to be trained by using the RGB sample image and the feature sample image;

[0214] A first determination module 604, configured to detect the result output by the style transfer model to be trained by using a loss function, and determine a loss value;

[0215] A second determination module 605, configured to, if the change in the loss value is less than a preset value, determine the style transfer model to be trained currently as the style transfer model.

[0216] Optionally, the style transfer model to be trained includes:

[0217] A downsampling module, configured to perform downsampling and feature extraction on the RGB sample image and the feature sample image to obtain a first feature map;

[0218] A first residual module, configured to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map;

[0219] A second residual module, which is used to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map;

[0220] An upsampling module; the upsampling module is used to perform upsampling and feature extraction on the third feature map, and obtain a stylized face image based on the result obtained by performing upsampling and feature extraction on the third feature map. The stylized face image is a face image obtained by performing style transfer on the sample face image.

[0221] Optionally, the style transfer model to be trained further includes a branch convolutional layer; the branch convolutional layer is used to perform downsampling and feature extraction on the RGB sample image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0222] The upsampling module includes a first sub-module and a second sub-module. The first sub-module is used to perform upsampling and feature extraction on the third feature map, and the second sub-module is used to add the output result of the first sub-module to the fourth feature map to obtain the stylized face image.

[0223] Optionally, the style transfer model to be trained further includes a parameter calculation layer, which is used to determine target parameters based on the first feature map and input the target parameters into the second residual module;

[0224] Wherein, the second normalization processing includes performing normalization processing based on the target parameters.

[0225] Optionally, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules and a last downsampling sub-module connected in sequence, where N is a positive integer. The downsampling sub-module includes:

[0226] A first convolutional layer, which is used to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map;

[0227] A first normalization layer, which is used to perform normalization processing on the first sub-feature map to obtain an output first result;

[0228] Wherein, the first image corresponding to the first downsampling sub-module is the RGB sample image and the feature sample image, and the corresponding first result is a second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is a third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

[0229] Optionally, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes:

[0230] An upsampling layer for upsampling the input second image to obtain a fourth sub-feature map;

[0231] A second convolutional layer for feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map;

[0232] A second normalization layer for normalizing the fifth sub-feature map to obtain an output second result;

[0233] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the stylized face image.

[0234] The training device 600 of the style transfer model provided by the embodiments of the present invention can implement Figure 4 each process implemented by the method embodiments shown, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.

[0235] The embodiments of the present invention further provide an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704,

[0236] The memory 703 is used to store a computer program;

[0237] The processor 701, when executing the program stored in the memory 703, implements the following steps:

[0238] Obtain the red, green, and blue RGB images of the first face image;

[0239] Obtain the feature image corresponding to the target area in the first face image; where the target area is used to represent the area where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features;

[0240] Input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is the face image obtained by performing style transfer on the first face image;

[0241] Or,

[0242] When the processor 701 executes the program stored in the memory 703, the following steps are implemented:

[0243] Obtain the RGB sample image of the sample face image;

[0244] Obtain the feature sample image corresponding to the target area in the sample face image; wherein, the target area is used to represent the area where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features;

[0245] Use the RGB sample image and the feature sample image to iteratively train the style transfer model to be trained;

[0246] Use a loss function to detect the result output by the style transfer model to be trained and determine the loss value;

[0247] If the change in the loss value is less than a preset value, determine the currently trained style transfer model to be trained as the style transfer model.

[0248] Optionally, the style transfer model to be trained includes:

[0249] A downsampling module, which is used to downsample and extract features from the RGB sample image and the feature sample image to obtain a first feature map;

[0250] A first residual module, which is used to extract features from the first feature map and perform a first normalization process to obtain a second feature map;

[0251] A second residual module, which is used to perform a stylization process and a second normalization process on the second feature map to obtain a third feature map;

[0252] An upsampling module; the upsampling module is used to upsample and extract features from the third feature map, and obtain a stylized face image based on the result obtained by upsampling and extracting features from the third feature map, where the stylized face image is the face image obtained by performing style transfer on the sample face image.

[0253] Optionally, when the processor 701 further executes the program stored in the memory 703, the following steps are implemented:

[0254] Obtain a set of key points in the first face image, where the set of key points includes coordinate points for representing the contour information of the target feature;

[0255] Generate a feature image corresponding to the target region based on the set of key points of the target region.

[0256] Optionally, when the processor 701 is further used to execute the program stored on the memory 703, the following steps are implemented:

[0257] Perform face segmentation on the first face image to determine a first region and a second region in the first face image, where the first region is used to represent the region where the portrait feature is located, and the second region is used to represent the region where the background feature is located;

[0258] Perform color adjustment on the first region to obtain a third face image, where the third face image is a face image obtained by performing color adjustment on the first region of the first face image;

[0259] Synthesize and render the third face image and the second face image to obtain a fourth face image.

[0260] Optionally, the style transfer model includes:

[0261] A downsampling module, which is used to perform downsampling and feature extraction on the RGB image and the feature image to obtain a first feature map;

[0262] A first residual module, which is used to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map;

[0263] A second residual module, which is used to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map;

[0264] An upsampling module, which is used to perform upsampling and feature extraction on the third feature map, and obtain the second face image based on the result obtained by performing upsampling and feature extraction on the third feature map.

[0265] Optionally, the style transfer model further includes a branch convolutional layer; the branch convolutional layer is used to perform downsampling and feature extraction on the RGB image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0266] The upsampling module includes a first sub-module and a second sub-module. The first sub-module is used to perform upsampling and feature extraction on the third feature map, and the second sub-module is used to add the output result of the first sub-module to the fourth feature map to obtain the second face image.

[0267] Optionally, the style transfer model includes a parameter calculation layer. The parameter calculation layer is used to determine target parameters based on the first feature map and input the target parameters into the second residual module;

[0268] Wherein, the second normalization process includes performing normalization processing based on the target parameters.

[0269] Optionally, the downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. Among them, the downsampling sub-module includes:

[0270] A first convolutional layer, which is used to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map;

[0271] A first normalization layer, which is used to perform normalization processing on the first sub-feature map to obtain an output first result;

[0272] Wherein, the first image corresponding to the first downsampling sub-module is the RGB image and the feature image, and the corresponding first result is the second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is the third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

[0273] Optionally, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes:

[0274] An upsampling layer, which is used to perform upsampling on the input second image to obtain a fourth sub-feature map;

[0275] A second convolutional layer, which is used to perform feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map;

[0276] A second normalization layer, which is used to perform normalization processing on the fifth sub-feature map to obtain an output second result;

[0277] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the middle upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last middle upsampling sub-module, and the corresponding second result is the second face image.

[0278] Optionally, the style transfer model to be trained further includes a branch convolutional layer; the branch convolutional layer is used to downsample and extract features from the RGB sample image to obtain a fourth feature map, and input the fourth feature map into the upsampling module;

[0279] The upsampling module includes a first sub-module and a second sub-module. The first sub-module is used to upsample and extract features from the third feature map, and the second sub-module is used to add the output result of the first sub-module and the fourth feature map to obtain the stylized face image.

[0280] Optionally, the style transfer model to be trained further includes a parameter calculation layer. The parameter calculation layer is used to determine target parameters based on the first feature map and input the target parameters into the second residual module;

[0281] Among them, the second normalization process includes normalization processing based on the target parameters.

[0282] Optionally, the downsampling module includes a first downsampling sub-module, N middle downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. Among them, the downsampling sub-module includes:

[0283] A first convolutional layer, which is used to downsample and extract features from the input first image to obtain a first sub-feature map;

[0284] A first normalization layer, which is used to normalize the first sub-feature map to obtain the output first result;

[0285] Among them, the first image corresponding to the first downsampling sub-module is the RGB sample image and the feature sample image, and the corresponding first result is the second sub-feature map; the first image corresponding to the middle downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is the third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last middle downsampling sub-module, and the corresponding first result is the first feature map.

[0286] Optionally, the upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes:

[0287] An upsampling layer for upsampling the input second image to obtain a fourth sub-feature map;

[0288] A second convolutional layer for feature extraction of the fourth sub-feature map to obtain a fifth sub-feature map;

[0289] A second normalization layer for normalizing the fifth sub-feature map to obtain an output second result;

[0290] Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the stylized face image.

[0291] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0292] The communication interface is used for communication between the above terminal and other devices.

[0293] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0294] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it can also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0295] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, the computer is enabled to execute the steps of the image generation method or the steps of the training method of the style transfer model described in any one of the above embodiments.

[0296] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is enabled to execute the steps of the image generation method or the steps of the training method of the style transfer model described in any one of the above embodiments.

[0297] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).

[0298] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0299] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.

[0300] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. An image generation method, characterized in that, Including: Obtain the Red-Green-Blue (RGB) image of the first face image; Obtain the feature image corresponding to the target region in the first face image; wherein, the target region is used to represent the region where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features; Input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is the face image obtained by performing style transfer on the first face image; The style transfer model includes: A downsampling module, which is used to downsample and extract features from the RGB image and the feature image to obtain a first feature map; A first residual module, which is used to extract features and perform a first normalization process on the first feature map to obtain a second feature map; A second residual module, which is used to perform a stylization process and a second normalization process on the second feature map to obtain a third feature map; An upsampling module, which is used to upsample and extract features from the third feature map, and obtain the second face image based on the result obtained by upsampling and extracting features from the third feature map.

2. The method according to claim 1, wherein The obtaining the feature image corresponding to the target region in the first face image includes: Obtain the key point set in the first face image, where the key point set includes coordinate points for representing the contour information of the target feature; Generate the feature image corresponding to the target region based on the key point set of the target region.

3. The method according to claim 1, characterized in that The method further includes: Perform face segmentation on the first face image to determine the first region and the second region in the first face image, where the first region is used to represent the region where the portrait feature is located, and the second region is used to represent the region where the background feature is located; Perform color adjustment on the first region to obtain a third face image, where the third face image is the face image obtained by performing color adjustment on the first region of the first face image; Perform compositing rendering on the third face image and the second face image to obtain a fourth face image.

4. The method according to claim 1, characterized in that, The style transfer model further includes a branch convolutional layer; the branch convolutional layer is used to downsample and extract features from the RGB image to obtain a fourth feature map, and input the fourth feature map into the upsampling module; The upsampling module includes a first sub-module and a second sub-module. The first sub-module is used to upsample and extract features from the third feature map, and the second sub-module is used to add the output result of the first sub-module and the fourth feature map to obtain the second face image.

5. The method according to claim 1, wherein The style transfer model includes a parameter calculation layer, which is used to determine target parameters based on the first feature map and input the target parameters into the second residual module; Wherein, the second normalization process includes performing normalization processing based on the target parameters.

6. The method according to claim 1, characterized in that The downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. Among them, the downsampling sub-module includes: A first convolutional layer, which is used to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map; A first normalization layer, which is used to perform normalization processing on the first sub-feature map to obtain an output first result; Among them, the first image corresponding to the first downsampling sub-module is the RGB image and the feature image, and the corresponding first result is the second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is the third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

7. The method according to claim 1, characterized in that, The upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. Among them, the upsampling sub-module includes: An upsampling layer, which is used to perform upsampling on the input second image to obtain a fourth sub-feature map; A second convolutional layer, which is used to perform feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map; A second normalization layer, which is used to perform normalization processing on the fifth sub-feature map to obtain an output second result; Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the intermediate upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last intermediate upsampling sub-module, and the corresponding second result is the second face image.

8. A training method for a style transfer model, characterized in that, Including: Obtain the RGB sample image of the sample face image; Obtain the feature sample image corresponding to the target area in the sample face image; where the target area is used to represent the area where the target feature is located, and the target feature includes face features and at least one of the following: eye features, mouth features, and nose features; Iteratively train the to-be-trained style transfer model using the RGB sample image and the feature sample image; Use the loss function to detect the result output by the to-be-trained style transfer model to determine the loss value; If the change in the loss value is less than the preset value, then determine the currently trained to-be-trained style transfer model as the style transfer model; The to-be-trained style transfer model includes: A downsampling module, which is used to perform downsampling and feature extraction on the RGB sample image and the feature sample image to obtain a first feature map; A first residual module, which is used to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map; The second residual module is configured to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map; An upsampling module; the upsampling module is configured to perform upsampling and feature extraction on the third feature map, and obtain a stylized face image based on the result obtained by performing upsampling and feature extraction on the third feature map, where the stylized face image is a face image obtained by performing style transfer on the sample face image.

9. The method according to claim 8, wherein The to-be-trained style transfer model further includes a branch convolutional layer; the branch convolutional layer is configured to perform downsampling and feature extraction on the RGB sample image to obtain a fourth feature map, and input the fourth feature map into the upsampling module; The upsampling module includes a first sub-module and a second sub-module. The first sub-module is configured to perform upsampling and feature extraction on the third feature map, and the second sub-module is configured to add the output result of the first sub-module to the fourth feature map to obtain the stylized face image.

10. The method according to claim 8, characterized in that, The to-be-trained style transfer model further includes a parameter calculation layer, and the parameter calculation layer is configured to determine a target parameter based on the first feature map and input the target parameter into the second residual module; Wherein, the second normalization processing includes performing normalization processing based on the target parameter.

11. The method according to claim 8, wherein The downsampling module includes a first downsampling sub-module, N intermediate downsampling sub-modules, and a last downsampling sub-module connected in sequence, where N is a positive integer. The downsampling sub-module includes: A first convolutional layer configured to perform downsampling and feature extraction on the input first image to obtain a first sub-feature map; A first normalization layer configured to perform normalization processing on the first sub-feature map to obtain an output first result; Wherein, the first image corresponding to the first downsampling sub-module is the RGB sample image and the feature sample image, and the corresponding first result is a second sub-feature map; the first image corresponding to the intermediate downsampling sub-module is the first result output by the previous downsampling sub-module, and the corresponding first result is a third sub-feature map; the first image corresponding to the last downsampling sub-module is the first result output by the last intermediate downsampling sub-module, and the corresponding first result is the first feature map.

12. The method according to claim 8, wherein The upsampling module includes a first upsampling sub-module, M intermediate upsampling sub-modules, and a last upsampling sub-module connected in sequence, where M is a positive integer. The upsampling sub-module includes: An upsampling layer configured to perform upsampling on the input second image to obtain a fourth sub-feature map; A second convolutional layer configured to perform feature extraction on the fourth sub-feature map to obtain a fifth sub-feature map; A second normalization layer configured to perform normalization processing on the fifth sub-feature map to obtain an output second result; Among them, the second image corresponding to the first upsampling sub-module is the third feature map, and the corresponding second result is the sixth sub-feature map; the second image corresponding to the middle upsampling sub-module is the second result output by the previous upsampling sub-module, and the corresponding second result is the seventh sub-feature map; the second image corresponding to the last upsampling sub-module is the second result output by the last middle upsampling sub-module, and the corresponding second result is the stylized face image.

13. An image generation device, characterized in that, Including: A first acquisition module, configured to acquire a red, green, and blue (RGB) image of a first face image; A second acquisition module, configured to acquire a feature image corresponding to a target area in the first face image; wherein, the target area is used to represent an area where a target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature; An input module, configured to input the RGB image and the feature image into a pre-trained style transfer model for style transfer to obtain a second face image, where the second face image is a face image obtained by performing style transfer on the first face image; The style transfer model includes: A downsampling module, configured to downsample and perform feature extraction on the RGB image and the feature image to obtain a first feature map; A first residual module, configured to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map; A second residual module, configured to perform style processing and second normalization processing on the second feature map to obtain a third feature map; An upsampling module, configured to upsample and perform feature extraction on the third feature map, and obtain the second face image based on the result obtained by upsampling and performing feature extraction on the third feature map.

14. A training device for a style transfer model, characterized in that, Including: A third acquisition module, configured to acquire an RGB sample image of a sample face image; A fourth acquisition module, configured to acquire a feature sample image corresponding to a target area in the sample face image; wherein, the target area is used to represent an area where a target feature is located, and the target feature includes a face feature and at least one of the following: an eye feature, a mouth feature, and a nose feature; An iterative training module, configured to iteratively train a style transfer model to be trained by using the RGB sample image and the feature sample image; A first determination module, configured to detect the result output by the style transfer model to be trained by using a loss function, and determine a loss value; A second determination module, configured to, if the change in the loss value is less than a preset value, determine the currently trained style transfer model to be trained as the style transfer model; The style transfer model to be trained includes: A downsampling module, configured to downsample and perform feature extraction on the RGB sample image and the feature sample image to obtain a first feature map; A first residual module, configured to perform feature extraction and first normalization processing on the first feature map to obtain a second feature map; The second residual module is configured to perform stylization processing and second normalization processing on the second feature map to obtain a third feature map; An upsampling module; the upsampling module is configured to perform upsampling and feature extraction on the third feature map, and obtain a stylized face image based on the result obtained by performing upsampling and feature extraction on the third feature map, where the stylized face image is a face image obtained by performing style transfer on the sample face image.

15. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing programs; The processor is configured to implement the steps of the method described in any one of claims 1-12 when executing the program stored on the memory.

16. A readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, it implements the steps of the method described in any one of claims 1-12.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN112102154A

  • Skin rendering method for 3D model and equipment

    CN113870404A