Stylized image generation method, device and image processing equipment
Multi-scale feature extraction and hidden vector correction are performed through variational autoencoder, and stylized images are generated by stylized image generators, which solves the problems of insufficient image stylized conversion quality and insufficient identity information retention ability in the prior art, and achieves high-quality stylized image generation.
Patent Information
- Application Number
- CN202210207313.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-03-03
AI Technical Summary
The prior art is difficult to effectively improve the quality of generated images in image stylized conversion, and the identity information retaining ability of the stylized image is insufficient.
The input image is extracted by a variational autoencoder multi-scale feature, calculate the mean and variance of the hidden vectors of each feature dimension, sample and obtain the target space hidden vector, and generate the stylized image using the stylized image generator. At the same time, the concept of spatial hidden vector correction is introduced to improve the identity information retention ability of stylized images.
The quality of the generated stylized images is improved, making them closer to the real feature data distribution of the input image, and the identity information retention ability of the stylized images is improved.
Smart Images

Figure CN114612289B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field related to graphic image processing, and in particular to a stylized image generation method, device and image processing equipment. Background Art
[0002] Image stylization, also known as image style transfer, is a technology that can transfer a distinctive image style (such as artistic features) to another image, so that the original image retains the original content while having a unique artistic style, such as cartoon, comic, oil painting, watercolor, ink and other styles. For example, in a typical application scenario of image stylization technology, a face image input by a user can be stylized and converted to a face image of a specific style. For example, a face image input by a user can be stylized and converted to a face image of Disney style, Marvel style, anime style, etc., to meet the specific needs of users.
[0003] As users' demand for image stylization increases, how to improve the quality of stylized images obtained by image stylization conversion has always been a research topic that relevant personnel in this field have been committed to studying. Summary of the invention
[0004] Based on the above, in order to solve at least some of the above problems, in a first aspect, an embodiment of the present application provides a method for generating a stylized image, the method comprising:
[0005] Inputting the image to be converted into a variational autoencoder trained in advance, and processing the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted;
[0006] Obtaining a target space latent vector of the image to be converted in each feature dimension according to the mean and variance of the latent vector in each feature dimension;
[0007] The target space latent vector is input into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
[0008] In a possible implementation of this embodiment, obtaining the target space latent vector of the image to be converted in each feature dimension according to the mean and variance of the latent vector in each feature dimension includes:
[0009] Sampling the mean and variance of the latent vector of each feature dimension by the variational autoencoder to obtain a first spatial latent vector of the image to be converted in each feature dimension;
[0010] The first spatial latent vector of the image to be converted in each feature dimension is corrected to obtain the second spatial latent vector of the image to be converted in each feature dimension as the target spatial latent vector.
[0011] Among them, the correction formula for correcting the first spatial latent vector of the image to be converted in each feature dimension is:
[0012]
[0013] Where s+ is the second spatial latent vector, s is the first spatial latent vector, x represents the input image, ε θ represents the variational autoencoder, represents the stylized image generator, the The stylized image generated by the stylized image generator and the image to be converted are consistent at the pixel level. represents the semantic feature loss of the image to be converted, F represents the VGG network, which is used to calculate the perceptual similarity of the image to be converted from the first space, and w vgg is the preset weight parameter, The difference between the initial value s0 of the second spatial latent vector and the first spatial latent vector needs to be within a set range, which means that a small adjustment is made with s0 as the initial value and the set number of iterations is performed.
[0014] In a possible implementation of this embodiment, the variational autoencoder includes a plurality of encoding layers cascaded in sequence, a plurality of decoding layers cascaded in sequence, and fully connected layers respectively connected to the decoding layers, each encoding layer is connected to a corresponding decoding layer, and each decoding layer is connected to one of the fully connected layers;
[0015] The image to be converted is input from a first coding layer among the multiple coding layers, and each coding layer sequentially performs downscaling coding processing on the image to be converted to obtain a coding feature map and outputs the obtained coding feature map to the coding layer of the next layer and the decoding layer corresponding to the coding layer;
[0016] The input of the first decoding layer is the output of its corresponding encoding layer, and the input of each other decoding layer is the output of the previous decoding layer plus the output of its corresponding encoding layer. Each decoding layer decodes its input data and outputs feature maps of different dimensions, and processes the feature maps through the fully connected layer to obtain the mean and variance of latent vectors of different feature dimensions corresponding to the image to be converted;
[0017] Among the multiple decoding layers, starting from the first decoding layer, the dimensions of the feature maps output by each decoding layer decrease step by step.
[0018] In a possible implementation of this embodiment, the method further includes a step of training the variational autoencoder, the step comprising:
[0019] Acquire a first training data set, where the first training data set includes a plurality of sample images;
[0020] Inputting the sample images into the variational autoencoder in sequence to obtain a target space latent vector corresponding to the sample images;
[0021] Inputting the target space latent vector corresponding to the sample image into the trained stylized image generator to obtain a stylized image corresponding to the sample image;
[0022] A loss function of the variational autoencoder is calculated according to the stylized image and the sample image, and model parameters of the variational autoencoder are adjusted according to the loss function until a training convergence condition is met.
[0023] In a possible implementation of this embodiment, the loss function includes pixel level loss, semantic similarity loss, and identity information loss. The expression of the loss function is as follows:
[0024] L=L rec +W per L per +W kl L arc ,in:
[0025] represents the pixel level loss, x is the sample image, ε θ represents the variational autoencoder, represents a trained stylized image generator, and L2 represents the image distance between the sample image and the stylized image generated by the stylized image generator;
[0026] represents the semantic similarity loss, L lpips Represents the semantic feature similarity between the sample image and the stylized image generated by the sample image after semantic feature extraction by the semantic extraction model;
[0027] Represents the identity information loss, L arc represents the similarity between identity features obtained by performing identity information recognition on the sample image and the stylized image generated from the sample image;
[0028] The w per 、w id 、w kl are respectivelyper , L id , L arc Set the weight parameter.
[0029] In a possible implementation of this embodiment, the method further includes a step of pre-training the stylized image generator, the step including:
[0030] Acquire a stylized image dataset, where the stylized image dataset includes a plurality of stylized sample images;
[0031] Inputting the stylized sample images into the stylized image generator to be trained in sequence, obtaining generated images corresponding to the stylized sample images, and calculating the loss function value of the stylized image generator;
[0032] Optimizing the network parameters of the stylized image generator according to the loss function value until the calculated loss function value satisfies the training convergence condition, thereby obtaining a trained stylized image generator;
[0033] The stylized image generator includes convolutional layers corresponding to different image resolutions respectively. When optimizing the network parameters of the stylized image generator, the network parameters of the convolutional layers corresponding to image resolutions smaller than the set resolution remain unchanged.
[0034] The calculation formula of the loss function value of the stylized image generator is as follows:
[0035] in:
[0036]
[0037] x~p d represents the distribution of the stylized image dataset, represents the distribution of a data set consisting of stylized images generated by the stylized image generator according to each of the stylized sample images, D is a discriminator, Represents the discriminator gradient calculation operator for the stylized sample image.
[0038] In a second aspect, this embodiment further provides a stylized image generation device, which is applied to an image processing device, and the stylized image generation device includes:
[0039] A residual calculation module, used for inputting the image to be converted into a variational autoencoder trained in advance, and processing the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted;
[0040] A latent vector processing module, used to obtain the target space latent vector of the image to be converted in each feature dimension according to the mean and variance of the latent vector of each feature dimension;
[0041] The image generation module is used to input the target space latent vector into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
[0042] In a third aspect, this embodiment further provides an image processing device, comprising a machine-readable storage medium and one or more processors, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the one or more processors, the above-mentioned method is implemented.
[0043] Based on the above content of the embodiments of the present application, compared with the prior art, the stylized image generation method, device and image processing device provided by the embodiments of the present application perform multi-scale feature extraction on the input image through a variational autoencoder (for example, a variational autoencoder with a feature pyramid residual network structure), and first calculate (estimate) the mean and variance of the latent vector of each feature dimension of the input image based on the feature extraction result, and then obtain the latent vector of the target space (such as S space) based on the mean and variance sampling, and finally generate the corresponding stylized image based on the latent vector by the stylized image generator. In this way, the features extracted by multi-scale feature extraction of the image will express the image more accurately, which can improve the image quality of the generated stylized image.
[0044] Furthermore, this embodiment introduces the concept of spatial latent vector correction. For example, the S spatial latent vector (first spatial latent vector) is corrected to obtain the S+ spatial latent vector (second spatial latent vector) for generating a stylized image, so that the generated stylized image has a better ID (identity information) retention capability, and the generated stylized image is closer to the real feature data distribution of the input image, thereby further improving the image quality of the generated stylized image. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 It is a flowchart of the stylized image generation method provided in an embodiment of the present application.
[0047] Figure 2It is a schematic diagram provided in an embodiment of the present application for describing the working process and principle of the variational autoencoder in this embodiment.
[0048] Figure 3 It is a flowchart for training the variational autoencoder provided in an embodiment of the present application.
[0049] Figure 4 It is a schematic diagram for describing the exemplary process of training the variational autoencoder mentioned above.
[0050] Figure 5 It is a schematic diagram of the training process of the stylized image generator used in this embodiment.
[0051] Figure 6 It is a schematic diagram of the comparison before and after the stylized image generator provided in the embodiment of the present application is trained.
[0052] Figure 7 It is a schematic diagram of an exemplary process of using a trained variational autoencoder and stylized image generation to perform stylized processing on an image to be converted to generate a stylized image of a corresponding style in an embodiment of the present application.
[0053] Figure 8 It is a schematic diagram of an image processing device provided in an embodiment of the present application.
[0054] Fig. 9 It is a schematic diagram of the functional modules of the stylized image generation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0056] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0057] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0058] Before introducing the relevant technical solutions provided by the embodiments of the present application in detail, in order to more clearly understand the relevant technical solutions of the embodiments of the present application, the relevant technical terms involved in the embodiments of the present application are first explained.
[0059] First, before exemplarily introducing the relevant embodiments provided in the present application, the relevant terms involved in the present application are explained.
[0060] Hidden space (style adversarial network StyleGAN2), generally includes Z space, W space, W+ space and S space.
[0061] Z space is the most original input space, which is generally a standard normal distribution or uniform distribution and can be called a random noise space.
[0062] W-space, the latent space obtained by transforming Z-space using a series of fully connected layers, is generally believed to reflect the learned disentanglement properties better than Z-space.
[0063] The W+ space is similar to the W space in construction method, but the potential vectors fed to each layer of the generator are different. It is often used for style mixing and image inversion.
[0064] S space, based on W space, performs further transformation. For each layer of the generator, a different affine transformation (Affine Transformations, a linear transformation followed by a translation operation) is used to map ω∈W into a channel-level style parameter s.
[0065] The inventor of the present application has found through research that in a possible technology for image stylization, it is possible to convert an input image (e.g., a face image) into a stylized image of a set style (e.g., a Disney-style image) based on transfer learning and convolutional layer swapping technology. This method generally uses an iteratively optimized inverse mapping method to calculate the latent space vector of the input image, and then generates a stylized image based on the latent space vector. This method has a slow calculation speed and may not be able to maintain the distribution of the latent vector space of the input image itself, thereby affecting the quality of the generated stylized image.
[0066] In view of this, the embodiment of the present application innovatively proposes a stylized image generation scheme, which performs multi-scale feature extraction on the input image through a variational autoencoder (for example, a variational autoencoder with a feature pyramid residual network structure), and first calculates (estimates) the mean and variance of the latent vector of each feature dimension of the input image based on the feature extraction result, and then obtains the latent vector of the target space (such as S space) based on the mean and variance sampling, and finally generates the corresponding stylized image based on the latent vector by the stylized image generator. In this way, the generated stylized image can be closer to the real feature data distribution of the input image, and the generation quality and effect of the stylized image can be improved. The specific implementation method of the embodiment of the present application will be described exemplarily in conjunction with the accompanying drawings below.
[0067] like Figure 1 , which is a flow chart of the stylized image generation method provided in the embodiment of the present application. It should be understood that the order of some of the steps included in the stylized image generation method provided in the present embodiment can be interchanged according to actual needs during actual implementation, or some of the steps can be omitted or deleted, and this embodiment does not specifically limit this.
[0068] The following is a detailed description of each step of the stylized image generation method of this embodiment by way of example. Figure 1 As shown, the method may include the relevant contents described in the following steps S100 to S300.
[0069] Step S100: input the image to be converted into a variational autoencoder trained in advance, and process the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted.
[0070] In a possible implementation of this embodiment, the variational autoencoder may include a plurality of encoding layers cascaded in sequence, a plurality of decoding layers cascaded in sequence, and fully connected layers respectively connected to the decoding layers, each encoding layer is connected to a corresponding decoding layer, and each decoding layer is connected to one of the fully connected layers.
[0071] The input of the first decoding layer is the output of its corresponding encoding layer, and the input of each other decoding layer is the output of the decoding layer above it plus the output of its corresponding encoding layer. Each decoding layer decodes its input data and outputs a feature map of different dimensions, and processes the feature map through the fully connected layer to obtain the mean and variance of the latent vectors of different feature dimensions corresponding to the image to be converted.
[0072] Among them, in this embodiment, among the multiple decoding layers, starting from the first decoding layer, the dimensions of the feature maps output by each decoding layer are gradually reduced.
[0073] For example Figure 2 As shown in FIG. 1 , it is a schematic diagram for describing the workflow and principle of the variational autoencoder provided by an embodiment of the present application. As an example, the multiple encoding layers included in the variational autoencoder may be five encoding layers cascaded in sequence, such as Figure 2 As shown in FIG. 1 , E1, E2, E3, E4, and E5, E1 is used as the first encoding layer and E5 is used as the last encoding layer. The multiple decoding layers included in the variational autoencoder may be five decoding layers cascaded in sequence, such as Figure 2 D1, D2, D3, D4, D5 are shown, where D1 is the first decoding layer and D5 is the last decoding layer. Encoding layers E1, E2, E3, E4, E5 are connected to decoding layers D5, D4, D3, D2, D1 respectively.
[0074] Among them, the image to be converted is input from the first coding layer among the multiple coding layers, and each coding layer sequentially downscales the image to be converted to obtain a coding feature map and outputs the obtained coding feature map to the next coding layer and the decoding layer corresponding to the coding layer.
[0075] For example, combined with Figure 2 As shown, the image to be converted Z1 may be an image with a scale of H (such as a pixel resolution of 1024), which is input from the first coding layer E1, and the first coding layer E1 may downscale the image to be converted Z1 by downsampling. For example, the image to be processed may be processed into a feature map with a scale of H / 2 (such as a pixel resolution of 512) by 2 times downsampling, and the processed feature map is input into the next coding layer E2 and its corresponding decoding layer D5.
[0076] The second encoding layer E2 can downscale the feature map output by the encoding layer E1 by downsampling. For example, the feature map output by the encoding layer E1 can also be downscaled by 2 times downsampling, and the feature map obtained after processing is input into the next encoding layer E3 and its corresponding decoding layer D4.
[0077] The third encoding layer E3 can downscale the feature map output by the encoding layer E2 by downsampling. For example, the feature map output by the encoding layer E2 can also be downscaled by 2 times downsampling to obtain a feature map with a scale of H / 4, and the feature map obtained after processing is input into the next encoding layer E4 and its corresponding decoding layer D3.
[0078] The fourth encoding layer E4 can downscale the feature map output by the encoding layer E3 by downsampling. For example, the feature map output by the encoding layer E3 can be downscaled by 2 times downsampling to obtain a feature map with a scale of H / 8, and the feature map obtained after processing is input into the next encoding layer E5 and its corresponding decoding layer D2.
[0079] The fifth encoding layer E5 is the last encoding layer, and the feature map output by the encoding layer E4 can be downscaled by downsampling. For example, the feature map output by the encoding layer E4 can be downscaled by 2 times downsampling to obtain a feature map with a scale of H / 16, and the feature map obtained after processing is input to its corresponding decoding layer D1.
[0080] Correspondingly, the decoding layer D1, decoding layer D2, decoding layer D3, decoding layer D4, and decoding layer D5 respectively decode the input feature map (for example, upsampling can be performed) to obtain the feature map of each feature dimension of the image to be converted, and transmit them to the corresponding fully connected layer (FC as shown in the figure). Among them, the output of each of the decoding layer D1, decoding layer D2, decoding layer D3, and decoding layer D4 is also used as the input of the corresponding next decoding layer. For example, the output of decoding layer D1 will be used as the input of decoding layer D2, the output of decoding layer D2 will be used as the input of decoding layer D3, the output of decoding layer D3 will be used as the input of decoding layer D4, and the output of decoding layer D4 will be used as the input of decoding layer D5.
[0081] The dimensions of the feature maps output by the decoding layer D1, the decoding layer D2, the decoding layer D3, the decoding layer D4, and the decoding layer D5 are reduced step by step. For example, the feature dimensions of the feature maps output by the decoding layer D1, the decoding layer D2, the decoding layer D3, the decoding layer D4, and the decoding layer D5 are 15*512, 3*256, 3*128, 3*64, and 2*32, respectively. In this embodiment, each encoding layer may include at least one convolutional layer (such as three downsampling convolutional layers), and each decoding layer may also include at least one convolutional layer (such as three upsampling convolutional layers). Therefore, each decoding layer can output feature map data corresponding to the three convolutional layers. For example, based on Figure 2 In this example, the present embodiment includes five decoding layers, and accordingly, the variational autoencoder can have an output of 5 feature dimensions.
[0082] The fully connected layer processes the input feature maps respectively to obtain the mean of the latent vectors of different feature dimensions corresponding to the image to be converted (such as Figure 2 Z μ ) and variance (such as Figure 2 Z σ ). Then according to the mean of the latent vector (such as Figure 2 Z μ ) and variance (such as Figure 2 Z σ ) obtains the latent vector S and inputs it into the stylized image generator to generate the corresponding stylized image Z2.
[0083] Step S200, obtaining the target space latent vector of the image to be converted in each feature dimension according to the mean and variance of the latent vector in each feature dimension.
[0084] Among them, in a possible implementation of this embodiment, each of the fully connected layers can process the mean and variance of the latent vectors of each feature dimension obtained to obtain the target space latent vector of the image to be converted in each feature dimension.
[0085] Based on step S200, the mean and variance of the latent vector of each feature dimension may be sampled by the variational autoencoder to obtain the first spatial latent vector of the image to be converted in each feature dimension. For example, the mean and variance may be sampled by the fully connected layer to obtain the first spatial latent vector of the image to be converted in each feature dimension. The first spatial latent vector may be an S spatial latent vector.
[0086] The formula for sampling the mean and variance may be: s is the first spatial latent vector.
[0087] Then, the first spatial latent vector of the image to be converted in each feature dimension is corrected to obtain the second spatial latent vector of the image to be converted in each feature dimension as the target spatial latent vector. The second spatial latent vector may be a latent vector obtained by correcting the S spatial latent vector, which may be defined as the S+ spatial latent vector in this embodiment.
[0088] In some possible implementations, the target space latent vector may also be directly the S space latent vector. However, the inventors have found through experimental research that in the process of image stylization conversion, for example, when converting from the real domain to the stylized domain, the ID retention ability of the character is relatively poor, so that the stylized image generated after the stylization conversion is not consistent with the character ID information (such as facial features) in the input real image in terms of subtle details. Based on the discovery of this technical problem, in the embodiment of the present application, the concept of space latent vector correction is introduced, for example, the S space latent vector is corrected so that the stylized image obtained in the subsequent stylized image generation has a better ID retention ability.
[0089] As a possible example, in this embodiment, the formula for correcting the first spatial latent vector of the image to be converted in each feature dimension may be:
[0090]
[0091] Where s+ is the second spatial latent vector, s is the first spatial latent vector, x represents the image to be converted, ε θ represents the variational autoencoder, represents the stylized image generator, the Characterizing the difference between the stylized image generated by the stylized image generator and the image to be converted at the pixel level, represents the semantic feature loss of the image to be converted, F represents the VGG network, which is used to calculate the perceptual similarity of the image to be converted from the first space, and w vgg is the preset weight parameter, The difference between the initial value s0 representing the second spatial latent vector and the first spatial latent vector needs to be within a set range, and a small adjustment is made with s0 as the initial value and the set number of iterations is performed.
[0092] Step S300: input the target space latent vector into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
[0093] In summary, in this embodiment, the image to be converted is sent from the bottom to the different coding layers of the variational autoencoder, and downscaling is performed, and feature extraction is performed on the feature maps of different scales corresponding to the image to be converted, and a feature pyramid can be obtained. The feature map corresponding to the bottom of the feature pyramid is a high scale (such as high resolution), and the feature map corresponding to the top is a low scale (such as low resolution). The higher the level, the smaller the image and the smaller the scale. In this way, the feature dimension can be increased based on the scale change of the image to construct a high-dimensional feature.
[0094] At the same time, the low-level coding layer of the variational autoencoder in this example can focus more on the detailed features of the image, and the high-level coding layer can focus on the deep semantic information of the image. The encoder with this structure can express the image more accurately when performing multi-scale feature extraction on the image.
[0095] Secondly, in this embodiment, the mean and variance of the latent vector of each feature dimension of the image are predicted based on the extracted features, and then the target space latent vector is obtained based on the predicted mean and variance and input into the stylized image generator to generate the corresponding stylized image, which can improve the quality of the generated stylized image.
[0096] Among them, since the variational autoencoder of this embodiment performs residual prediction based on the feature pyramid for subsequent style image generation, the variational autoencoder of this embodiment can also be defined as a "feature pyramid residual variational autoencoder".
[0097] Furthermore, the variational autoencoder provided in this embodiment can be obtained in advance through training. Figure 3 As shown, the step of training the variational autoencoder may include the following steps S310-S340, which are exemplarily introduced as follows.
[0098] Step S310: Acquire a first training data set, where the first training data set includes a plurality of sample images.
[0099] In this embodiment, the sample image may be a pre-collected face image with face features. The variational autoencoder may be trained based on the reconstruction principle and implemented in conjunction with any stylized image generator obtained through pre-training.
[0100] Step S320: input the sample images into the variational autoencoder in sequence to obtain the spatial latent vectors corresponding to the sample images.
[0101] In this embodiment, the spatial latent vector corresponding to the sample image may be the S spatial latent vector, and no correction is required during the training process.
[0102] Step S330: input the spatial latent vector corresponding to the sample image into the trained stylized image generator to obtain a stylized image corresponding to the sample image.
[0103] In this embodiment, the trained stylized image generator may be a generator for generating stylized images of any specific style. For example, in this embodiment, the stylized image generator used in the training process of the variational autoencoder may be StyleGANv2-ADA, and the style of the stylized image generated by it may be the original style of the sample image, or another style different from the original style of the sample image. During the training process, the original style may be used. At the same time, during the training process, the model parameters (such as weights) of the stylized image generator remain fixed, and only the model parameters of the variational autoencoder need to be adjusted so that the stylized image generated (or reconstructed) by the stylized image generator is consistent with the sample image as the training target of the variational autoencoder. In this way, the z obtained by the encoder can be μ (the mean of the latent vector) and z σ (The variance of the latent vector) can satisfy the normal distribution as much as possible. For example, we can define the mean and variance obtained by sampling to satisfy the following Kullback-Leibler divergence loss formula:
[0104]
[0105] D kl is the Kullback-Leiler divergence formula, p(z) represents the normal distribution, and z in the expanded formula is μ,i represents the mean of the latent vector z of the i-th dimension, z σ,i is the variance of the latent vector z in the i-th dimension.
[0106] Step S340, calculating the loss function of the variational autoencoder according to the stylized image and the sample image, and adjusting the model parameters of the variational autoencoder according to the loss function until a training convergence condition is met, thereby obtaining a trained variational autoencoder.
[0107] In this embodiment, the loss function may include pixel level loss, semantic similarity loss, identity information loss, and the expression of the loss function is as follows:
[0108] L=L rec +W per L per +W kl L arc .
[0109] in: represents the pixel level loss, x is the sample image, ε θ represents the variational autoencoder, represents a pre-trained stylized image generator, and L2 represents the image distance between the sample image and the stylized image generated by the stylized image generator. For example, the image distance may represent the similarity between the sample image and the stylized image generated by the stylized image generator, and may be represented by cosine distance or Euclidean distance. The training process needs to make the image distance between the two as close as possible or equal to 0.
[0110] represents the semantic similarity loss, L lpips Represents the semantic feature similarity between the sample image and the stylized image of the generated sample image calculated after semantic feature extraction by the semantic extraction model. In this embodiment, the semantic extraction model can be any model that can realize image semantic extraction, for example, it can be a lpips (Learned Perceptual Image Patch Similarity) model.
[0111] Represents the identity information loss, L arcRepresents the similarity between identity features obtained by performing identity information recognition on the sample image and the stylized image of the generated sample image, for example, the similarity between facial features obtained by performing face recognition on the sample image and the stylized image of the generated sample image, which can also be represented by cosine similarity.
[0112] The w per 、w id 、w kl are respectively per , L id , L arc Set the weight parameter.
[0113] The above training process of the variational autoencoder can be referred to Figure 4 In detail, in this embodiment, the sample image P0 can be input into the variational autoencoder to be trained to predict the mean and variance, and the mean z of the latent vector of each feature dimension of the sample image P0 can be obtained. μ and variance z σ , then based on the mean z μ and variance z σ Sampling is performed to obtain the corresponding latent vector S, and the latent vector S is input into the existing (pre-trained) stylized image generator for stylized image generation to obtain the stylized image P1. Finally, the loss function value of the variational autoencoder is calculated based on the stylized image P1 and the sample image P0 (used to characterize the consistency of the reconstructed / generated images P1 and P0), and then the model parameters (or weights) of the variational autoencoder are further iteratively updated or iteratively optimized based on the loss function value to complete the training of the variational autoencoder. Among them, the stylized image generator used in the training process can be a generator for generating any kind of stylized image, for example, it can also be a generator in which the style of the input image and the generated image is consistent, and this embodiment does not limit this.
[0114] Secondly, for the stylized image generator used in step S300, Figure 5 The training process shown is implemented, and the training process may include the following steps S510-S530, which are exemplarily introduced as follows.
[0115] Step S510: obtaining a stylized image dataset, where the stylized image dataset includes a plurality of stylized sample images.
[0116] For example, in a possible implementation, a plurality of (eg, 110) images of a set style (eg, Marvel character portrait style) may be crawled from the Internet as the stylized image dataset.
[0117] Step S520: input the stylized sample images into the stylized image generator to be trained in sequence, obtain the generated images corresponding to the stylized sample images, and calculate the loss function value of the stylized image generator.
[0118] In this embodiment, the loss function value can be constructed according to the similarity between the generated image corresponding to the stylized sample image and the stylized sample image. For example, the corresponding loss function value can be obtained according to the Euclidean distance, cosine similarity, Pearson similarity, etc. between the two. The loss function value can be used to characterize whether the image reconstructed by the stylized image generator from the stylized sample image is consistent with or nearly consistent with the original image.
[0119] Step S530, optimizing the network parameters of the stylized image generator according to the loss function value until the calculated loss function value meets the training convergence condition, thereby obtaining a trained stylized image generator.
[0120] The training convergence condition may be that the loss function value is less than a set threshold, or the number of iterative training reaches a set number of training times.
[0121] The stylized image generator includes convolutional layers corresponding to different image resolutions respectively. When optimizing the network parameters of the stylized image generator, the network parameters of the convolutional layers corresponding to image resolutions smaller than the set resolution remain unchanged.
[0122] For example Figure 6 As shown, in this embodiment, the stylized image generator to be trained (the pre-trained model in the figure) may include multiple convolutional layers of different resolutions, such as convolutional layers of resolutions such as 8×8, 16×16, 32×32, 64×64, 128×128, 256×256, 512×512, and 1024×1024. Taking the stylized image generator for face portraits as an example, low-resolution convolutional layers such as 8×8 and 16×16 are used to control geometric structure (geometry) information such as face orientation and face shape. When converting face stylization, an important step from the real domain to the stylized domain is to maintain the consistency of face orientation (pose) and face shape with the input image. Therefore, in this embodiment, the network parameters of the two low-resolution convolutional layers of 8×8 and 16×16 can be fixed, and only the network parameters of the high-resolution convolutional layers such as 32×32 to 1024×1024 need to be adjusted. In other embodiments, the specific convolutional layer whose network parameters need to be kept fixed during the training process can be determined according to actual conditions and is not limited to the above examples.
[0123] In a possible implementation of this embodiment, the calculation formula of the loss function value of the stylized image generator is as follows:
[0124] in:
[0125]
[0126] x~p d represents the distribution of the stylized image dataset, represents the distribution of a data set consisting of stylized images generated by the stylized image generator according to each of the stylized sample images, D is a discriminator, Represents the discriminator gradient calculation operator for the stylized sample image.
[0127] The above-mentioned training process of stylized image generation can be an unsupervised image stylization training method based on small samples. This method can quickly develop generators of different styles, such as martial arts style, medieval style, optimization style, etc., through small samples (such as 110 samples) and low cost, providing a low-cost and high-efficiency standardized process for new special effects in various application scenarios of image stylization transformation (such as live broadcast and short video application scenarios).
[0128] The process of stylizing the image to be converted using the trained variational autoencoder and stylized image generation to generate a stylized image of the corresponding style can be found in Figure 7 shown.
[0129] In detail, in this embodiment, the image to be converted M0 can be input into the trained variational autoencoder to predict the mean and variance, and the mean z of the latent vector of each feature dimension of the image to be converted M0 can be obtained. μ and variance z σ , then based on the mean z μ and variance z σ Sampling is performed to obtain the corresponding latent vector S. Then, the latent vector S (latent vector in the first space) is corrected to obtain the corrected latent vector S+ (latent vector in the second space). Finally, the corrected latent vector S+ is input into the trained stylized image generator to generate a stylized image, and the generated stylized image M1 is obtained. Based on the above description, during the training process of the variational autoencoder, the training of the variational autoencoder can be completed without correcting the latent vector S of the sample image in the training process, which can speed up the training process and improve the training efficiency. In the application process after the training is completed, the latent vector S (latent vector in the first space) corresponding to the input image (image to be converted) is corrected before generating a stylized image, which can further enhance the ID retention capability of the object to be converted (such as a face image) during the stylized conversion process, and further enable the generated stylized image to retain the subtle details of the input image.
[0130] See also Figure 8 As shown, Figure 8 is a schematic diagram of an image processing device 100 for implementing the above-mentioned stylized image generation method provided in an embodiment of the present application. In detail, the image processing device 100 may include one or more processors 110, a machine-readable storage medium 120, and a stylized image generation device 130. The processor 110 and the machine-readable storage medium 120 may be communicatively connected via a system bus. The machine-readable storage medium 120 stores machine-executable instructions, and the processor 110 implements the above-described stylized image generation method by reading and executing the machine-executable instructions in the machine-readable storage medium 120. In this embodiment, the image processing device 100 may be, but is not limited to, a computer device with image processing capabilities such as a personal computer, a laptop computer, a smart phone, a tablet computer, a server, a cloud service platform, etc.
[0131] The machine-readable storage medium 120 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The machine-readable storage medium 120 is used to store a program, and the processor 110 executes the program after receiving an execution instruction.
[0132] The processor 110 may be an integrated circuit chip with signal processing capability. The processor may be, but is not limited to, a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.
[0133] Please refer to Fig. 9, is a functional module diagram of the stylized image generating device 130. In this embodiment, the stylized image generating device 130 may include one or more software function modules running on the image processing device 100, and these software function modules may be stored in the machine-readable storage medium 120 in the form of computer programs, so that when these software function modules are called and executed by the processor 130, the stylized image generating method described in the embodiment of the present application can be implemented.
[0134] In detail, the stylized image generating device 130 may include a residual calculation module 131 , a latent vector processing module 132 , and an image generating module 133 .
[0135] The residual calculation module 131 is used to input the image to be converted into a variational autoencoder trained in advance, and process the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted.
[0136] In this embodiment, the residual calculation module 131 is used to execute step S100 in the above method embodiment. For more details about the residual calculation module 131, please refer to the above specific description of step S100, which will not be repeated here.
[0137] The latent vector processing module 132 is used to obtain the target space latent vector of the image to be converted in each feature dimension according to the mean and variance of the latent vector in each feature dimension.
[0138] In this embodiment, the latent vector processing module 132 can be used to execute step S200 in the above method embodiment. For more details about the residual calculation module 132, please refer to the above description of the specific content of step S200, which will not be repeated here.
[0139] The image generation module 133 is used to input the target space latent vector into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
[0140] In this embodiment, the image generation module 133 is used to execute step S300 in the above method embodiment. For more details about the image generation module 133, please refer to the above description of the specific content of step S300, which will not be repeated here.
[0141] In summary, the stylized image generation method, device and image processing device provided in the embodiments of the present application perform multi-scale feature extraction on the input image through a variational autoencoder (e.g., a variational autoencoder with a feature pyramid residual network structure), and first calculate (estimate) the mean and variance of the latent vector of each feature dimension of the input image based on the feature extraction result, and then obtain the latent vector of the target space (such as S space) based on the mean and variance sampling, and finally the stylized image generator generates the corresponding stylized image based on the latent vector. In this way, the features extracted by performing multi-scale feature extraction on the image will express the image more accurately, which can improve the image quality of the generated stylized image.
[0142] Furthermore, this embodiment introduces the concept of spatial latent vector correction. For example, the S spatial latent vector (first spatial latent vector) is corrected to obtain the S+ spatial latent vector (second spatial latent vector) for generating a stylized image, so that the generated stylized image has a better ID (identity information) retention capability, and the generated stylized image is closer to the real feature data distribution of the input image, thereby further improving the image quality of the generated stylized image.
[0143] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of boxes in the block diagram and / or the flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0144] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0145] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0146] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0147] The above are only various implementations of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for generating a stylized image, characterized in that: Applied to an image processing device, the method comprises: Inputting the image to be converted into a variational autoencoder trained in advance, and processing the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted; Sampling the mean and variance of the latent vector of each feature dimension by the variational autoencoder to obtain a first spatial latent vector of the image to be converted in each feature dimension; Correcting the first spatial latent vector of the image to be converted in each feature dimension to obtain the second spatial latent vector of the image to be converted in each feature dimension as the target spatial latent vector; The target space latent vector is input into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
2. The method for generating a stylized image according to claim 1, characterized in that: The correction formula for correcting the first spatial latent vector of the image to be converted in each feature dimension is: Where s+ is the second spatial latent vector, s is the first spatial latent vector, x represents the input image, represents the variational autoencoder, represents the stylized image generator, the The stylized image generated by the stylized image generator and the image to be converted are consistent at the pixel level. represents the semantic feature loss of the image to be converted, F represents the VGG network, which is used to calculate the perceptual similarity of the image to be converted from the first space, and w vgg is the preset weight parameter, Represents the initial value of the second space latent vector and the first space latent vector The gap needs to be within the set range, representing Make small adjustments to the initial value and iterate a set number of times.
3. The method for generating a stylized image according to claim 1, characterized in that: The variational autoencoder includes a plurality of encoding layers cascaded in sequence, a plurality of decoding layers cascaded in sequence, and fully connected layers respectively connected to the decoding layers, each encoding layer is connected to a corresponding decoding layer, and each decoding layer is connected to one of the fully connected layers; The image to be converted is input from a first coding layer among the multiple coding layers, and each coding layer sequentially performs downscaling coding processing on the image to be converted to obtain a coding feature map and outputs the obtained coding feature map to the coding layer of the next layer and the decoding layer corresponding to the coding layer; The input of the first decoding layer is the output of its corresponding encoding layer, and the input of each other decoding layer is the output of the previous decoding layer plus the output of its corresponding encoding layer. Each decoding layer decodes its input data and outputs feature maps of different dimensions, and processes the feature maps through the fully connected layer to obtain the mean and variance of latent vectors of different feature dimensions corresponding to the image to be converted; Among the multiple decoding layers, starting from the first decoding layer, the dimensions of the feature maps output by each decoding layer decrease step by step.
4. The method for generating a stylized image according to claim 3, characterized in that: The method further comprises the step of training the variational autoencoder, the step comprising: Acquire a first training data set, where the first training data set includes a plurality of sample images; Inputting the sample images into the variational autoencoder in sequence to obtain a target space latent vector corresponding to the sample images; Inputting the target space latent vector corresponding to the sample image into the trained stylized image generator to obtain a stylized image corresponding to the sample image; A loss function of the variational autoencoder is calculated according to the stylized image and the sample image, and model parameters of the variational autoencoder are adjusted according to the loss function until a training convergence condition is met.
5. The method for generating a stylized image according to claim 4, characterized in that: The loss function includes pixel level loss, semantic similarity loss, and identity information loss. The expression of the loss function is as follows: ,in: , represents the pixel level loss, x is the sample image, represents the variational autoencoder, represents the trained stylized image generator, L 2 represents the image distance between the sample image and the stylized image generated by the stylized image generator; , represents the semantic similarity loss, L lpips Represents the semantic feature similarity between the sample image and the stylized image generated by the sample image after semantic feature extraction by the semantic extraction model; , representing the loss of identity information, L arc represents the similarity between identity features obtained by performing identity information recognition on the sample image and the stylized image generated from the sample image; Said w per , w id , w kl are respectively L per , L id , L arc Set the weight parameter.
6. The method for generating a stylized image according to any one of claims 1 to 5, characterized in that: The method further comprises the step of pre-training the stylized image generator, the step comprising: Acquire a stylized image dataset, where the stylized image dataset includes a plurality of stylized sample images; Inputting the stylized sample images into the stylized image generator to be trained in sequence, obtaining generated images corresponding to the stylized sample images, and calculating the loss function value of the stylized image generator; Optimizing the network parameters of the stylized image generator according to the loss function value until the calculated loss function value satisfies the training convergence condition, thereby obtaining a trained stylized image generator; The stylized image generator includes convolutional layers corresponding to different image resolutions respectively. When optimizing the network parameters of the stylized image generator, the network parameters of the convolutional layers corresponding to image resolutions smaller than the set resolution remain unchanged.
7. The method for generating a stylized image according to claim 6, characterized in that: The calculation formula of the loss function value of the stylized image generator is as follows: ,in: , ; represents the distribution of the stylized image dataset, represents the distribution of a data set consisting of stylized images generated by the stylized image generator according to each of the stylized sample images, D is the discriminator, Represents the discriminator gradient calculation operator for the stylized sample image.
8. A stylized image generation device, applied to an image processing device, characterized in that: The stylized image generating device comprises: A residual calculation module, used for inputting the image to be converted into a variational autoencoder trained in advance, and processing the image to be converted by the variational autoencoder to obtain the mean and variance of the latent vector of each feature dimension of the image to be converted; A latent vector processing module is used to sample the mean and variance of the latent vector of each feature dimension through the variational autoencoder to obtain the first spatial latent vector of the image to be converted in each feature dimension; correct the first spatial latent vector of the image to be converted in each feature dimension to obtain the second spatial latent vector of the image to be converted in each feature dimension as the target spatial latent vector; The image generation module is used to input the target space latent vector into a trained stylized image generator for image stylization conversion to generate a stylized image with a set image style.
9. An image processing device, characterized in that: The method comprises a machine-readable storage medium and one or more processors, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the one or more processors, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image processing method, device and equipment and storage medium
CN111583165A
Seal cutting work customized design generation device through utilizing generative adversarial network
CN112132916A