An Image Animation Method Based on Cascade Generative Adversarial Network

Through cascading generation of adversarial networks and Vgg16 network feature extraction, the problem of long and high cost of image style transfer in deep learning methods is solved, and efficient image animation style conversion is achieved.

CN114140317BActive Publication Date: 2025-07-29HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111446222.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-07-29
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The existing deep learning methods are time-consuming, resource-cost and cumbersome in image style transfer, making it difficult to efficiently complete image style conversion, especially animation style transfer.

Method used

The cascade generation adversarial network is adopted. Through the adversarial training of the generator and the discriminator, the generator is optimized layer by layer to generate high-reality reconstructed pictures, and combined with the Vgg16 network to extract features to realize the migration of reference picture style features.

Benefits of technology

It realizes the low-cost and efficient conversion of target images into animation style, improving the efficiency and effect of image animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140317B_ABST
    Figure CN114140317B_ABST
Patent Text Reader

Abstract

The present invention discloses an image animation method based on a cascaded generative adversarial network. First, preprocess the target image and the reference image respectively to obtain a set of low-resolution images; then construct a reference image generative adversarial network and iteratively train the reference image generative adversarial network multiple times; then input the downsampled target image into the trained reference image generative adversarial network model to learn the style features of the reference image; at the same time, in each sub-generative adversarial network layer, design an additional discriminator to obtain a target image generative adversarial network, and repeatedly iteratively train each sub-generative adversarial network of the target image generative adversarial network model until the training ends, and the generator generates a target image with the anime style of the reference image. Compared with traditional methods, the present invention not only has low cost and high efficiency, but also can better complete the conversion of the anime style of the given target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and specifically provides an image cartoonization method based on a cascaded generative adversarial network. Background Art

[0002] As an important part of the field of image processing, image style transfer has received extensive attention and research. Image style transfer means that for each given input image, the image will have a specific artistic style. Through the image transfer method, the artistic style of the image can be converted on the premise of keeping the main content of the image unchanged, so as to present different visual effects.

[0003] Focusing on the problem of image style transfer, many image style transfer methods have been developed. The traditional approach is to analyze and establish a model of a specific type of image style for a set of images with a certain specific style. When an input target image whose style needs to be changed is input, the image is continuously modified to adapt to the image style model. However, this method can only complete the transfer of a specific type of image style under the condition of consuming a large amount of time and resource costs. With the rapid development of the deep learning field, currently, deep learning-based image style transfer has received extensive attention and research. The deep learning-based method mainly completes image processing by designing relevant neural network models. Usually, a convolutional neural network module is used in the network model to extract features of the image, and operations such as data normalization and non-linear transformation are accompanied to further process the feature data. When a target image and a reference image providing other image styles are input into the network model, the model can effectively extract the content and style features of the input image and the reference image by coupling multiple layers of neural network modules, and design a relevant optimization function according to relevant constraint conditions and the error between the processing result and the ideal result to optimize the processing result, and transfer the style features of the reference image to the target image. Through multiple iterations of optimization, the performance of the network model is continuously improved, and finally the effect of target image style transfer is achieved. However, commonly used deep learning methods usually require a dataset containing a large number of images for training, which takes a long time, and sometimes they are also limited by pre-trained models, and the training model is relatively cumbersome and complex, and the algorithm efficiency needs to be improved. Summary of the Invention

[0004] In view of the problems existing in the existing methods, the present invention proposes an image cartoonization method based on a cascaded generative adversarial network. Compared with other deep learning models, the generative adversarial network has unique advantages in fitting the data features of the input pictures. The input of this network model includes the target picture and the reference cartoonized picture. The model mainly includes two sub-networks, namely the generator G (Generator) and the discriminator D (Discriminator). Among them, the generator G is used to generate a highly realistic reconstructed picture. Its input is random noise. The generator network processes and maps the input noise, fits the data features of the reference cartoonized picture, and finally generates a reconstructed picture with a high similarity to the reference picture through multiple iterations of training, so that the discriminator determines it as a real picture. The goal of the discriminator is to optimize the network parameters through training and accurately distinguish whether the input picture is a real picture or a reconstructed picture generated by the generator. In the "adversarial" training of the generator G and the discriminator D, the performance of both modules is continuously optimized. After the training is completed, the generator can well fit the feature information of the reference picture. When the content of the input target picture and the reference picture is similar (such as both are portrait pictures, landscape pictures, etc.), the style features of the reference picture can be transferred to the input target picture. To achieve the above goals, the main solutions and implementation steps of the present invention are as follows:

[0005] Step 1: Preprocess the real target picture and the reference picture respectively to obtain a low-resolution picture set of the target picture and the reference picture.

[0006] For a given target picture I real_d and the provided reference picture I real_s with cartoon style, perform downsampling on the two pictures respectively according to the same downsampling parameter α until the size of the pictures is 1 / 10 of the original pictures, and a low-resolution picture set of the target picture and the reference picture is obtained respectively.

[0007] Step 2: Construct a reference picture generative adversarial network;

[0008] Arrange the low-resolution picture set of the reference picture in the order from the original resolution picture to the lowest resolution, and design a layer of sub-generative adversarial network for each picture.

[0009] Step 3: Start training the reference picture generative adversarial network from the sub-generative adversarial network of the lowest resolution layer. Input random noise and generate a reconstructed picture through the generator. Input the reconstructed picture and the low-resolution reference picture into the discriminator in turn. For the reference picture, the expected output value of the discriminator is as close to 1 as possible, otherwise it is expected to be close to 0.

[0010] Step 4: Iteratively train the sub-generative adversarial network of the lowest resolution layer of the reference image generative adversarial network multiple times. Under the constraint of the optimization function, continuously improve the authenticity of the reconstructed image generated by the generator to make it approximate the reference image.

[0011] Step 5: Upsample the reconstructed image generated by the lowest resolution layer, add random noise, and input it into the sub-generative adversarial network of a higher resolution layer. Iteratively train it according to the method of training the sub-generative adversarial network of the lowest resolution layer to generate a reconstructed image with a higher resolution. Keep training until the highest layer sub-generative adversarial network to generate a realistic reconstructed image with the same original resolution as the reference image.

[0012] Step 6: Input the downsampled target image into the trained reference image generative adversarial network model to learn the style features of the reference image. At the same time, in each layer of the sub-generative adversarial network, design an additional discriminator to obtain the target image generative adversarial network and further maintain the content features of the target image.

[0013] Step 7: Repeatedly and iteratively train each layer of the sub-generative adversarial network of the target image generative adversarial network model until the training ends, and the generator generates a target image with the anime style of the reference image.

[0014] Further, the specific method of Step 2 is as follows:

[0015] Arrange the low-resolution image set of the reference image in the order from the original resolution image to the lowest resolution, and design a layer of sub-generative adversarial network for each image.

[0016] The generative adversarial network includes a generator and a discriminator. The generator is composed of convolutional blocks consisting of multiple convolutional kernels-normalization-nonlinear transformation. Each convolutional block has two types of convolutional kernels with sizes of 5*5 and 3*3, which are used to extract large and small feature blocks respectively. By combining the extracted feature blocks, the input features are fully utilized to improve the feature extraction effect. The feature blocks are input into the next convolutional block. The generative adversarial network structure of each layer is the same, only the parameter scale of the network is different. The generator can learn the features of the reference image through the network model and the optimization function. The discriminator D is composed of five identical convolutional blocks, and each layer contains three parts: convolution-normalization-nonlinear activation function.

[0017] The generator is composed of convolutional blocks consisting of multiple convolutional kernels-normalization-nonlinear transformation. Each convolutional block has two types of convolutional kernels with sizes of 5*5 and 3*3, which are used to extract large and small feature blocks respectively. By combining the extracted feature blocks, the input features are fully utilized to improve the feature extraction effect. The feature blocks are input into the next convolutional block. The discriminator D is composed of five convolutional blocks, and each layer contains three parts: convolution-normalization-nonlinear activation function, which is used for feature extraction.

[0018]

[0019] The convolution kernel size of the discriminator D convolutional block is 3*3. Except for the fifth layer which uses the sigmoid function (the output is between 0 and 1), the activation function used in the other layers is the f leaky (x acti ) function. For the generator G, the convolutional blocks of the first layer and the fifth layer are consistent with the discriminator structure. In the convolutional blocks of the middle three layers, for the input feature matrix x fea , the convolutional kernels of sizes 3*3 and 5*5 are used simultaneously to extract small and large feature blocks of the picture, and then the convolution results are concatenated according to the channel dimension, as shown in formulas (2)-(3). Among them, x w*d*c1 represents that its picture dimension is w*d*c1, conv i*i represents the convolutional block with a convolutional kernel of i*i, and O(x fea ) represents the output feature.

[0020] cat(x w*d*c1 , x w*d*c2 ) = x w,d,c1+c2 (2)

[0021] O(x fea ) = conv 1*1 (x fea + cat(conv 3*3 (x fea ), conv 5*5 (x fea ))) (3)

[0022] Furthermore, the specific method of step 4 is as follows:

[0023] Iteratively train the lowest-resolution layer sub-generative adversarial network of the reference picture generative adversarial network multiple times. Under the constraint of the optimization function, continuously improve the authenticity of the reconstructed picture generated by the generator to make it approximate the reference picture. During the iterative training process, the Vgg16 network model is used to extract the features of the reconstructed / reference picture, and these features are used as part of the input of the optimization function to improve the training effect of the sub-generative adversarial network.

[0024] As shown in formula (4), under the constraint of the optimization function, through multiple iterations, improve the authenticity of the generator's reconstructed picture I real_n,j .

[0025] f loss_s = L GP (G j , D j ) + 2L rec (G j ) + Lstyle (I real_n,j ,I real_s,j ) (4)

[0026]

[0027]

[0028] where z represents random noise, and the subscript j represents the j-th layer of the sub-generative adversarial network, 0 ≤ j ≤ J, L GP represents the WGAN-GP loss function, and L rec is the reconstruction loss. L style represents a common style loss function. For formula (6), gr(I) represents calculating the Gram matrix of the features of the input image I extracted using the Vgg16 network. If the size of the feature map is smaller than the size requirement of the Vgg16 network, the feature map is copied and padded and extended with pixel value 1. M l represents the width / height of the feature map, and N l represents the number of convolution kernels of this convolutional layer.

[0029] Furthermore, the specific method of step 6 is as follows:

[0030] After the reference image generative adversarial network model is trained, the downsampled target image is input into the trained reference image generative adversarial network model to learn the style features of the reference image. At the same time, in each layer of the sub-generative adversarial network, an additional discriminator D ex is designed. The structure of this discriminator is the same as that of the discriminator in the reference image generative adversarial network, and the input is the content features of the reconstructed image and the target image. Its function is to discriminate whether the input features are the features of the target image. The acquisition of the input features is shown in formula (7):

[0031] D ex (x) = cat(vg i (X)) (7)

[0032] where X represents the input image, vg i (X) represents the features obtained after the input image passes through the i-th (1 ≤ i ≤ 5) convolutional model layer of Vgg16, and cat represents concatenating the output features of different layers i in the channel dimension. Through multiple iterations of training, when the input target image learns the style features of the reference image, the content features of the target image are further maintained.

[0033] The beneficial effects of the present invention are as follows:

[0034] The present invention proposes an image cartoonization method based on a cascaded generative adversarial network. Compared with traditional methods, it not only has low cost and high efficiency, but also can better complete the conversion of the given target picture into a cartoon style through this method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a reference picture generation adversarial network model for an embodiment of the present invention;

[0036] Figure 2 It is a target picture generation adversarial network model for an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The method of the present invention will be further described below in conjunction with the drawings and embodiments.

[0038] An image cartoonization method based on a cascaded generative adversarial network comprises the following steps:

[0039] Step 1: Image preprocessing. For a given target picture I real_d and a reference picture I providing a cartoon style real_s , the dimensions of the pictures are all L*W*D, where W = H = 256, representing the width and height of the image respectively, and D = 3 represents the picture channel dimension. The target picture and the reference picture are respectively downsampled continuously according to the same downsampling parameter α, where α ∈ (0.6, 0.8), until the size of the picture is about 1 / 10 of the original picture, and a low-resolution picture set I real_d,j and I real_s,j can be obtained respectively, where 0 ≤ j ≤ 8, and the smaller j is, the higher the resolution of the input picture of the sub-generative adversarial network layer it is located in.

[0040] Step 2: Design a reference picture generation adversarial network (as Figure 1 shown). Starting from the lowest-resolution subgraph of the reference picture, each subgraph corresponds to a sub-generative adversarial network layer, and each layer contains two parts: a generator and a discriminator.

[0041] The generator is composed of convolutional blocks consisting of multiple convolutional kernels-normalization-nonlinear transformation. Each convolutional block has two types of convolutional kernels with sizes of 5*5 and 3*3, which are used to extract large and small feature blocks respectively. By combining the extracted feature blocks, the input features are fully utilized to improve the feature extraction effect. The feature blocks are input into the next convolutional block. The discriminator D is composed of five convolutional blocks, and each layer contains three parts: convolution-normalization-nonlinear activation function, which is used for feature extraction.

[0042]

[0043] The convolution kernel size is 3*3. Except for the fifth layer which uses the sigmoid function (the output is between 0 and 1), the activation function used in the other layers is the f leaky (x acti ) function. For the generator G, the convolution blocks in the first and fifth layers are consistent with the discriminator structure. In the convolution blocks of the middle three layers, for the input feature matrix x fea , convolution kernels of sizes 3*3 and 5*5 are used simultaneously to extract small and large feature blocks of the picture, and then the convolution results are concatenated according to the channel dimension, as shown in formulas (2)-(3). Among them, x w*d*c1 represents that its picture dimension is w*d*c1, conv i*i represents the convolution block with the convolution kernel of i*i, and O(x fea ) represents the output feature.

[0044] cat(x w*d*c1 , x w*d*c2 ) = x w,d,c1+c2 (2)

[0045] O(x fea ) = Conv 1*1 (x fea + cat(conv 3*3 (x fea ), conv 5*5 (x fea ))) (3)

[0046] Step 3: Start training from the sub-generative adversarial network layer with the lowest resolution. The input of this layer is random noise. The generator processes it to generate a reconstructed picture. The generator expects the reconstructed picture to be highly realistic so that when it is input into the discriminator, the expected output value is close to 1. The reconstructed picture and the reference picture are sequentially input into the discriminator. For the reference picture, the discriminator expects the output value to be as close to 1 as possible, otherwise it expects to be close to 0.

[0047] Step 4: Iteratively train the sub-generative adversarial network of the lowest resolution layer of the reference picture generative adversarial network multiple times. Under the constraint of the optimization function, continuously improve the realism of the reconstructed picture generated by the generator to make it approximate the reference picture. During the iterative training process, the Vgg16 network model is used to extract the features of the reconstructed / reference pictures, and these features are used as part of the input of the optimization function to improve the training effect of the sub-generative adversarial network.

[0048] As shown in formula (4), under the constraint of the optimization function, through multiple iterations, improve the realism of the reconstructed picture I real_n,j of the generator.

[0049] f loss_s = L GP (G j , Dj ) + 2L rec (G j ) + L style (I real_n,j ,I real_s,j ) (4)

[0050]

[0051]

[0052] where \(z\) represents random noise, the subscript \(j\) represents the \(j\)-th layer of the sub-generative adversarial network, and \(L\) GP represents the WGAN-GP loss function, and \(L\) rec is the reconstruction loss. \(L\) style represents the common style loss function. For formula (6), \(gr(I)\) represents calculating the Gram matrix of the features of the input image \(I\) extracted using the Vgg16 network. If the size of the feature map is smaller than the size requirement of the Vgg16 network, the feature map is replicated and padded and extended with pixel value 1. \(M\) l represents the width / height of the feature map, and \(N\) l represents the number of convolution kernels in this convolutional layer.

[0053] Step 5: According to Steps 3 - 4, starting from the lowest resolution layer, after training is completed, upsample the reconstructed image, add random noise, and input it into the sub-generative adversarial network at a higher resolution layer. Iteratively train the sub-generative adversarial network at a higher resolution layer in the same way as training the sub-generative adversarial network at the lowest resolution layer until the training process of the overall network model is completed, and obtain a realistic reconstructed image with the same resolution as the original reference image.

[0054] Step 6: The content of the target image is roughly the same as that of the reference image. After the reference image generative adversarial network model training is completed, input the downsampled target image into the trained reference image generative adversarial network model to learn the style features of the reference image. At the same time, in each layer of the sub-generative adversarial network, design an additional discriminator \(D\) ex , the structure of this discriminator is the same as that of the discriminator in the reference image generative adversarial network, and the input is the content features of the reconstructed image and the target image. Its function is to discriminate whether the input features are the features of the target image (if so, the expected output value is 1, otherwise the expected output value is 0), and obtain the target image generative adversarial network model (as Figure 2 shown). The acquisition of the input features is as shown in formula (7):

[0055] D ex (x) = cat(vg i (X)) (7)

[0056] where \(X\) represents the input image, and \(vg\)i (X) represents the feature obtained after the input image passes through the i-th (1 ≤ i ≤ 5) convolutional model layer of Vgg16, and cat represents the channel dimension concatenation of the output features of different layers i, as shown in Equation (2). Through multiple iterative trainings, the input target image further maintains the content features of the target image while learning the style features of the reference image.

[0057] Step 7: Repeatedly iterate and train each sub-generative adversarial network of the target image generative adversarial network model multiple times until the training ends, and the generator generates a target image with the anime style of the reference image.

Claims

1. An image animation method based on a cascaded generative adversarial network, characterized in that, The steps are as follows: Step 1: Preprocess the real target image and the reference image respectively to obtain a set of low-resolution images of the target image and the reference image; For a given target image I real_d and a provided reference image I with anime style real_s , the two images are respectively downsampled continuously according to the same downsampling parameter α until the size of the images is close to 1 / 10 of the original images, and a low-resolution image set for the target image and the reference image is obtained respectively; Step 2: Construct a reference image generation adversarial network; Arrange the set of low-resolution images of the reference image in the order from the original resolution image to the lowest resolution, and design a sub-generation adversarial network for each image; Step 3: Start training the reference image generation adversarial network from the sub-generation adversarial network of the lowest resolution layer. Input random noise, and generate a reconstructed image through the generator; sequentially input the reconstructed image and the low-resolution reference image of the corresponding level into the discriminator; For the reference image, the expected output value of the discriminator is as close to 1 as possible, otherwise it is expected to be close to 0; Step 4: Iteratively train the sub-generation adversarial network of the lowest resolution layer of the reference image generation adversarial network multiple times; Under the constraint of the optimization function, continuously improve the authenticity of the reconstructed image generated by the generator to make it approximate the reference image; Step 5: Upsample the reconstructed image generated by the lowest resolution layer, add random noise and input it into the sub-generation adversarial network of a higher resolution layer, and iteratively train according to the method of training the sub-generation adversarial network of the lowest resolution layer to generate a reconstructed image with higher resolution; Keep training until the highest layer sub-generation adversarial network to generate a realistic reconstructed image with the same original resolution as the reference image; Step 6: Input the downsampled target image into the trained reference image generation adversarial network model to learn the style features of the reference image; At the same time, in each sub-generation adversarial network, design an additional discriminator to obtain the target image generation adversarial network and further maintain the content features of the target image; The specific method of Step 6 is as follows: After the reference image generation adversarial network model is trained, input the downsampled target image into the trained reference image generation adversarial network model to learn the style features of the reference image; Meanwhile, in each layer of the sub-generative adversarial network, an additional discriminator D is designed ex , and the structure of this discriminator is the same as that of the discriminator in the reference image generation adversarial network. The input is the content features of the reconstructed image and the target image, and its function is to determine whether the input features are the features of the target image; The acquisition of the input feature is shown in formula (7); D ex (x) = cat(vg i (X)) (7) Among them, X represents the input image, and vg i (X) represents the feature obtained after the input image passes through the i-th (1 ≤ i ≤ 5) convolutional model layer of Vgg16. Cat represents the channel dimension splicing of the output features of different layers i. Through multiple iterative trainings, when the input target image learns the style features of the reference image, the content features of the target image are further maintained. Step 7: Repeatedly and iteratively train each sub-generation adversarial network of the target image generation adversarial network model until the training ends, and the generator generates a target image with the anime style of the reference image.

2. The image cartoonization method based on a cascaded generative adversarial network according to claim 1, characterized in that, The specific method of Step 2 is as follows: Arrange the set of low-resolution images of the reference image in the order from the original resolution image to the lowest resolution, and design a sub-generation adversarial network for each image; The described generation adversarial network includes a generator and a discriminator. The generator is composed of convolutional blocks composed of multiple convolutional kernels-normalization-nonlinear transformation. Each convolutional block has two types of convolutional kernels with sizes of 5*5 and 3*3. The 5*5 is used to extract large feature blocks, and the 3*3 is used to extract small feature blocks. By combining the extracted feature blocks, the input features are fully utilized to improve the feature extraction effect; input the feature blocks into the next convolutional block; the structure of each layer of the generation adversarial network is the same, only the parameter scale of the network is different; among them, the generator can learn the features of the reference image through the network model and the optimization function. The discriminator D is composed of five identical convolutional blocks, and each layer contains three parts: convolution-normalization-nonlinear activation function; The convolution kernel size of the discriminator D convolution block is 3*3. Except for the fifth layer which uses the sigmoid function, the activation function used in the remaining layers is the f leaky (x acti ) function; for the generator G, the convolution blocks of the first and fifth layers are consistent with the discriminator structure. In the convolution blocks of the middle three layers, for the input feature matrix x fea , both 3*3 and 5*5 convolution kernels are used to extract small and large feature blocks of the image, and then the convolution results are concatenated according to the channel dimension, as shown in formulas (2)-(3); where, x w*d*c1 represents that the image dimension is w*d*c1, conv i*i represents a convolution block with a convolution kernel of i*i, and O(x fea ) represents the output feature; cat(x w*d*c1 ,x w*d*c2 )=x w,d,c1+c2 (2) O(x fea ) = conv 1*1 (x fea + cat(conv 3*3 (x fea ), conv 5*5 (x fea ))) (3).

3. The image anime-style conversion method based on a cascaded generative adversarial network according to claim 2, wherein The specific method of Step 4 is as follows: Iteratively train the lowest-resolution layer sub-generative adversarial network of the reference image generation adversarial network multiple times; under the constraint of the optimization function, continuously improve the authenticity of the reconstructed images generated by the generator to make it approximate the reference image; During the iterative training process, use the Vgg16 network model to extract the features of the reconstructed / reference images, and these features are used as part of the input of the optimization function to improve the training effect of the sub-generative adversarial network; As shown in formula (4), under the constraints of the optimization function, the authenticity of the generated reconstructed image I is improved through multiple iterations. real_n,j ; f loss_s = L GP (G j , D j ) + 2L rec (G j ) + L style (I real_n,j , I real_s,j ) (4) where z represents random noise, the subscript j represents the j-th layer of the sub-generative adversarial network, 0 ≤ j ≤ J, L GP represents the WGAN-GP loss function, L rec is the reconstruction loss; L style represents the common style loss function; for formula (6), gr(I) represents calculating the Gram matrix of the features of the input image I extracted using the Vgg16 network. If the size of the feature map is smaller than the size requirement of the Vgg16 network, the feature map is copied and padded and extended with pixel value 1; M l represents the width / height of the feature map, N l represents the number of convolutional kernels of this convolutional layer.

Citation Information

Patent Citations

  • Image super-division method based on progressive residual generative adversarial network

    CN114140323A