Artistic style migration method based on aesthetic consciousness adversarial learning network

By leveraging RPAD and ASA modules with VGG-19 network and attention mechanisms, the method enhances artistic style transfer by integrating artistic features and content preservation, resulting in more realistic and aesthetically pleasing outputs.

CN120318089APending Publication Date: 2025-07-15HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210397.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing style transfer methods lack artistic sense and insufficient retention of content features when implementing style transfer, and cannot effectively integrate the style model.

Method used

The art style transfer method based on aesthetic awareness adversarial learning network is adopted, and artistic features are learned using the residual pyramid art discriminator (RPAD), combined with the art style attention module (ASA) and the channel attention mechanism, enhance the fusion of content and style characteristics, remove unnecessary style information through the spatial attention mechanism, and use the pre-trained VGG-19 network for feature extraction and decoding.

Benefits of technology

The generated stylized images are closer to real artistic paintings, maintaining good content retention and artistic sense, improving the performance of the generator, and being able to generate more harmonious stylized results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318089A_ABST
    Figure CN120318089A_ABST
Patent Text Reader

Abstract

The invention discloses an artistic style migration method based on an aesthetic consciousness adversarial learning network, and relates to the field of computer vision and deep learning. The method comprises the following steps: 1, zooming and cutting a content image and a style image, and correspondingly extracting content features and style features from the content image and the style image; and 2, learning artistic features from the stylized image, carrying out artistic feature judgment on the generated stylized image, guiding the generator and improving the performance of the generator. 3, enhancing content features of the content image, and removing style information in the content image; fusing the style features of the style image with the artistic features, and enhancing the style features of the style image; and carrying out attention fusion on the enhanced content features and the style features to realize style migration. And 4, decoding the stylized features, and restoring the original size of the image. And 5, carrying out model training. And 6, outputting a stylized result. By applying the scheme, the technical problems of lack of artistic feeling and insufficient content feature retention during style migration in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and specifically to an artistic style transfer method based on an aesthetic awareness adversarial learning network. Background Art

[0002] The goal of style transfer is to apply the features of a style image to a content image while ensuring that the main information of the content image remains unchanged. The core difficulty of this task lies in effectively extracting style features from the style image and successfully integrating these style features into the content image. With the development of deep learning, the style transfer task has become increasingly mature and is widely used in artistic creation, film and games, augmented reality, cultural heritage protection, and medical images, etc.

[0003] Existing style transfer methods fall into two main categories: global statistics-based methods and local patch-based methods. When integrating style features, global statistics-based methods can lead to a significant loss of original style details; although local patch-based methods can retain the details of style features, local matching of deep features can produce discordant artifacts. These can all be summarized as insufficient integration of style patterns, that is, unrealistic aesthetics. Some later researchers have proposed many solutions to the aesthetic problem, but these methods either have difficulty effectively transmitting aesthetic information from the style, resulting in a lack of artistic sense, or insufficient retention of content features, and still cannot solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide an artistic style transfer method based on an aesthetic awareness adversarial learning network, which can solve the technical problems of the prior art in realizing style transfer, such as lack of artistic sense and insufficient retention of content features.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions.

[0006] An artistic style transfer method based on an aesthetic awareness adversarial learning network includes the following steps.

[0007] Step 1: Scale and crop the content image and the style image participating in the training, and use the pre-trained VGG-19 network to extract content features from the content image and style features from the style image.

[0008] Step 2: Use the Residual Pyramid Artistic Discriminator (RPAD) to learn artistic features from the style image, and judge the artistic features of the generated stylized image according to the learned artistic features. Guide the generator according to the judgment result and improve its performance.

[0009] Step 3: Use the Artistic Style Attention (ASA) module to enhance the content features of the content image through the spatial attention mechanism and remove the style information in the content image simultaneously.

[0010] Fuse the style features of the style image with the artistic features learned by the residual pyramid art discriminator through the channel attention mechanism to enhance the style features of the style image.

[0011] Perform attention fusion on the enhanced content features and style features to achieve style transfer and obtain stylized features.

[0012] Step 4: Decode the obtained stylized features to restore the original size of the image and obtain the stylized image.

[0013] Step 5: Conduct model training to complete the preset number of iterations.

[0014] Step 6: Output the stylized result according to the specified content image and style image.

[0015] Furthermore, in Step 1, scale the content image and style image participating in the training to a resolution of 512×512 and randomly crop them to a resolution of 256×256.

[0016] Furthermore, in Step 2, the residual pyramid art discriminator serves as an artistic feature extractor, downsamples the input style image by 1x, 2x, and 4x respectively, and uses the multi-scale residual discriminant block (Residual Dense Block, RDB) in the residual pyramid art discriminator to learn artistic features from the input image.

[0017] Furthermore, in Step 3, use the Whitening and Coloring Transform (WCT) module to remove style-related information such as colors in the content image.

[0018] Furthermore, in Step 4, use an encoder with a structure symmetric to the VGG-19 network structure, replace the downsampling in VGG-19 with upsampling, and restore the image to the same size as the original image.

[0019] Furthermore, in Step 5, during model training, conduct two-stage training, with 80,000 iterations in each stage, set the learning rate to 0.0001, use the Adam optimizer, and set the batch size to 2.

[0020] Furthermore, select natural scene images containing styles as the content images and artistic paintings as the style images.

[0021] Furthermore, the content images are selected from the Microsoft COCO2017 dataset, and the style images are selected from the Wikiart dataset.

[0022] After adopting the above technical solution, the present invention has the following beneficial effects: 1. In the present invention, the residual pyramid art discriminator is used to learn general art features from the style images, and the learned art features are used to guide the generator and improve the performance of the generator, so that works closer to real art paintings can be generated; 2. In the present invention, the art style attention module adaptively fuses content features, style features and art features through the spatial attention mechanism and the channel attention mechanism, and can better integrate the style patterns; 3. The present invention proposes an aesthetic-aware adversarial learning network for style transfer, which can achieve good content retention while maintaining the artistic sense, and can generate works closer to real art paintings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of the overall principle of the present invention.

[0024] Figure 2 is Figure 1 a schematic diagram of the overall principle of the art style attention (ASA) module in

[0025] Figure 3 is Figure 1 a schematic diagram of the structure of the residual pyramid art discriminator (RPAD) module in

[0026] Figure 4 is Figure 3 a schematic diagram of the structure of the residual discriminant block (RBD) in

[0027] Figure 5 is a comparison diagram of the visualization results in the embodiments of the present invention.

[0028] Figure 6 is a comparison diagram of the ablation experiment in the embodiments of the present invention.

[0029] Figure 7 is a quantitative comparison result diagram of the present invention and six currently popular style transfer methods. DETAILED DESCRIPTION OF THE INVENTION

[0030] In order to make the objectives, technical solutions and advantages of the present invention clearer, the features and performance of an art style transfer method based on an aesthetic awareness adversarial learning network in the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0031] Please refer to the attached Figures 1 to 4, an artistic style transfer method based on an aesthetic awareness adversarial learning network, includes the following steps.

[0032] Step 1, scale and crop the content image and the style image participating in the training, and use the pre-trained VGG-19 network to extract content features from the content image and style features from the style image.

[0033] When scaling and cropping, scale the content image and the style image participating in the training to a resolution of 512×512, and randomly crop them to a resolution of 256×256.

[0034] Step 2, use the Residual Pyramid Artistic Discriminator (RPAD) to learn artistic features from the style image, and judge the artistic features of the generated stylized image according to the learned artistic features. Guide the generator according to the judgment result and improve its performance.

[0035] When learning artistic features, the Residual Pyramid Artistic Discriminator is used as an artistic feature extractor to perform downsampling on the input style image by 1 time, 2 times, and 4 times respectively, and use the multi-scale residual discriminant block (Residual Dense Block, RDB) in the Residual Pyramid Artistic Discriminator to learn artistic features from the input image Step 3, use the Artistic Style Attention (ASA) module to enhance the content features of the content image through the spatial attention mechanism, and at the same time remove the style information in the content image.

[0036] When removing style information, use the Whitening and Coloring Transform (WCT) module to remove style-related information such as colors in the content image.

[0037] Fuse the style features of the style image with the artistic features learned by the Residual Pyramid Artistic Discriminator through the channel attention mechanism to enhance the style features of the style image.

[0038] Perform attention fusion on the enhanced content features and style features to achieve style transfer and obtain stylized features.

[0039] Step 4, decode the obtained stylized features to restore the original size of the image and obtain a stylized image.

[0040] When decoding, use an encoder with a structure symmetric to the VGG-19 network structure, replace the downsampling in VGG-19 with upsampling, and restore an image with the same size as the original image.

[0041] Step 5: Conduct model training to complete the preset number of iterations.

[0042] During model training, two-stage training is carried out. In each stage, 80,000 iterations are performed. The learning rate is set to 0.0001, the Adam optimizer is used, and the batch size is set to 2.

[0043] For the content images, natural scene images containing styles are selected from the Microsoft COCO2017 dataset; for the style images, artistic paintings are selected from the Wikiart dataset.

[0044] Step 6: Output the stylized result according to the specified content image and style image.

[0045] Specifically in implementation, the present invention proposes an artistic style transfer method based on the Aesthetic-Aware Adversarial Learning Network for Artistic Style Transfer (AAALNet), which specifically includes the following steps.

[0046] Step 1: Feature extraction.

[0047] First, the pictures participating in the training are scaled to a resolution of 512×512, then randomly cropped to a resolution of 256×256, and then the pre-trained VGG-19 network is used for feature extraction to extract content features and style features from the content image and the style image respectively.

[0048] Step 2: Learn artistic features and discriminate the stylized result.

[0049] An artistic painting that can be accepted and liked by most people must have features that are generally acceptable, such as harmonious patterns and textures. To make the stylized result present more harmonious patterns and textures, the present invention introduces the Residual Pyramid Artistic Discriminator (RPAD). On the one hand, as an artistic feature extractor, RPAD downsamples the input image by 1 time, 2 times, and 4 times respectively, and uses the multi-scale residual discriminant block (Residual DenseBlock, RDB) in RPAD to learn the generally acceptable artistic features from a large number of artistic paintings. On the other hand, RPAD also acts as a discriminator to judge whether the finally generated stylized result belongs to a real artistic painting, so as to guide the generator and improve its performance.

[0050] Specifically, the discriminator judges the authenticity of the images generated by the generator and produces a loss value (Ladv). This loss value reflects the gap between the generated images and the real images. The generator adjusts its parameters according to this loss value to generate images closer to the real images. Lc (content loss), by minimizing the content loss, ensures the content similarity between the generated image and the content image. Lca (attention-based content loss), by separately calculating the self-attention maps of the content image and the generated image and minimizing the distance between the two self-attention maps, further preserves the details in the content image. Ls (style loss), by minimizing the style loss, ensures the style similarity between the generated image and the style image.

[0051] Step 3, feature enhancement and fusion.

[0052] To achieve style transfer, the present invention introduces an Artistic Style Attention (ASA) module. ASA first enhances the content features through a spatial attention mechanism, where the Whitening and Coloring Transform (WCT) module is used to remove style-related information such as colors in the content image to avoid affecting the stylization result. Then ASA uses a channel attention mechanism to fuse the style features and artistic features to enhance the style features, so as to achieve aesthetically conscious style transfer. Finally, the enhanced content features and style features are fused through attention to obtain the final stylized features.

[0053] Step 4, decoding.

[0054] Since the obtained stylized features are just a series of high-dimensional vectors, a decoder is needed to decode these stylized features to restore the original size of the image. Since the pre-trained VGG-19 is used as the encoder, to ensure that the restored image has the same size as the original image, the structure of the decoder should be symmetric to that of the encoder, that is, the downsampling in VGG-19 is replaced by upsampling.

[0055] Step 5, perform model training.

[0056] In the initial stage of model training, the performance of the discriminator and the generator is not good enough, so the artistic features extracted by RPAD are not effective. Therefore, this solution adopts two-stage training. Both stages are iterated 80,000 times, the learning rate is set to 0.0001, the Adam optimizer is used, and the batch size is set to 2. The content images use the COCO2017 dataset from Microsoft, with approximately 118,000 images, which contains rich natural scenes. The style images come from the WikiArt dataset, with approximately 80,000 images, which contains a large number of artistic painting styles. The experiment is conducted using the Pytorch framework on an NVIDIA RTX2080Ti 11GB GPU.

[0057] Step Six: Output the stylized result.

[0058] The specific dataset introduction and experiments are as follows.

[0059] (1) Dataset The COCO dataset is a widely used dataset, with approximately 118,000 images. There are a total of 80 categories, which contain images of various real objects in a large number of daily life scenes.

[0060] The WikiArt dataset is an image dataset focusing on artistic styles and content, with approximately 80,000 images, which contain artworks from different historical periods and artistic styles, covering the works of a large number of famous painters, such as Van Gogh, Picasso, etc., involving 27 artistic styles and 45 artistic genres.

[0061] (2) Evaluation Metrics This invention uses content fidelity (CF), global effect (GE), and local pattern (LP) to measure the stylization quality of this invention.

[0062] Among them, CF measures the degree of retention of content features in the content image in the stylized result. GE measures the quality of the stylized result according to global effects such as global color and overall texture. LP measures the quality of the stylized result according to the diversity and similarity of local style patterns. Generally, the result of GE + LP is used to measure the stylization quality.

[0063] In addition, this invention uses the content features of the Relu4_1 and Relu5_1 layers in VGG-19 to calculate the content loss to further measure the degree of retention of content features. To further measure the aesthetic quality of the stylized image, this invention uses Neural Image Assessment (NIMA) to evaluate the quality of the stylized image. NIMA is an image quality assessment method without a reference image. It uses a CNN to predict the distribution law of human evaluation, which is highly correlated with human perception.

[0064] (3)Qualitative Evaluation Taking Figure 5 the second row in it as an example, the content picture (Content) is a photo of a person's head, and the style picture (Style) is a sketched head portrait on a white background.

[0065] Although AesUST and ArtFlow can well preserve the features of the content picture, the presented background obviously does not conform to the white background of the style picture. It should be noted that for the implementation of ArtFlow, the present invention uses the ArtFlow+AdaIN model.

[0066] AdaATTN and AesFA will introduce obvious disharmonious artifacts and patterns. For example, there are messy lines at the bridge of the nose in the result generated by AdaAttN, and there are disharmonious artifacts in the background of the result generated by AesFA. Moreover, AesFA also has deficiencies in content retention. For example, the lines at the bridge of the nose are directly missing.

[0067] Although TSSAT maintains a good content retention rate, it introduces disharmonious black shadows in the stylized result.

[0068] Although IECAST has good overall results, the strokes in the results it generates are relatively rough and do not transfer complex textures.

[0069] Figure 5 In the results shown in the third row, the several comparison methods all lack a certain degree of artistic sense and realism more or less. For example, AdaAttN has insufficient retention of the stamens in the petals, the result generated by TSSAT is hazy and unclear, and the strokes of IECAST for outlining the petals are too thick. In contrast, the AAALNet proposed by the present invention controls the degree of stylization just right while retaining the content, and is closer to the paintings created by real artists.

[0070] (4)Ablation Experiment The present invention conducts ablation experiments to verify the effectiveness of the proposed ASA module, RPAD module and the introduced loss function.

[0071] Figure 6 The results of the ablation experiment are shown.

[0072] It can be seen from the results in columns (c) and (d) that without RPAD and ASA, the overall quality of style transfer will decline, and some repeated grid-like spots will appear in the image, such as the scenery in the first row and the hair part of the person in the second row. In addition, the lines of the nose are messy in the part marked by the red box in the second row, and artifacts appear in the part marked by the yellow box.

[0073] As can be seen from the results in column (e), in the absence of the attention-based content loss Lca, the hand lines are missing at the positions marked by the yellow boxes in the second row.

[0074] In column (f), using all components, the resulting stylized image has finer lines, complete details retained, and is closer to a real work of art.

[0075] In addition, Figure 6 The quantitative metrics at the bottom also verify the above analysis from another aspect.

[0076] (5) Quantitative evaluation To further verify the effectiveness of the proposed AAALNet of the present invention, CF, GE+LP, Content loss, and NIMA are used to measure several comparison methods.

[0077] For the proposed AAALNet of the present invention and six other methods, 100 stylized images are generated using 10 content images and 10 style images respectively.

[0078] Figure 7 The average scores of CF, GE+LP, Content loss, and NIMA are shown. For CF, GE+LP, and NIMA, the higher the better, and for Content loss, the lower the better. The bold font indicates the best score, and the underlined font indicates the second-best score. It can be clearly observed that the proposed AAALNet of the present invention achieves the best results in terms of CF, Content loss, and NIMA, and also achieves the second-highest result in terms of GE+LP. This further indicates that the AAALNet of the present invention achieves good content retention while maintaining aesthetic features.

[0079] It should be noted that the parts not described in detail in this solution are all prior arts. The above embodiments are only used to illustrate the present invention, but the present invention is not limited to the above embodiments. Any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention all fall within the protection scope of the present invention.

Claims

1. An art style transfer method based on an aesthetic awareness adversarial learning network, characterized in that: It includes the following steps: Step 1: Scale and crop the content image and the style image participating in the training, and use the pre-trained VGG-19 network to extract content features from the content image and style features from the style image. Step 2: Use the residual pyramid art discriminator to learn art features from the style image, and judge the art features of the generated stylized image according to the learned art features. Guide the generator according to the judgment result and improve its performance. Step 3: Use the art style attention module to enhance the content features of the content image through the spatial attention mechanism, and at the same time remove the style information in the content image. Fuse the style features of the style image with the art features learned by the residual pyramid art discriminator through the channel attention mechanism to enhance the style features of the style image. Perform attention fusion on the enhanced content features and style features to achieve style transfer and obtain stylized features. Step 4: Decode the obtained stylized features to restore the original size of the image and obtain a stylized image. Step 5: Conduct model training and complete a preset number of iterations. Step 6: Output the stylized result according to the specified content image and style image.

2. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, characterized in that: In Step 1, the content image and style image participating in the training are scaled to a resolution of 512×512 and randomly cropped to a resolution of 256×256.

3. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, wherein: In Step 2, the residual pyramid art discriminator is used as an art feature extractor to perform downsampling on the input style image by 1 time, 2 times, and 4 times respectively, and use the multi-scale residual discriminant block in the residual pyramid art discriminator to learn art features from the input image.

4. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, wherein: In Step 3, a whitening and coloring transformation module is used to remove style-related information such as colors in the content image.

5. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, characterized in that: In Step 4, an encoder with a structure symmetric to the VGG-19 network structure is used to replace the downsampling in VGG-19 with upsampling to restore an image with the same size as the original image.

6. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, wherein: In Step 5, two-stage training is performed during model training. Each stage is iterated 80,000 times, the learning rate is set to 0.0001, the Adam optimizer is used, and the batch size is set to 2.

7. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 1, characterized in that: The content image is selected as a natural scene image containing style, and the style image is selected as an art painting.

8. The artistic style transfer method based on the aesthetic awareness adversarial learning network according to claim 7, wherein: The content image is selected from the Microsoft COCO2017 dataset, and the style image is selected from the Wikiart dataset.