Image style migration method and device, equipment, storage medium and program product
By recoding the style image features and fusing them with the content image features, the problem of requiring a large amount of sample data in the prior art to improve the learning effect of the style transfer model is solved, and an efficient and low-cost style transfer effect is achieved.
Patent Information
- Application Number
- CN202510120958.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the training effect and training speed of the style transfer model have improved, but a large amount of sample data is still needed to improve the learning effect of generating adversarial networks, which increases the difficulty of network implementation.
By recoding the style image features and fusing them with the content image features, a style transfer image is generated without increasing the number of samples to augment the style feature space.
Without increasing the number of samples, the style feature space is expanded, the utilization efficiency of image features is improved, the training cost is reduced, and the high-quality style transfer effect is achieved.
Smart Images

Figure CN120047309A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to an image style transfer method, device, equipment, storage medium, and program product. Background Art
[0002] In recent years, with the development of dedicated computing hardware and the enhancement of computing power, both the training effect and training speed of style transfer models have made great progress, and the application fields of style transfer have become more extensive.
[0003] In related technologies, a Generative Adversarial Network (GAN) extracts the data distribution from a known domain through adversarial learning between a generator and a discriminator based on a known data domain. The effect of its style transfer depends on the quantity and quality of the known domains, and a large amount of sample data needs to be collected to improve its learning effect, which undoubtedly increases the difficulty of network implementation. Summary of the Invention
[0004] To solve the above technical problems existing in related technologies, embodiments of this application propose an image style transfer method, device, equipment, storage medium, and program product.
[0005] Embodiments of this application provide an image style transfer method, and the method includes:
[0006] Extract content image features and style image features from the image to be processed, and obtain first semantic information corresponding to the content image features and second semantic information corresponding to the style image features;
[0007] Preprocess the content image features and the style image features according to the first semantic information and the second semantic information to obtain preprocessed content image features and preprocessed style image features;
[0008] Recode the preprocessed style image features to obtain recoded style image features;
[0009] Fuse the preprocessed content image features and the recoded style image features to obtain a fused feature;
[0010] Decode the fused feature to obtain a style transfer image.
[0011] In some embodiments, preprocessing the content image features and the style image features according to the first semantic information and the second semantic information to obtain preprocessed content image features and preprocessed style image features includes: performing semantic alignment on the first semantic information and the second semantic information to obtain aligned semantic information; performing pre-fusion on the content image features and the style image features according to the aligned semantic information to obtain pre-fused features; and extracting content features and image features from the pre-fused features to obtain the preprocessed content image features and the preprocessed style image features.
[0012] In some embodiments, after obtaining the preprocessed content image features and the preprocessed style image features, the method further includes: obtaining third semantic information corresponding to the preprocessed content image features and fourth semantic information corresponding to the preprocessed style image features; and fusing the preprocessed content image features and the re-encoded style image features to obtain a fused feature, including: fusing the preprocessed content image features and the re-encoded style image features according to the third semantic information and the fourth semantic information to obtain a fused feature.
[0013] In some embodiments, re-encoding the preprocessed style image features to obtain re-encoded style image features includes: using a variational autoencoder network to re-encode the preprocessed style image features to obtain re-encoded style image features.
[0014] In some embodiments, extracting content image features and style image features from a to-be-processed image and obtaining first semantic information corresponding to the content image features and second semantic information corresponding to the style image features includes: using a plurality of content feature extraction layers to extract features from the to-be-processed image to obtain output features of each content feature extraction layer in the plurality of content feature extraction layers, and obtaining semantic features having the same size as the output features of each content feature extraction layer; the content image features include the output features of each content feature extraction layer, and the first semantic information includes semantic features corresponding to the output features of each content feature extraction layer; using a plurality of style feature extraction layers to extract features from the to-be-processed image to obtain output features of each style feature extraction layer in the plurality of style feature extraction layers, and obtaining semantic features having the same size as the output features of each style feature extraction layer; the style image features include the output features of each style feature extraction layer, and the second semantic information includes semantic features corresponding to the output features of each style feature extraction layer.
[0015] In some embodiments, decoding the fused feature to obtain a style transfer image includes: sequentially processing the fused feature by a plurality of upsampling modules to obtain the style transfer image; the method further includes: generating semantic information for evaluating the image transfer quality of the style transfer image according to the output features of each upsampling module in the plurality of upsampling modules.
[0016] An embodiment of the present application further provides an image style transfer device, where the device includes:
[0017] A first processing module, configured to extract a content image feature and a style image feature from an image to be processed, and obtain first semantic information corresponding to the content image feature and second semantic information corresponding to the style image feature;
[0018] A second processing module, configured to preprocess the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain a preprocessed content image feature and a preprocessed style image feature; re-encode the preprocessed style image feature to obtain a re-encoded style image feature; and fuse the preprocessed content image feature and the re-encoded style image feature to obtain a fused feature;
[0019] A decoding module, configured to decode the fused feature to obtain a style transfer image.
[0020] An embodiment of the present application further provides an electronic device, where the electronic device includes a processor and a memory for storing a computer program that can run on the processor; wherein, the processor is configured to run the computer program to execute any one of the above image style transfer methods.
[0021] An embodiment of the present application further provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above image style transfer methods.
[0022] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any one of the above image style transfer methods.
[0023] It can be seen that the embodiments of the present application can augment the style feature space without increasing the number of samples by re-encoding the style image features, and there is no need to collect a large amount of sample data to improve the learning effect of the generative adversarial network, thereby improving the utilization efficiency of image features. Description of the Drawings
[0024] Figure 1 is a flowchart of an image style transfer method according to an embodiment of the present application;
[0025] Figure 2 It is a flowchart for implementing image style transfer and network loss calculation provided by an embodiment of the present application;
[0026] Figure 3 It is a schematic structural diagram of an upsampling module provided by an embodiment of the present application;
[0027] Figure 4 It is a flowchart of another image style transfer method provided by an embodiment of the present application;
[0028] Figure 5 It is a schematic structural diagram of an image style transfer device provided by an embodiment of the present application;
[0029] Figure 6 It is a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present application. Specific implementation manners
[0030] With the development of dedicated artificial intelligence computing chips and artificial intelligence algorithms, some artificial intelligence applications that require huge calculations to verify their feasibility have become feasible. One important application is style transfer. As a long-popular image processing direction, style transfer has two major goals: retaining content features and transferring style features. It can generate an image with the style of the style image and the content of the original image by extracting and fusing features in the original image and the style image, and has many important applications in fields such as data augmentation, information protection, image translation, and artistic creation.
[0031] Traditional style transfer networks can only extract the global data style from the source domain and then attach the global style to the target image domain. In some cases, this global style cannot describe the style of the image, and rashly transferring it will cause the effect of transferring the image filter, and cannot achieve a style transfer that includes details. The related technology has proposed the very amazing Generative Adversarial Network StyleGAN, but the network structure is very complex and the training cost is very high, and it requires a supercomputer to train, which is obviously unaffordable for individual developers.
[0032] Generative adversarial networks are based on known data domains. Through the adversarial learning of generators and discriminators, they extract data distributions from the known domains. The effect of style transfer depends on the quantity and quality of the known domains. A large number of known-domain images need to be collected to improve the learning effect, which undoubtedly increases the difficulty of network implementation. In recent years, generative adversarial networks have developed rapidly, and a number of variants of generative adversarial networks have emerged, integrating a series of excellent algorithmic ideas into generative adversarial networks, giving new life to traditional generative adversarial networks. For example, StyleGAN integrates multiple algorithmic ideas, achieving amazing image generation. The CycleGAN (Cycle Generative Adversarial Network) realizes efficient multi-domain mutual conversion by cleverly adjusting the structure of the GAN network and using cycle transformation consistency.
[0033] In related technologies, the goal of style transfer is no longer limited to cross-domain style conversion. This goal can also be the diversity of transferred styles. Related technologies have proposed style transfer networks that can achieve arbitrary transfer of image styles. For example, the Adaptive Instance Normalization (AdaIN) network based on a forward neural network has been proposed in related technologies. By calculating the features of the content image and the style image and fusing them, this network can achieve arbitrary style transfer. The principle of achieving arbitrary transfer of image styles is as follows: Through the learning and layering of the representations within the image, the content features and style features of the image are automatically separated, and then recombined to form a new-style image. By recombining the style contents of different-style images to generate new images, many new image domains can be produced. For example, by migrating a photo of a building to its elevation view, it can assist architects in architectural design modeling; by migrating an interior decoration plan to an undecorated room, it can help interior designers with interior decoration; by migrating the pathological images of healthy biological tissues to the pathological images of diseased biological tissues, a new dataset for medical diagnosis can be produced.
[0034] When training traditional machine learning models, a large amount of data is usually required to extract data distributions. When the dataset is very small, training is often hindered. For example, the FFHQ (Flickr-Faces-Hight-Quality) dataset used by traditional image style conversion networks has 70,000 high-definition face images with a resolution of 1024×1024 pixels. The AFHQ (Animal Faces Hight Quality) dataset used by traditional image style conversion networks has animal face images with a resolution of 512×512 pixels. The MS-COCO (Microsoft Common Objects in Context) dataset used by traditional image style conversion networks has 118,000 training images. Such data volumes are required to support large models to train better results.
[0035] Traditional style transfer networks consider the style in the content and style images as a whole, learn the overall distribution of the style domain, and add the style to the content through a neural network to complete style transfer. Generative adversarial networks can obtain different types of content and style features in the content domain and style domain through adversarial learning between the generator and the discriminator, and fuse them respectively to obtain a better and more hierarchical style transfer effect. These two methods have their own advantages and disadvantages. In the scheme of using traditional style transfer networks, relatively good texture transfer effects can be obtained through learning with relatively few samples, but the generated image style is more holistic and the hierarchy is relatively blurred; in the scheme of using generative adversarial networks, high-quality image styles with strong layering can be generated, but there are certain requirements for the number of training samples. It can be seen that there are two difficulties in training a high-quality style transfer network: 1) being able to extract content and style features from few samples; 2) being able to fuse the main and object features of content and style to achieve a hierarchical effect.
[0036] In related technologies, there are already schemes to achieve high-quality style transfer through a large amount of learning. For example, image style transfer can be achieved based on StyleGAN. Although the StyleGAN network is very versatile and provides many possibilities for generating high-quality images, StyleGAN requires a huge amount of computation. For individuals, it is extremely difficult to adjust its network structure, and the cost of retraining the network is also relatively high.
[0037] In view of the above technical problems existing in related technologies, the technical solutions of the embodiments of the present application are proposed. The embodiments of the present application can be applied to scenarios such as artificial intelligence, image processing, style transfer, art creation, and dataset generation. In the embodiments of the present application, a small-sample training image style transfer network can be used to generate high-quality images through the trained image style transfer network, which has the advantage of low training cost. Small-sample learning can quickly generalize to new tasks with only a small number of samples by using prior knowledge or some augmentation methods. The main implementation methods of small-sample learning include reinforcement learning, transfer learning, meta-learning, etc., and the most widely used method is meta-learning.
[0038] Exemplarily, the ideas of few-shot learning generally include the following: The first idea is to guide the model to generate images of corresponding styles by adjusting the network hidden feature representation mode. The methods adopted include but are not limited to modifying the hidden features using additional information and directly manipulating the hidden feature dimensions; the second idea is to achieve style transfer by gradually fine-tuning the weights of different layers of the pre-trained model; the third idea is to achieve zero-shot style transfer by modifying the surface features for style-semantic alignment; the fourth idea is to perform style transfer by using the method of generative image processing with GAN; the fifth idea is to perform standard sample image style transfer by augmenting the few-shot feature space, which has the characteristics of being concise and easy to implement. The embodiments of the present application can adopt the above fifth idea to implement the training of the image style transfer network and the style transfer of images.
[0039] The following further details the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the examples provided herein are only used to explain the embodiments of the present application and are not used to limit the embodiments of the present application. In addition, the examples provided below are partial examples for implementing the present application, rather than providing all the embodiments for implementing the present application. Without conflict, the technical solutions described in the embodiments of the present application can be implemented in any combined manner.
[0040] It should be noted that in the embodiments of the present application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a method or device including a series of elements not only includes the elements specifically recited, but also includes other elements not explicitly listed, or further includes elements inherent to the implementation of the method or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of other related elements in the method or device including the element (such as steps in the method or units in the device, and the unit can be part of a circuit, part of a processor, part of a program or software, etc.).
[0041] The image style transfer method provided by the embodiments of the present application includes a series of steps, but the image style transfer method provided by the embodiments of the present application is not limited to the recited steps. Similarly, the image style transfer device provided by the embodiments of the present application includes a series of modules, but the device provided by the embodiments of the present application is not limited to including the explicitly recited modules, and may further include modules required for obtaining relevant information or processing based on the information.
[0042] Figure 1 is a flowchart of an image style transfer method according to an embodiment of the present application, as Figure 1 shown, this process includes:
[0043] Step 101: Extract content image features and style image features from the image to be processed, and obtain the first semantic information corresponding to the content image features and the second semantic information corresponding to the style image features.
[0044] In the embodiments of the present application, the image to be processed can be collected by a photographing device, or obtained from the storage space of an electronic device, or obtained through a network.
[0045] In some embodiments of the present application, one implementation of this step is:
[0046] Use multiple content feature extraction layers to extract features from the image to be processed, obtain the output features of each content feature extraction layer in the multiple content feature extraction layers, and obtain semantic features having the same size as the output features of each content feature extraction layer; the content image features include the output features of each content feature extraction layer, and the first semantic information includes the semantic features corresponding to the output features of each content feature extraction layer.
[0047] Use multiple style feature extraction layers to extract features from the image to be processed, obtain the output features of each style feature extraction layer in the multiple style feature extraction layers, and obtain semantic features having the same size as the output features of each style feature extraction layer; the style image features include the output features of each style feature extraction layer, and the second semantic information includes the semantic features corresponding to the output features of each style feature extraction layer.
[0048] Here, the multiple content feature extraction layers can be network layers in a Visual Geometry Group (VGG) network or other networks, and the multiple style feature extraction layers can be network layers in a VGG network or other networks. The network layers in the VGG network can include Rectified Linear Unit (ReLU); exemplarily, this step can be implemented by a feature data encoding module, and the feature data encoding module can include the above-mentioned multiple content feature extraction layers and multiple style feature extraction layers.
[0049] Exemplarily, the image to be processed can be used as an input image and input to the first content feature extraction layer of the feature data encoding module. The first content feature extraction layer is used to process the image to be processed to obtain the output features of the first content feature extraction layer; after obtaining the output features of the first content feature extraction layer, the first semantic features corresponding to the output features of the first content feature extraction layer can be obtained.
[0050] Input the output features of the first content feature extraction layer into the second content feature extraction layer of the feature data encoding module, and use the second content feature extraction layer to process the image to be processed to obtain the output features of the second content feature extraction layer; after obtaining the output features of the second content feature extraction layer, perform downsampling on the first semantic features to obtain the second semantic features corresponding to the output features of the second content feature extraction layer.
[0051] Input the output features of the second content feature extraction layer into the third content feature extraction layer of the feature data encoding module, and use the third content feature extraction layer to process the image to be processed to obtain the output features of the third content feature extraction layer; after obtaining the output features of the third content feature extraction layer, perform downsampling on the second semantic features to obtain the semantic features corresponding to the output features of the third content feature extraction layer.
[0052] Input the output features of the third content feature extraction layer into the fourth content feature extraction layer of the feature data encoding module, and use the fourth content feature extraction layer to process the image to be processed to obtain the output features of the fourth content feature extraction layer; after obtaining the output features of the fourth content feature extraction layer, perform downsampling on the third semantic features to obtain the semantic features corresponding to the output features of the fourth content feature extraction layer.
[0053] The content image features include the output features of the first content feature extraction layer to the fourth content feature extraction layer, and the first semantic information includes the first semantic feature to the fourth semantic feature.
[0054] Similarly, referring to the way of obtaining content image features, style image features can be obtained; referring to the way of obtaining the first semantic information, the second semantic information can be obtained.
[0055] In the embodiments of the present application, by aligning the output features of different content feature extraction layers with the corresponding semantic features, pairs of content image features and semantic features can be formed; by aligning the output features of different style feature extraction layers with the corresponding semantic features, pairs of style image features and semantic features can be formed; according to the pairs of content image features and semantic features, and the pairs of style image features and semantic features, the network can be accurately guided to extract and transfer regional style features.
[0056] Step 102: Preprocess the content image features and style image features according to the first semantic information and the second semantic information to obtain the preprocessed content image features and the preprocessed style image features.
[0057] Step 103: Recode the preprocessed style image features to obtain the recoded style image features.
[0058] In some embodiments, a Variational Auto Encoder (VAE) network can be used to re-encode the pre-processed style image features to obtain the re-encoded style image features. By re-encoding the style image features, it is possible to augment the style feature space without increasing the style images.
[0059] Step 104: Fuse the pre-processed content image features and the re-encoded style image features to obtain fused features.
[0060] In one example, the pre-processed content image features can be denoted as content features, and the re-encoded style image features can be denoted as style features. The semantic information of the content features is characterized by a content mask, and the semantic information of the style features is characterized by a style mask. Referring to Figure 2 , in a semantic-based feature calculation fusion layer, the content features and the style features can be fused according to the content mask and the style mask to obtain fused features.
[0061] Step 105: Decode the fused features to obtain a style transferred image.
[0062] In the embodiments of the present application, referring to Figure 2 , the fused features can be decoded by a decoder to obtain an output image, and the output image is the style transferred image described above.
[0063] In some embodiments, the decoder can include N upsampling modules, where N is an integer greater than 1. The decoder can be obtained by replacing the pooling modules of the VGG19 network with upsampling modules. The implementation manner of this step is: using multiple upsampling modules to process the fused features in sequence to obtain a style transferred image.
[0064] Here, each content feature extraction layer (or each style feature extraction layer) corresponds to an upsampling module. The order of the content feature extraction layers in the feature data encoding module is opposite to the order of the upsampling modules in the decoder. Specifically, when the number of content feature extraction layers in the feature data encoding module is 4, the upsampling module corresponding to the 4th content feature extraction layer in the feature data encoding module is the 1st upsampling module arranged in sequence in the decoder, the upsampling module corresponding to the 3rd content feature extraction layer in the feature data encoding module is the 3rd upsampling module arranged in sequence in the decoder, the upsampling module corresponding to the 2nd content feature extraction layer in the feature data encoding module is the 3rd upsampling module arranged in sequence in the decoder, and the upsampling module corresponding to the 1st content feature extraction layer in the feature data encoding module is the 4th upsampling module arranged in sequence in the decoder.
[0065] Exemplarily, the fused features can be input into the first upsampling module, and the first upsampling module processes the fused features to obtain the output features of the first upsampling module; after obtaining the output features of the first upsampling module, semantic features having the same size as the output features of the first upsampling module can be acquired.
[0066] The output features of the first upsampling module are input into the second upsampling module, and the second upsampling module processes the fused features to obtain the output features of the second upsampling module; after obtaining the output features of the second upsampling module, the semantic features corresponding to the output features of the first upsampling module are upsampled to obtain semantic features having the same size as the output features of the second upsampling module.
[0067] The output features of the second upsampling module are input into the third upsampling module, and the third upsampling module processes the fused features to obtain the output features of the third upsampling module; after obtaining the output features of the third upsampling module, the semantic features corresponding to the output features of the second upsampling module are upsampled to obtain semantic features having the same size as the output features of the third upsampling module.
[0068] The output features of the third upsampling module are input into the fourth upsampling module, and the fourth upsampling module processes the fused features to obtain the output features of the fourth upsampling module; after obtaining the output features of the fourth upsampling module, the semantic features corresponding to the output features of the third upsampling module are upsampled to obtain semantic features having the same size as the output features of the fourth upsampling module.
[0069] The output image of the decoder can be obtained according to the output features of the fourth upsampling module.
[0070] In some embodiments, semantic information for evaluating the image migration quality of the style transfer image can also be generated according to the output features of each upsampling module in each upsampling module. Exemplarily, the semantic information for evaluating the image migration quality of the style transfer image may include: semantic features corresponding to the output features of each upsampling module. It can be seen that the embodiments of the present application can accurately generate semantic information for image migration quality by combining the output features of each upsampling module, so that the evaluation of the style transfer effect can be more easily achieved.
[0071] Exemplarily, in order to optimize the upsampling effect of the upsampling module, the bilinear interpolation upsampling module in the related art can be changed to a dynamic upsampling module. In the dynamic upsampling module, the parameters of upsampling are dynamically adjusted through a non-linear parameter adjustment method, so as to better perform upsampling. Refer to Figure 3, each dynamic upsampling module may include: a two-dimensional transposed convolution (ConvTranspose2d) layer, a two-dimensional convolution (Conv2d) layer, and a rectified linear unit; exemplarily, in the two-dimensional transposed convolution layer, the size of the convolution kernel is 2, the stride of the convolution kernel is 2, and the padding is 0; in the two-dimensional convolution layer, the size of the convolution kernel is 1, the stride of the convolution kernel is 1, and the padding is 0.
[0072] In practical applications, steps 101 to 105 may be implemented based on a processor, and the above-mentioned processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor.
[0073] It can be seen that the embodiments of the present application can augment the style feature space without increasing the number of samples by re-encoding the style image features, and there is no need to collect a large amount of sample data to improve the learning effect of the generative adversarial network, thus improving the utilization efficiency of image features.
[0074] In some embodiments of the present application, the process of preprocessing the content image features and the style image features according to the first semantic information and the second semantic information to obtain the preprocessed content image features and the preprocessed style image features includes:
[0075] Semantically align the first semantic information and the second semantic information to obtain the aligned semantic information; pre-fuse the content image features and the style image features according to the aligned semantic information to obtain the pre-fused features; extract the content features and the image features from the pre-fused features to obtain the preprocessed content image features and the preprocessed style image features.
[0076] In the embodiments of the present application, an asymmetric semantic preprocessing module may be introduced. The asymmetric semantic preprocessing module can automatically semantically align the first semantic information and the second semantic information to obtain the aligned semantic information, and the content image features and the style image features can be fused through the aligned semantic information, thereby generating the pre-fused features with semantically symmetric information.
[0077] To address the generation of unexpected calculation processes, the embodiments of this application have improved the way of introducing semantic information. By analyzing the asymmetry between content semantics and style semantics, a completely symmetric semantic information (i.e., the aligned semantic information) is matched and generated. And through this semantic information graph, the content image features and style image features are fused to form a pre-fused feature.
[0078] The pre-fused feature S can be represented by formula (1).
[0079]
[0080] Where, represents the part corresponding to semantic i of the pre-fused feature S. In the case of dual semantics, the dual semantics are denoted as semantic 0 and semantic 1. When the second semantic information corresponding to the style image feature lacks semantic 0, the part corresponding to semantic 1 of the pre-fused feature S and the part corresponding to semantic 0 of the pre-fused feature S can be represented by formula (2) and formula (3).
[0081]
[0082] Where, represents the part corresponding to semantic 1 in the second semantic information of the style image feature, represents the part corresponding to semantic 0 in the first semantic information of the content image feature.
[0083] When the second semantic information corresponding to the style image feature lacks semantic 1, the part corresponding to semantic 1 of the pre-fused feature S and the part corresponding to semantic 0 of the pre-fused feature S can be represented by formula (4) and formula (5).
[0084]
[0085] Where, represents the part corresponding to semantic 1 in the first semantic information of the content image feature, represents the part corresponding to semantic 0 in the second semantic information of the style image feature.
[0086] In some embodiments, after obtaining the preprocessed content image features and the preprocessed style image features, the third semantic information corresponding to the preprocessed content image features and the fourth semantic information corresponding to the preprocessed style image features may also be obtained; the process of fusing the preprocessed content image features and the re-encoded style image features may include: fusing the preprocessed content image features and the re-encoded style image features according to the third semantic information and the fourth semantic information to obtain the fused features.
[0087] It can be seen that the embodiments of the present application can pre-align the semantic information corresponding to the content image features and the style image features respectively, which can solve the problem of semantic asymmetry in actual application scenarios, and is beneficial to obtaining the content image features and the style image features for realizing high-quality semantic style transfer of images based on the aligned semantic information.
[0088] In the embodiments of the present application, the AdaIN network can be improved. By calculating the regional feature information of different semantic regions respectively and fusing them respectively for semantic styleization. In addition, the method of the stochastic AdaIN network is borrowed, and the calculated style features are re-encoded using the feature re-encoding layer to augment the style feature space without increasing the number of samples, so as to optimize the style transfer effect.
[0089] The style fusion formula of AdaIN is as follows in formula (6):
[0090]
[0091] Among them, c represents the content image features, s represents the style image features, σ(·) represents calculating the standard deviation, and μ(·) represents calculating the channel mean.
[0092] The loss function L of AdaIN AdaIN can be represented by formula (7).
[0093] L AdaIN = L c + λ s L s (7)
[0094] Among them, L c represents the content loss, L s represents the style loss, and λ s represents a hyperparameter that controls the relative weight of the content loss and the style loss.
[0095] In the related art, the style fusion formula of the RAIN network is as follows in formula (8):
[0096]
[0097] Among them, f(·) represents the data processing operation of the variational autoencoder network.
[0098] The loss function L of the RAIN network RAIN can be represented by formula (9).
[0099] L RAIN = L c + λ s L s + λ k L KL + λ r L Rec (9)
[0100] Among them, L KL represents the KL divergence between the features of the image without image style transfer and the features of the image after image style transfer, L Rec represents the image reconstruction loss, and λ k is a hyperparameter representing the weight of L KL , and λ r is a hyperparameter representing the weight of L Rec .
[0101] In the embodiments of the present application, AdaIN and the RAIN network can be combined to achieve the fusion of image features. The part T i (c, s, M c , M s ) corresponding to semantic i in the fusion features can be calculated according to formula (10).
[0102]
[0103] Among them, M c represents the third semantic information corresponding to the content image features, and M s represents the second semantic information corresponding to the style image features. represents the standard deviation of the part corresponding to semantic i in the second semantic information in the preprocessed style image features. represents the channel mean of the part corresponding to semantic i in the second semantic information in the preprocessed style image features. represents the standard deviation of the part corresponding to semantic i in the third semantic information in the preprocessed content image features. represents the channel mean of the part corresponding to semantic i in the third semantic information in the preprocessed content image features.
[0104] In the case of dual semantics, the fusion feature T(c, s, M c , M s ) can be calculated according to formula (11).
[0105] T(c,s,M c ,M s )=M c ×T 1 (c,s,M c ,M s )+(1-M c )×T 0 (c,s,M c ,M s ) (11)
[0106] It can be seen that, through the above formula (10) and formula (11), the preprocessed content image features and the re-encoded style image features can be fused to obtain fused features, which can be used for decoding by the decoder.
[0107] In some embodiments, steps 104 to 105 may be implemented by a pre-trained neural network, see Figure 2 , when the image to be processed is training data of a neural network, the loss of the neural network can be calculated based on the content features, the content mask, the style features, the style mask and the above output image.
[0108] In the embodiment of the present application, when the above-described semantic-based feature calculation fusion layer is used for feature fusion in actual applications, the neural network will not be trained by only calculating the loss of AdaIN. It needs to be adaptively improved to reduce the features within the semantic area to retain more style features.
[0109] The embodiment of the present application proposes a loss function that includes the fusion of local features and global features. The loss of the neural network includes content loss and style loss. In order to preserve the global content information and regional content information of the content image as much as possible, a new content loss function L is proposed. pc , L pc It can be calculated by formula (12) to formula (14).
[0110]
[0111] L identity =1-similarity(classify(g(c)),classify(g(t))) (13)
[0112] L pc =L p +λ identity L identity (14)
[0113] Among them, t is the image generated by the decoder, Mmse(·) represents the Masked Mean Squared Error, and the Masked Mean Squared Error can be obtained by calculating the mean squared error loss of different semantic parts of the image and adding them up with weights. The meaning of g; represents the part corresponding to semantic i in the content image features after preprocessing and the Masked Mean Squared Error of g. Refer to Figure 2 , the output image, content features, content mask, style features, and style mask of the decoder can be input into the Masked Mean Squared Error module, and the Masked Mean Squared Error module processes its own input data to obtain L p .
[0114] L identity is the identity similarity loss, and similarity(classify(g(c)),classify(g(t))) represents the cosine similarity between the type of the content image features after preprocessing and the type of t. Refer to Figure 2 , a VGG classifier can be used to classify the content image features after preprocessing and t, and then calculate the identity similarity loss L identity ; λ identity is the hyperparameter representing the weight of L identity .
[0115] A new style loss function L ps is proposed in the embodiment of the present application, and L ps can be calculated by formula (15).
[0116]
[0117] Among them, represents the standard deviation of the part corresponding to semantic j in the output features of the (i + 1)-th upsampling module, represents the standard deviation of the part corresponding to semantic j in the output features of the (i + 1)-th style feature extraction layer; represents the channel mean of the part corresponding to semantic j in the output features of the (i + 1)-th upsampling module, represents the channel mean of the part corresponding to semantic j in the output features of the (i + 1)-th style feature extraction layer; mse(·) represents calculating the mean squared error, and λ j represents the weight parameter corresponding to semantic j. Refer to Figure 2 , the output image, style features, and style mask of the decoder can be used to calculate the value of L ps .
[0118] Figure 4 is the flowchart of another image style transfer method according to the embodiment of the present application, as shown in Figure 4As shown, the process includes:
[0119] Step 401: Obtain content image features, style image features, first semantic information, and second semantic information.
[0120] Step 402: Pre-fuse the content image features and the style image features to obtain pre-fused features.
[0121] Step 403: Obtain pre-processed content image features, pre-processed style image features, third semantics, and fourth semantic information.
[0122] Step 404: Re-encode the pre-processed style image features to obtain re-encoded style image features.
[0123] Step 405: Fuse the pre-processed content image features and the re-encoded style image features to obtain fused features.
[0124] Step 406: Decode the fused features through a decoder to obtain a style transfer image.
[0125] The implementation manners of Steps 401 to 406 have been described in the foregoing content and will not be elaborated here.
[0126] In summary, the embodiment of the present application redesigned on the basis of the Random AdaIN network, added a mask for distinguishing different semantics of images, and realized semantic partition fusion of image styles by performing semantic separation calculations on image features based on the mask. Then, the fused features are decoded by a decoder, and then the decoded image is separately calculated for semantic separation loss with the content image and the style image to optimize the decoding effect. By repeating this process, high-quality style transfer images can be generated. In addition, in general semantic partition fusion scenarios, semantic asymmetry may occur. In response to this situation, the embodiment of the present application further proposes a targeted and somewhat general asymmetric semantic image adaptive style transfer method to solve the semantic asymmetry problem that appears in actual application scenarios and achieve high-quality semantic style transfer of images, thereby providing technical support for related application fields.
[0127] Embodiments of the present application can utilize a semantic image style transfer method to finally fuse and decode image features according to semantics, integrate and utilize semantic information during the image style transfer process, reduce the occupation and consumption of computing resources, improve the efficiency of image feature utilization, and optimize the effect of the style transfer image. Further, embodiments of the present application are adaptively modified based on the random AdaIN method according to the image and its semantic information, extract the fused features and their semantic information, convert the features into a form convenient for neural network calculation, increase the information density of the feature matrix, and achieve the unity of image features and their semantic features, improve the computing efficiency, and enhance the effect of style transfer through the neural network. Further, embodiments of the present application preprocess the asymmetric parts in the semantic features through an asymmetric semantic preprocessing module to prevent contamination of the features during the neural network calculation process, realize asymmetric semantic style transfer, and reduce the probability of manual intervention.
[0128] Embodiments of the present application utilize an improved AdaIN method to process image features and their semantic information, and use an asymmetric semantic processing module to handle the situation of semantic asymmetry, realizing the image style transfer of asymmetric semantics, innovatively integrating image semantic features into the image processing process, effectively utilizing the semantic information of the image, and reducing the probability of manual intervention. Therefore, embodiments of the present application have the advantages of high feature utilization rate and good processing effect in image processing and image style transfer. Embodiments of the present application can be applied to the method of artificial intelligence image processing, can effectively improve the effect of image processing, and improve user satisfaction.
[0129] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0130] Figure 5 It is a structural schematic diagram of the image style transfer device according to an embodiment of the present application, as Figure 5 shown, the device includes:
[0131] A first processing module 501, configured to extract content image features and style image features from the image to be processed, and obtain first semantic information corresponding to the content image features and second semantic information corresponding to the style image features;
[0132] The second processing module 502 is configured to preprocess the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain a preprocessed content image feature and a preprocessed style image feature; re-encode the preprocessed style image feature to obtain a re-encoded style image feature; and fuse the preprocessed content image feature and the re-encoded style image feature to obtain a fused feature.
[0133] The decoding module 503 is configured to decode the fused feature to obtain a style transferred image.
[0134] In some embodiments, the second processing module 502 is configured to preprocess the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain a preprocessed content image feature and a preprocessed style image feature, including:
[0135] Semantically align the first semantic information and the second semantic information to obtain aligned semantic information;
[0136] Pre-fuse the content image feature and the style image feature according to the aligned semantic information to obtain a pre-fused feature;
[0137] Extract the content feature and the image feature from the pre-fused feature to obtain the preprocessed content image feature and the preprocessed style image feature.
[0138] In some embodiments, the second processing module 502 is further configured to, after obtaining the preprocessed content image feature and the preprocessed style image feature, acquire third semantic information corresponding to the preprocessed content image feature and fourth semantic information corresponding to the preprocessed style image feature;
[0139] In some embodiments, the second processing module 502 is configured to fuse the preprocessed content image feature and the re-encoded style image feature to obtain a fused feature, including:
[0140] Fuse the preprocessed content image feature and the re-encoded style image feature according to the third semantic information and the fourth semantic information to obtain a fused feature.
[0141] In some embodiments, the second processing module 502 is configured to re-encode the preprocessed style image feature to obtain a re-encoded style image feature, including: re-encoding the preprocessed style image feature by using a variational autoencoder network to obtain a re-encoded style image feature.
[0142] In some embodiments, the first processing module 502 is configured to extract content image features and style image features from the image to be processed, and obtain first semantic information corresponding to the content image features and second semantic information corresponding to the style image features, including:
[0143] Using multiple content feature extraction layers to extract features from the image to be processed, obtaining output features of each content feature extraction layer in the multiple content feature extraction layers, and obtaining semantic features having the same size as the output features of each content feature extraction layer; the content image features include the output features of each content feature extraction layer, and the first semantic information includes semantic features corresponding to the output features of each content feature extraction layer;
[0144] Using multiple style feature extraction layers to extract features from the image to be processed, obtaining output features of each style feature extraction layer in the multiple style feature extraction layers, and obtaining semantic features having the same size as the output features of each style feature extraction layer; the style image features include the output features of each style feature extraction layer, and the second semantic information includes semantic features corresponding to the output features of each style feature extraction layer.
[0145] In some embodiments, the decoding module 503 is configured to decode the fused features to obtain a style transfer image, including:
[0146] Using multiple upsampling modules to process the fused features in sequence to obtain the style transfer image;
[0147] The method further includes: generating semantic information for evaluating the image transfer quality of the style transfer image according to the output features of each upsampling module in the multiple upsampling modules.
[0148] In practical applications, the first processing module 501, the second processing module 502, and the decoder 503 may be implemented based on a processor.
[0149] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0150] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0151] Correspondingly, the embodiments of the present application further provide a computer program product, and the computer program product includes computer-executable instructions for implementing any one of the image style transfer methods provided by the embodiments of the present application.
[0152] Correspondingly, the embodiments of the present application further provide a computer storage medium, and computer-executable instructions are stored on the computer storage medium for implementing any one of the image style transfer methods provided by the above embodiments.
[0153] The embodiments of the present application also provide an electronic device. Figure 6 As a schematic diagram of the composition structure of an electronic device provided by the embodiments of the present application, as Figure 6 shown, the electronic device 60 may include:
[0154] A memory 601 for storing executable instructions;
[0155] A processor 602 for implementing any one of the above image style transfer methods when executing the executable instructions stored in the memory 601.
[0156] The above processor 602 may be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.
[0157] The above computer-readable storage medium, the memory 602 can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0158] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0159] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated here.
[0160] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0161] The features disclosed in the product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0162] The features disclosed in the method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0164] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.
Claims
1. A method for image style transfer, characterized in that: The method comprises: Extracting content image features and style image features from the image to be processed, and obtaining first semantic information corresponding to the content image features and second semantic information corresponding to the style image features; Preprocessing the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain a preprocessed content image feature and a preprocessed style image feature; Re-encoding the pre-processed style image features to obtain re-encoded style image features; Fusing the preprocessed content image features and the re-encoded style image features to obtain fused features; The fused features are decoded to obtain a style transfer image.
2. The method according to claim 1, characterized in that: The preprocessing of the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain the preprocessed content image feature and the preprocessed style image feature includes: Performing semantic alignment on the first semantic information and the second semantic information to obtain aligned semantic information; Pre-fusing the content image features and the style image features according to the aligned semantic information to obtain pre-fused features; Content features and image features are extracted from the pre-fused features to obtain the pre-processed content image features and the pre-processed style image features.
3. The method according to claim 2, characterized in that After obtaining the preprocessed content image feature and the preprocessed style image feature, the method further includes: obtaining third semantic information corresponding to the preprocessed content image feature and fourth semantic information corresponding to the preprocessed style image feature; The fusing the preprocessed content image features and the re-encoded style image features to obtain fused features includes: The preprocessed content image feature and the re-encoded style image feature are fused according to the third semantic information and the fourth semantic information to obtain a fused feature.
4. The method according to any one of claims 1 to 3, characterized in that: The re-encoding the pre-processed style image features to obtain the re-encoded style image features includes: re-encoding the pre-processed style image features using a variational autoencoder network to obtain the re-encoded style image features.
5. The method according to any one of claims 1 to 3, characterized in that: The step of extracting content image features and style image features from the image to be processed, and obtaining first semantic information corresponding to the content image features and second semantic information corresponding to the style image features includes: Using multiple content feature extraction layers to extract features from the image to be processed, obtaining output features of each of the multiple content feature extraction layers, and obtaining semantic features having the same size as the output features of each content feature extraction layer; the content image features include the output features of each content feature extraction layer, and the first semantic information includes semantic features corresponding to the output features of each content feature extraction layer; Using multiple style feature extraction layers to extract features from the image to be processed, obtaining output features of each of the multiple style feature extraction layers, and obtaining semantic features having the same size as the output features of each style feature extraction layer; the style image features include the output features of each style feature extraction layer, and the second semantic information includes semantic features corresponding to the output features of each style feature extraction layer.
6. The method according to any one of claims 1 to 3, characterized in that: The decoding of the fusion feature to obtain a style transfer image includes: Using multiple upsampling modules to process the fusion features in sequence to obtain the style transfer image; The method further includes: generating semantic information for evaluating image migration quality of the style migration image according to output features of each upsampling module in the multiple upsampling modules.
7. An image style transfer device, characterized in that: The device comprises: A first processing module, configured to extract content image features and style image features from the image to be processed, and obtain first semantic information corresponding to the content image features and second semantic information corresponding to the style image features; A second processing module is used to preprocess the content image feature and the style image feature according to the first semantic information and the second semantic information to obtain preprocessed content image features and preprocessed style image features; re-encode the pre-processed style image features to obtain re-encoded style image features; and fuse the pre-processed content image features and the re-encoded style image features to obtain a fused feature; The decoding module is used to decode the fusion features to obtain a style transfer image.
8. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is used to run the computer program to execute the image style transfer method according to any one of claims 1 to 6.
9. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image style transfer method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the image style transfer method according to any one of claims 1 to 6.
Citation Information
Cited By
Interactive icon automatic generation system based on artificial intelligence
CN120672895A